Power load prediction method based on improved DWT-LSTM
Through the improved DWT-LSTM method, the accuracy and reliability problems caused by the nonlinear relationship of factors in power load prediction are solved, and high-precision and stable power load prediction are achieved, supporting the optimized operation and sustainable development of the power system.
Patent Information
- Application Number
- CN202510283808.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is difficult to achieve high accuracy and high reliability in power load prediction, especially due to the complex nonlinear relationships of factors such as climatic conditions and social activities, which leads to the deviation of the prediction results and poor data quality.
The improved DWT-LSTM method is used to decompose the power load data through discrete wavelet transformation, combine the long and short-term memory network LSTM model, and the LSTM model is constructed using the input variables after DWT decomposed, and the model weight is adjusted through an optimization algorithm to minimize prediction errors.
It improves the accuracy and robustness of power load prediction, can identify power consumption patterns and electricity consumption changes in advance, ensure the stability of power supply, optimize the combination of clean energy and traditional energy, reduce carbon emissions, and support the efficient operation and sustainable development of power companies.
Smart Images

Figure BDA0005306542640000032 
Figure BDA0005306542640000033 
Figure BDA0005306542640000034
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power technology, and in particular to a power load forecasting method based on improved DWT-LSTM. Background Technique
[0002] Power load forecasting is the basis for the optimal operation of power systems, aiming to predict future power demands by analyzing historical power data and related influencing factors. With the rapid development of new energy and the complexity of power markets, traditional forecasting methods have been difficult to meet the high-precision requirements. This has made the application of advanced technologies such as machine learning in power load forecasting more prominent.
[0003] Power load is affected by various factors, including environmental conditions (such as temperature and humidity), social activities (such as holidays and festivals), and economic factors. Machine learning methods such as long short-term memory network (LSTM) and support vector machine (SVM) can be used to capture these complex relationships. These models can effectively process non-linear and time-series data, making the prediction of power load more accurate.
[0004] In addition, the diversity and integrity of data are crucial for prediction results. An effective data collection mechanism needs to cover multi-dimensional information, such as historical power load data, environmental parameters, and social event records. By constructing a rich feature set and combining advanced machine learning algorithms, high-performance power load forecasting can be achieved in different geographical and social environments. Combining the comprehensive analysis of various influencing factors can provide strong support for the sustainable development of power systems.
[0005] In the prior art, in the research of power load forecasting, power load is affected by various factors, including climate conditions, social activities, etc., and there are complex non-linear relationships between these factors, making it impossible to achieve high-precision and high-reliability forecasting. Moreover, the data quality is poor, resulting in the prediction results deviating further from the actual situation. And power load data has time series characteristics, with time lag characteristics and potential seasonality and trends. Therefore, the load fluctuates greatly over time, resulting in lower prediction accuracy.
[0006] The patent with the publication number CN114219139B discloses a DWT-LSTM power load forecasting method based on the attention mechanism. The discrete wavelet decomposition method is used to decompose the collected original load data to obtain load components of different scales. The attention mechanism is introduced to adaptively assign weights according to the importance of each scale load component, completing the preprocessing of the load data. The improved PSO particle swarm algorithm is used to optimize the parameters of the LSTM long short-term memory neural network model to obtain an optimized LSTM model. The weighted load components are respectively substituted into the optimized LSTM model for training to obtain the LSTM load forecasting models of each load component. The load component obtained after the attention mechanism processing is used as the input, and the load forecasting value of each component is obtained by inputting it into the LSTM load forecasting model of the corresponding component. Then, the load forecasting values of each component are accumulated, which is the forecasting value of the power load at the next moment. It assigns weights according to the importance of each scale load, but it lacks the collection of power load-related variables and has limitations in variable data processing. Summary of the Invention
[0007] The present invention proposes a power load forecasting method based on improved DWT-LSTM, which solves the problem in the prior art research of power load forecasting that power load is affected by various factors, including climate conditions, social activities, etc., and there are complex non-linear relationships among these factors, making it impossible to achieve high-precision and high-reliability forecasting.
[0008] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0009] A power load forecasting method based on improved DWT-LSTM includes the following steps:
[0010] Step S1: Collect power load data and its related input variables, including temperature, humidity, and weather conditions. Divide the collected original data into a training set and a test set, and perform data cleaning.
[0011] Step S2: Use the DWT method to decompose the cleaned input variables, decompose the signal into approximation coefficients and detail coefficients. The approximation coefficients represent the low-frequency components of the signal, and the detail coefficients capture the high-frequency changes of the signal.
[0012] Step S3: Combine the input variables after DWT decomposition to construct a long short-term memory network LSTM model. The LSTM model structure design includes an input layer, an LSTM layer, and a fully connected layer.
[0013] Step S4: Use the coefficient of each frequency band obtained by DWT decomposition as the input of the LSTM model. Use the training set to train the LSTM model, and adopt an optimization algorithm to adjust the model weights to minimize the prediction error.
[0014] Step S5: The model training is completed. The power load is predicted using the test set. According to the prediction results output by the LSTM model, they are merged with the frequency levels decomposed by DWT and converted into continuous data.
[0015] Further, in step S1, let the power load data be a time series {y1, y2, …, y N}, where y t represents the power load at the t-th moment; the temperature, humidity, and weather conditions are represented by , where m is the number of input variables, and the data set is the observations of N time steps. The data set D can be expressed as:
[0016] D = {(y1, x1), (y2, x2), …, (y N , x N )}
[0017] The data set D is divided into a training set D train and a test set D test , as follows:
[0018]
[0019] where N train is the serial number of the last group of data in the training set; the training set D train and the test set D test are subjected to data cleaning. The data cleaning process includes: missing value processing, outlier detection and removal, and data standardization.
[0020] Further, in step S2, the DWT method is used to decompose the cleaned input variables. A low-pass filter h[n] is used to filter the signal to obtain the low-frequency component (approximate coefficient):
[0021]
[0022] A high-pass filter g[n] is used to filter the signal to obtain the high-frequency component (detail coefficient):
[0023]
[0024] where a t and d t represent the low-frequency component and the high-frequency component at the t-th moment respectively. The low-frequency component is the approximate coefficient, and the high-frequency component is the detail coefficient. x n is the n values of the input sequence.
[0025] Further, the low-frequency component obtained from the first decomposition is filtered and downsampled again. Let the low-frequency component obtained from the first decomposition be Continue to perform secondary decomposition on it:
[0026] Perform a second low-pass filtering on the first low-frequency part to obtain a new low-frequency component:
[0027]
[0028] Perform a second high-pass filtering on the low-frequency part of the first layer to obtain a new high-frequency component:
[0029]
[0030] Among them is the low-frequency component obtained after the first decomposition of the nth input variable; it can be further decomposed, and the signal is decomposed into low-frequency and high-frequency components at multiple different frequency levels;
[0031] Approximation coefficient a t represents the overall trend of the signal, and the detail coefficient d t captures the detailed changes of the signal.
[0032] Furthermore, for the LSTM model in step S3, given time t, the calculation process of the LSTM model is as follows:
[0033] Forget gate: Controls which information in the previous moment's memory C t-1 needs to be discarded, and the output f of the forget gate t is calculated through the activation function Sigmoid:
[0034] f t = σ(W f · [h t-1 , x t + b f )
[0035] Among them, W f is the weight matrix of the forget gate, b f is the bias term, σ is the Sigmoid activation function, h t-1 is the hidden state of the previous time step, and x t is the input of the current time step;
[0036] Input gate: Controls the update degree of the input information at the current moment, and the output i of the input gate t is calculated through the Sigmoid function:
[0037] i t = σ(W i · [h t-1 , x t + b i )
[0038] Generate a candidate value simultaneously to update the memory:
[0039]
[0040] Memory cell update: Update the state of the memory cell according to the outputs of the forget gate and the input gate:
[0041]
[0042] where, C t is the state of the memory cell at the current time step;
[0043] Output gate: Control the output h t at the current moment and determine which information is passed to the next time step. The output o t of the output gate is calculated through the Sigmoid function:
[0044] o t = σ(W o · [h t-1 , x t + b o )
[0045] Hidden state h t is calculated through the following formula:
[0046] h t = o t · tanh(C t )
[0047] where, h t is the output at the current moment, and C t is the memory state at the current moment.
[0048] Furthermore, the structure of the LSTM model includes the following parts:
[0049] Input layer: Receive the data processed by DWT. Assume that the dimension of the input features is m, that is, the input data at each time step is an m-dimensional vector, then the size of the input layer is m;
[0050] LSTM layer: Composed of multiple LSTM cells. Each LSTM cell performs calculations according to the calculation process. The output of the LSTM layer is a sequence of hidden states h t containing time series information, with a length of T;
[0051] Fully connected layer: After the LSTM layer, add a fully connected layer to map the output of the LSTM to the final predicted value. This layer outputs a scalar representing the predicted power load:
[0052]
[0053] Among them, W y is the weight matrix of the fully connected layer, and b y is the bias term, and h t is the output of the LSTM layer.
[0054] Furthermore, the LSTM model uses the mean squared error MSE as the loss function for optimization, and the expression of the loss function is:
[0055]
[0056] Among them, is the predicted value of the model, y t is the true power load value, and T is the number of samples;
[0057] The gradient descent method is used in the optimization process, and the Adam optimizer updates the weights and biases of the model:
[0058]
[0059] Among them, θ t is the parameter of the model, η is the learning rate, is the gradient of the loss function with respect to the model parameters.
[0060] Furthermore, in step S4, the coefficient of each frequency band obtained by DWT decomposition is used as the input of the LSTM model. The signal is decomposed into low-frequency and high-frequency components at multiple different frequency levels. Let the coefficient of the frequency band obtained after DWT decomposition be:
[0061]
[0062] Among them, is the low-frequency component (approximation coefficient) of the first layer, is the high-frequency component (detail coefficient) of the i-th layer, and J is the number of decomposition layers; the input X t at each time step is a multi-dimensional vector representing the coefficients of each frequency band at this time point;
[0063] For each time step t, the input of the LSTM model is:
[0064]
[0065] Among them, X t is the input feature vector at time t, which contains signal features at multiple frequency levels;
[0066] The data is divided into small batches for training. Suppose there is data for N time steps, and the data dimension at each time step is m, that is, the number of frequency band coefficients. Then the data set is represented as an N×m matrix;
[0067] The data is divided into B batches, and each batch contains N b samples. The input matrix X of each batch batch has the shape of N b ×m; The specific form of batch data preparation is:
[0068]
[0069] During the training process, the input data X batch will be fed into the LSTM network for processing;
[0070] The training process of the LSTM model is based on the backpropagation method and the gradient descent method. The mean squared error MSE is used as the loss function, and this loss function is used to measure the difference between the predicted value and the true value of the model. The calculation formula is:
[0071]
[0072] Among them, is the predicted value of the model, y t is the true power load value, and T is the total number of samples;
[0073] The goal of training the LSTM model is to minimize the loss function L, and the parameters of the model are updated by the gradient descent method:
[0074]
[0075] Among them, θ t are the parameters of the LSTM model (such as the weight matrix and the bias term), η is the learning rate, is the gradient of the loss function with respect to the model parameters.
[0076] Furthermore, during the training process of the LSTM model, the Adam optimizer is used, and the update formula of the Adam optimizer is:
[0077]
[0078] Among them, m t and v t are the first-order moment estimate and the second-order moment estimate of the gradient respectively, β1 and β2 are the decay factors, η is the learning rate, and ∈ is a constant to prevent division by zero errors.
[0079] Furthermore, after the model training in step S5 is completed, the power load is predicted, and the process is as follows:
[0080] At the t-th moment, the input data is as follows:
[0081]
[0082] Among them, represents the first low-frequency (approximate) coefficient, is the high-frequency (detail) coefficient of the i-th layer. Through the forward propagation process of the LSTM model, the predicted power load value at each time point is obtained That is:
[0083]
[0084] Among them, f(X t , θ) represents the mapping function of the LSTM network, θ is the weight and bias in the LSTM model, and X t is the input feature at the t-th moment, is the predicted power load value output by the model;
[0085] Perform the inverse DWT. Let the frequency band coefficients in the prediction process be and Among them is the approximate coefficient, is the detail coefficient, and the inverse transformation formula is:
[0086]
[0087] Among them, J is the number of layers of DWT decomposition, is the low-frequency approximate coefficient, is the high-frequency detail coefficient of the i-th layer, is the predicted power load value after inverse transformation. Sum the high-frequency detail coefficients and low-frequency approximate coefficients of each layer to restore the original signal of the time series;
[0088] Perform prediction result time series restoration and inverse transformation multi-step prediction.
[0089] The positive effect of the present invention is that when using this method for power load prediction, the model extracts complex features in power load data by combining multi-level discrete wavelet transform DWT and long short-term memory network LSTM. Further, on specific holidays and in special electricity consumption scenarios, the electricity demand will change continuously under different scenarios. Through the improved DWT-LSTM model, it can predict and identify the electricity consumption patterns and changes in electricity consumption in various scenarios in advance. The power company can allocate the electricity load and electricity consumption in advance to ensure the stable operation of the power grid.
[0090] The model can detect peak electricity consumption periods related to holidays, enabling power companies to more accurately predict electricity demand on the eve of the Spring Festival. Consequently, power companies can adjust their power generation plans in advance, increase backup power sources, avoid potential power shortages and user power outages during holidays, and improve the stability of power operation.
[0091] When dealing with extreme weather, the improved DWT-LSTM model also has a certain degree of adaptability. During cold snaps, electricity demand often surges, especially in the early morning and evening. The model can identify the rapid growth trend of electricity load during cold snaps by analyzing historical load data and meteorological data. In actual operation, power companies can increase power generation capacity in a timely manner based on the model's prediction results to ensure the stability of power supply when a cold snap arrives.
[0092] In addition, this method can also assist power companies in making more scientific decisions regarding the integration of renewable energy. Through accurate prediction of electricity load, power companies can reasonably arrange the connection times of wind energy and solar energy, and optimize the combination of clean energy and traditional energy. This optimization can improve the overall efficiency of the power system, reduce carbon emissions, and promote the achievement of sustainable development goals.
[0093] The improved DWT-LSTM electricity load forecasting method can demonstrate certain technical effects in practical applications. By improving forecasting accuracy, enhancing robustness, and optimizing resource allocation, it provides support for the efficient operation and sustainable development of power companies. Detailed implementation manners
[0094] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0095] In the research of electricity load forecasting, a series of basic technical problems need to be solved to achieve high-precision and high-reliability forecasting. First of all, it is to model the non-linear relationship between electricity load and various external factors. Since electricity load is affected by multiple factors, including climate conditions, social activities, etc., and there are complex non-linear relationships among these factors. Therefore, it is particularly important to select a suitable model, such as the combination model of long short-term memory network LSTM and discrete wavelet transform DWT-LSTM, to capture these characteristics.
[0096] Secondly, considering the time series characteristics of power load data, it is necessary to effectively handle time lag features such as hours, days, and weeks to identify potential seasonality and trends. Feature identification can enhance the input dimension and information volume of the prediction model by extracting and constructing relevant features of environmental factors, social events, and historical load data.
[0097] Data quality and integrity also affect the prediction accuracy. To ensure data reliability, issues such as missing values and outliers need to be addressed to maintain data consistency; real-time prediction ability is a necessary requirement for a power load prediction system, and a balance needs to be achieved between computational efficiency and prediction accuracy.
[0098] Embodiment 1
[0099] A power load prediction method based on improved DWT-LSTM includes the following steps:
[0100] Step S1: Collect power load data and its related input variables, including temperature, humidity, and weather conditions. Divide the collected raw data into a training set and a test set, and perform data cleaning.
[0101] First, collect power load data and its related input variables, including temperature, humidity, and weather conditions. To ensure the generalization ability of the model on unseen data, the collected raw data is divided into a training set and a test set. The data cleaning process includes handling missing values and removing outliers, and at the same time performing standardization to improve data quality.
[0102] In step S1, let the power load data be a time series {y1, y2, …, y N}, where y t represents the power load at the t-th moment; in addition, external variables related to the power load, temperature, humidity, and weather conditions, need to be collected, and these variables are represented by , where m is the number of input variables, and the data set is the observations of N time steps. The data set D can be expressed as:
[0103] D = {(y1, x1), (y2, x2), …, (y N , x N )}
[0104] For model training and evaluation, the data set D is divided into a training set D train and a test set D test , usually divided in a ratio of 70%:30%, as follows:
[0105]
[0106] Among them, N train is the serial number of the last group of data in the training set; for the training set Dtrain and the test set D test Perform data cleaning, and the data cleaning process includes: missing value handling, outlier detection and removal, and data standardization.
[0107] For the missing values in the data, mean filling or interpolation methods are used for processing. If some data points are missing values, they can be filled with the mean of this column, and the formula is as follows:
[0108] If is missing, use mean filling
[0109] In addition, linear interpolation can also be used to complete the missing values, that is, infer the missing values through the data points at adjacent times:
[0110] Perform linear interpolation on the missing values.
[0111] The outliers in the data may affect the training effect of the model, so outlier detection and removal are required. According to statistical methods, calculate the Z-score of the data:
[0112] μ is the mean, σ is the standard deviation
[0113] When |Z t |>3, the data point x t is considered an outlier and can be removed or corrected.
[0114] In addition, the IQR (interquartile range) method can be used to detect outliers. Assume that the first quartile of the data set is Q1 and the third quartile is Q3, then IQR is:
[0115] IQR = Q3 - Q1
[0116] When the data point x t is less than Q1 - 1.5×IQR or greater than Q3 + 1.5×IQR, it is considered an outlier and needs to be removed or processed.
[0117] To improve the stability of model training, the data needs to be standardized. Perform Z-score standardization on all input variables and the formula is as follows:
[0118]
[0119] Among them, μ j and σ jThey are the mean and standard deviation of the j-th input variable respectively. The standardized data helps to accelerate the training process of the LSTM model and prevent some variables from having too much impact on the model due to their large scales.
[0120] After data cleaning, missing value handling, outlier removal, and standardization, the processed dataset D train and D test are obtained. These two datasets can be used for subsequent DWT decomposition and LSTM model training.
[0121] Step S2: Use the DWT method to decompose the cleaned input variables. The signal is decomposed into approximation coefficients and detail coefficients. The approximation coefficients represent the low-frequency components of the signal, while the detail coefficients capture the high-frequency variations of the signal.
[0122] Use the DWT method to decompose the cleaned input variables. DWT decomposes the signal into multiple frequency levels through high-pass (HP) and low-pass (LP) filters, thereby extracting different frequency band information of the signal. This process generates multiple frequency bands, usually five, which represent low-frequency and high-frequency features respectively.
[0123] Specifically, the signal is decomposed into approximation coefficient a and detail coefficient d, where the approximation coefficient represents the low-frequency component of the signal, and the detail coefficient captures the high-frequency variations of the signal.
[0124] The discrete wavelet transform DWT is an effective method for decomposing signals through filters, aiming to extract different frequency band information of the signal. It is particularly suitable for analyzing non-stationary signals such as power loads. DWT uses high-pass and low-pass filters to decompose the signal, and the decomposed signal contains low-frequency components - approximation coefficients and high-frequency components - detail coefficients.
[0125] Let the input signal be x = {x1, x2, …, x N}, where N is the length of the signal. DWT processes the signal in two steps: first, high-pass and low-pass filtering, and then downsampling the filtered signal.
[0126] In step S2, use the DWT method to decompose the cleaned input variables. Use the low-pass filter h[n] to filter the signal to obtain the low-frequency component - approximation coefficient:
[0127]
[0128] Use the high-pass filter g[n] to filter the signal to obtain the high-frequency component (detail coefficient):
[0129]
[0130] where, a t and dt respectively represent the low-frequency component and the high-frequency component at the t-th moment. The low-frequency component is the approximation coefficient, and the high-frequency component is the detail coefficient. x n is the n values of the input sequence.
[0131] DWT is decomposed in a recursive manner. The low-frequency component obtained from the first decomposition is filtered and downsampled again. Let the low-frequency component obtained from the first decomposition be Continue to perform a secondary decomposition on it:
[0132] Second low-pass filtering, perform low-pass filtering on the first low-frequency part to obtain a new low-frequency component:
[0133]
[0134] Second high-pass filtering, perform high-pass filtering on the low-frequency part of the first layer to obtain a new high-frequency component:
[0135]
[0136] where is the low-frequency component obtained after the first decomposition of the n-th input variable; this recursive decomposition can continue until the signal is decomposed into low-frequency and high-frequency components at multiple different frequency levels;
[0137] The approximation coefficient a t represents the overall trend of the signal, and the detail coefficient d t captures the detailed changes of the signal.
[0138] Assume that after J-layer decomposition, the signal will be decomposed into multiple frequency bands where is the low-frequency component after the first low-pass filtering, is the detail coefficient after the i-th layer of decomposition.
[0139] To improve the accuracy and efficiency of the power load forecasting model, the following optimization measures are taken during the DWT process:
[0140] In practical applications, choosing an appropriate wavelet basis is crucial for the decomposition effect. Commonly used wavelet bases include Haar wavelet, Daubechies wavelet, etc. For power load signals, using Daubechies wavelet (such as DB4) can often better capture the high-frequency components and local features of the signal.
[0141] Let the selected Daubechies wavelet be ψ(t), and its discretized form is:
[0142]
[0143] Among them, φ(t) is the scaling function and ψ(t) is the wavelet function. By selecting an appropriate wavelet basis, the accuracy and stability of signal decomposition can be improved.
[0144] The decomposition level J of DWT has an important impact on the prediction result. The higher the level, the more frequency bands after decomposition, which can capture the details of the signal more accurately, but at the same time, it will also increase the computational complexity. According to the actual needs, the optimal decomposition level J can be selected by methods such as cross-validation.
[0145] After decomposing the power load data by DWT, the characteristics of each frequency level can be obtained, and more abundant information can be provided for the training of the LSTM model.
[0146] Specifically, the DWT process can:
[0147] Extract multi-scale features of the signal: By decomposing the low-frequency and high-frequency components of different frequency bands, DWT can effectively capture the long-term trend and short-term fluctuations of the signal.
[0148] Reduce noise interference: High-frequency noise is usually decomposed into detail coefficients, and the LSTM model can ignore these detail coefficients to reduce the interference of noise.
[0149] Enhance prediction accuracy: Using the coefficients of each frequency band obtained by DWT decomposition as the input of LSTM improves the learning ability and prediction accuracy of the model.
[0150] Step S3: Combine the input variables after DWT decomposition to construct a long short-term memory network (LSTM) model. The structure design of the LSTM model includes an input layer, an LSTM layer, and a fully connected layer; the LSTM network has the advantage of processing time series data and can effectively capture the time-dependent relationship between input variables. By adjusting the number of units and time steps in the LSTM layer, the learning ability of the model is optimized.
[0151] The long short-term memory network (LSTM) is a special type of recurrent neural network (RNN) that can effectively capture and learn long-term dependencies in time series data. Compared with traditional RNNs, LSTM solves the problems of vanishing gradients and exploding gradients by introducing three gating mechanisms (input gate, forget gate, and output gate), so it has significant advantages in processing time series data such as power load.
[0152] The basic unit of the LSTM network includes an input gate, a forget gate, and an output gate, which control the flow of information in the network. The core structure of LSTM can be expressed by the following formula:
[0153] For the LSTM model in step S3, given time t, the calculation process of the LSTM model is as follows:
[0154] Forget Gate: Controls the memory C from the previous moment t-1 to determine which information needs to be discarded. The output f of the forget gate t is calculated through the activation function Sigmoid:
[0155] f t = σ(W f · [h t-1 , x t + b f )
[0156] where W f is the weight matrix of the forget gate, b f is the bias term, σ is the Sigmoid activation function, h t-1 is the hidden state of the previous time step, and x t is the input of the current time step;
[0157] Input Gate: Controls the update degree of the input information at the current moment. The output i of the input gate t is calculated through the Sigmoid function:
[0158] i t = σ(W i · [h t-1 , x t + b i ), and b i is the bias term of i t ;
[0159] Meanwhile, a candidate value is generated to update the memory:
[0160] b C is 's bias term;
[0161] Memory Cell Update: Updates the state of the memory cell according to the outputs of the forget gate and the input gate:
[0162]
[0163] where C t is the state of the memory cell at the current time step;
[0164] Output Gate: Controls the output h t at the current moment and determines which information is passed to the next time step. The output o of the output gate t is calculated through the Sigmoid function:
[0165] o t = σ(W o · [h t-1 , xt +b o ), b o is o t 's bias term;
[0166] The hidden state h t is calculated by the following formula:
[0167] h t = o t ·tanh(C t )
[0168] where h t is the output at the current time step, and C t is the memory state at the current time step.
[0169] The structure of the LSTM model includes the following parts:
[0170] Input layer: Receives the data processed by DWT. Assuming the input feature dimension is m, that is, the input data at each time step is an m-dimensional vector, then the size of the input layer is m;
[0171] LSTM layer: Consists of multiple LSTM cells. Each LSTM cell performs calculations according to the calculation process. The output of the LSTM layer is a sequence of hidden states h t , with a length of T;
[0172] Fully connected layer: After the LSTM layer, a fully connected layer is added to map the output of the LSTM to the final predicted value. This layer outputs a scalar representing the predicted power load:
[0173]
[0174] where W y is the weight matrix of the fully connected layer, b y is the bias term, and h t serves as the output of the LSTM layer here.
[0175] The training process of the LSTM model is carried out through the backpropagation algorithm. The goal is to minimize the prediction error. The LSTM model usually uses the mean squared error MSE as the loss function for optimization. The expression of the loss function is:
[0176]
[0177] where is the predicted value of the model, y t is the true power load value, T is the number of samples, and is the same value as the length T of the above hidden state sequence h t ;
[0178] The optimization process uses the gradient descent method, and the Adam optimizer updates the weights and biases of the model:
[0179]
[0180] where θ t is the parameter of the model, η is the learning rate, and is the gradient of the loss function with respect to the model parameters.
[0181] When training the LSTM model, the reasonable selection of hyperparameters is crucial for the performance of the model. Common hyperparameters include:
[0182] Number of LSTM units: Determines the size of the LSTM layer and affects the learning ability and computational complexity of the model.
[0183] Learning rate η: Affects the step size of gradient update. Too high a learning rate may cause the model to diverge, while too low a learning rate may result in a too slow convergence speed.
[0184] Batch size: The number of data samples used for each gradient update. A smaller batch size may make the model more stable but increase the training time; a larger batch size can speed up the training but may reduce the generalization ability of the model.
[0185] These hyperparameters are usually tuned through methods such as cross-validation to obtain the best model configuration.
[0186] After training, it is necessary to evaluate the prediction performance of the LSTM model. Commonly used evaluation metrics include:
[0187] Mean Squared Error (MSE): Measures the difference between the predicted value and the true value.
[0188]
[0189] Root Mean Squared Error (RMSE): Taking the square root of MSE can provide a more intuitive quantification of the error.
[0190]
[0191] Mean Absolute Percentage Error (MAPE): Measures the percentage error between the predicted value and the true value.
[0192]
[0193] The advantages of LSTM in power load forecasting are mainly reflected in the following aspects:
[0194] Long-Short Term Dependence Capturing Ability: LSTM can effectively capture long-term dependencies in power load data and is suitable for processing time-series data with strong temporal characteristics and complex variations such as power load.
[0195] Robustness: During the training process, LSTM can automatically select and learn effective features and has strong robustness to data noise and variations.
[0196] Strong Adaptability: LSTM can flexibly adapt to changes at different time scales and achieve a balance between high-frequency fluctuations and long-term trends.
[0197] Step S4: Use the coefficient of each frequency band obtained by DWT decomposition as the input of the LSTM model to ensure that the input data can reflect the frequency characteristics and time variations of the signal. Use the training set to train the LSTM model and adopt an optimization algorithm to adjust the model weights to minimize the prediction error. During the training process, use cross-validation technology to ensure the stability and robustness of the model;
[0198] After the processing of Discrete Wavelet Transform (DWT), the original power load data and related input variables have been decomposed into multiple frequency bands. To improve the prediction effect of the LSTM model, the coefficient of each frequency band after DWT decomposition is used as the input of the LSTM model. Specifically, DWT decomposes each time-series data into low-frequency (approximate coefficient) and high-frequency (detail coefficient) components at multiple frequency levels. These components contain different scale information of the signal, thus providing rich input features for the LSTM model.
[0199] In step S4, use the coefficient of each frequency band obtained by DWT decomposition as the input of the LSTM model. The signal is decomposed into low-frequency and high-frequency components at multiple different frequency levels. Let the coefficient of the frequency band obtained after DWT decomposition be:
[0200]
[0201] Among them, is the low-frequency component (approximate coefficient) of the first layer, is the high-frequency component (detail coefficient) of the i-th layer, and J is the number of decomposition layers; the input X t at each time step is a multi-dimensional vector representing the coefficient of each frequency band at this time point;
[0202] For each time step t, the input of the LSTM model is:
[0203]
[0204] Among them, X t is the input feature vector at time t, which contains signal features at multiple frequency levels;
[0205] To improve the training efficiency and stability of the LSTM model, the data is usually divided into small batches for training. Suppose there are N time steps of data, and the data dimension at each time step is m, that is, the number of frequency band coefficients. Then the data set is represented as an N×m matrix;
[0206] The data is divided into B batches, and each batch contains N b samples. The input matrix X of each batch batch has the shape of N b ×m; The specific form of batch data preparation is:
[0207]
[0208] During the training process, the input data X batch will be fed into the LSTM network for processing;
[0209] The training process of the LSTM model is based on the backpropagation method and the gradient descent method. Through the backpropagation algorithm, the gradient of the loss function with respect to the model parameters is calculated, and the optimization algorithm is used to update the model parameters, so as to minimize the prediction error of the model.
[0210] If the mean squared error MSE is used as the loss function, this loss function is used to measure the difference between the model prediction value and the true value. The calculation formula is:
[0211]
[0212] where, is the prediction value of the model, y t is the true power load value, and T is the total number of samples;
[0213] The goal of LSTM model training is to minimize the loss function L, and the parameters of the model are updated by the gradient descent method:
[0214]
[0215] where, θ t are the parameters of the LSTM model (such as the weight matrix and bias term), η is the learning rate, is the gradient of the loss function with respect to the model parameters.
[0216] The choice of optimization algorithm in the LSTM model training process has an important impact on the training effect. Commonly used optimization algorithms include Stochastic Gradient Descent (SGD), Adam optimizer, etc.
[0217] The Adam optimizer is a currently widely used optimization algorithm. It combines the advantages of the momentum method and the adaptive learning rate, and can accelerate the convergence speed and improve the training stability.
[0218] During the training process of the LSTM model, the Adam optimizer is used. The update formula of the Adam optimizer is as follows:
[0219]
[0220] where, m t and v t are the first and second moment estimates of the gradient respectively, β1 and β2 are decay factors, η is the learning rate, and ∈ is a constant to prevent division by zero errors.
[0221] In addition, hyperparameters such as Batch Size, Learning Rate, Number of LSTM Units, etc. also affect the training effect. These hyperparameters are usually tuned through methods such as cross-validation.
[0222] During the training process, the LSTM model may exhibit overfitting, resulting in good performance on the training set but poor performance on the test set. To prevent overfitting, the following strategies can be adopted:
[0223] Dropout: By randomly discarding a portion of neurons during the training process, reduce the network's dependence on specific neurons and avoid overfitting.
[0224] L2 regularization: Add an L2 regularization term to the loss function to constrain the magnitude of the model's weights and prevent the model from overfitting the training data.
[0225] The regularized loss function is:
[0226]
[0227] where, λ is the regularization coefficient and θ i are the model parameters.
[0228] During the training process of the LSTM model, a validation set can be used to evaluate the training effect of the model in real-time. The validation set is used to monitor the change in error during the training process to determine whether the model has overfitted. If the error of the validation set starts to increase during the training process, it indicates that overfitting may have occurred. At this time, the training can be stopped or the hyperparameters can be adjusted.
[0229] Common evaluation metrics include Mean Squared Error (MSE), Root Mean Squared Error (RMSE), Mean Absolute Percentage Error (MAPE), etc. These metrics help evaluate the performance of the model on the test set and thus judge the generalization ability of the model.
[0230] After the model training is completed, the test set is used to predict the power load. According to the prediction results output by the LSTM model, they are combined with the frequency levels decomposed by DWT and converted into a series of continuous data. This inverse transformation process ensures that the prediction results can reflect the changes in the actual power load in the time series.
[0231] Step S5: After the model training is completed, the test set is used to predict the power load. According to the prediction results output by the LSTM model, they are combined with the frequency levels decomposed by DWT and converted into continuous data.
[0232] In the training of step S4, the LSTM model has been trained and learned to extract temporal patterns from the input power load data and related features. Through the trained LSTM model, in step S5: Prediction and Inverse Transformation, this model is used to predict the power load.
[0233] When the model training in step S5 is completed, the prediction of the power load is carried out as follows:
[0234] Suppose at time t, the input data is:
[0235]
[0236] where, represents the first low-frequency (approximate) coefficient, is the high-frequency (detail) coefficient of the i-th layer. Through the forward propagation process of the LSTM model, the predicted power load values at each time point are obtained That is:
[0237]
[0238] where, f(X t ,θ) represents the mapping function of the LSTM network, θ is the weights and biases in the LSTM model, X t is the input feature at time t, is the predicted power load value output by the model;
[0239] After being predicted by the LSTM model, the predicted values we obtain are the combined results of the coefficient of each frequency band of the power load. In order to restore to the original power load time series, an inverse DWT transformation is required.
[0240] The key to the inverse transformation process lies in restoring each frequency band coefficient to a continuous power load value through weighted combination.
[0241] For a time step t, the predicted value output by the LSTM model can be restored to the predicted value of the original power load sequence through the inverse transformation formula.
[0242] Perform the inverse DWT. Let the frequency band coefficients in the prediction process be and where is the approximation coefficient, is the detail coefficient, and the inverse transformation formula is:
[0243]
[0244] where J is the number of layers of DWT decomposition, is the low-frequency approximation coefficient, is the high-frequency detail coefficient of the i-th layer, is the predicted power load value after the inverse transformation. Sum the high-frequency detail coefficients and low-frequency approximation coefficients of each layer to restore the original signal of the time series;
[0245] Perform the restoration of the predicted result time series and inverse transformation multi-step prediction.
[0246] Example 2
[0247] The difference between this example and Example 1 is that: perform the restoration of the predicted result time series and inverse transformation multi-step prediction.
[0248] To restore the predicted power load time series, we can predict multiple future time steps step by step. In the recursive prediction process, the model uses the predicted value of the previous moment as the input for the next moment. Therefore, the prediction process can be represented as an iterative process.
[0249] Recursive prediction: For time t, we use the model to predict the power load at time t + 1 and use this predicted value as the input for time t + 1 to continue the next prediction. That is:
[0250]
[0251] This recursive method may introduce a certain error in each step of the model prediction. Therefore, it is usually necessary to adjust and optimize the predictions of multiple time steps.
[0252] Direct prediction: Different from recursive prediction, the direct prediction strategy is to directly output the load prediction values for a certain period in the future through the LSTM model. This method does not rely on step-by-step prediction but obtains the prediction results for a period of time directly. For example, predict the power load for the next T pred time steps:
[0253]
[0254] This method can reduce the error accumulation in the recursive process but requires the model to have strong generalization ability.
[0255] For the case of multi-step prediction, an inverse transformation needs to be performed on each prediction result. Suppose we make a prediction for T pred steps, and the inverse transformation process is as follows:
[0256]
[0257] where, is the predicted power load value at time t + k, is the detail coefficient at time t + k, is the approximation coefficient at time t + k.
[0258] Once the power load prediction value is restored through the inverse transformation, the accuracy of the prediction result needs to be evaluated. Common evaluation metrics include mean squared error (MSE), root mean squared error (RMSE), mean absolute percentage error (MAPE), etc. The specific calculation formulas are as follows:
[0259] Mean squared error (MSE):
[0260]
[0261] where, is the predicted value, y t is the actual power load value, and T is the number of samples.
[0262] Root mean squared error (RMSE):
[0263]
[0264] Mean absolute percentage error (MAPE):
[0265]
[0266] Through these evaluation metrics, the accuracy of the LSTM model in predicting power load can be effectively measured.
[0267] To further analyze the prediction results, the power load values predicted by the model can be compared with the actual power load values and visualized through charts. Common visualization methods include:
[0268] Time series plot: Plot the predicted values and actual values in chronological order on the same chart to observe the change trends and deviations between the two.
[0269] Error distribution plot: Show the distribution of prediction errors and analyze the concentration degree and distribution characteristics of the errors.
[0270] Residual plot: Show the distribution of prediction residuals, and the residuals usually represent prediction errors.
[0271] Through these charts, the prediction ability of the model can be intuitively demonstrated, which helps to further optimize the model.
[0272] Example 3
[0273] The difference between this example and Example 2 is that the prediction performance of the model is evaluated using indicators such as the mean squared error (MSE), root mean squared error (RMSE), and mean absolute percentage error (MAPE). The superiority of the proposed DWT-LSTM method is verified by comparison with the benchmark model.
[0274] To effectively evaluate the prediction performance of the proposed improved DWT-LSTM power load forecasting method, appropriate evaluation indicators need to be selected. Common evaluation indicators include:
[0275] Mean Squared Error (MSE)
[0276] Root Mean Squared Error (RMSE)
[0277] Mean Absolute Percentage Error (MAPE)
[0278] These evaluation indicators can measure the difference between the prediction results and the actual values, thereby reflecting the accuracy, robustness, and stability of the model.
[0279] The mean squared error (MSE) is one of the most commonly used indicators for evaluating the prediction performance of regression models. MSE measures the prediction accuracy of the model by calculating the average of the sum of the squares of the differences between the predicted values and the actual values.
[0280] For a sample set T = {t1, t2, …, t n}, the error ∈ between the prediction result t of the model and the actual value y t is calculated as follows:
[0281]
[0282] where is the predicted value and y t is the actual value. The formula for the mean squared error (MSE) is:
[0283]
[0284] where n is the number of samples, is the predicted value at the t-th moment, and y t is the actual power load value at the t-th moment. The smaller the MSE value, the higher the prediction accuracy of the model.
[0285] Root Mean Square Error (RMSE) is the square root of the mean square error and has the same unit as the predicted value. RMSE can effectively reflect the difference between the predicted value and the actual value and is often used to evaluate the prediction accuracy of a model.
[0286] The formula for RMSE is:
[0287]
[0288] where n is the number of samples, is the predicted value, and y t is the actual value. Different from MSE, RMSE has better interpretability because its unit is the same as the original data. The smaller the RMSE value, the smaller the prediction error of the model and the higher the prediction accuracy.
[0289] Mean Absolute Percentage Error (MAPE) is an index to measure the accuracy of prediction results. It calculates the percentage of the prediction error in the actual value and can intuitively show the ratio of the prediction error of the model to the actual value.
[0290] The calculation formula for MAPE is:
[0291]
[0292] where, is the predicted value at the t-th moment, y t is the actual value at the t-th moment, and n is the number of samples. The smaller the MAPE value, the smaller the prediction error of the model and the higher the prediction accuracy.
[0293] To comprehensively evaluate the performance of the model, in addition to using MSE, RMSE, and MAPE alone, the results of different evaluation metrics can also be comprehensively analyzed. For example, the weighted average of these three metrics can be calculated, or the most appropriate evaluation criteria can be selected according to application requirements.
[0294] The comprehensive evaluation formula is as follows:
[0295] Performance Score = w1 × MSE + w2 × RMSE + w3 × MAPE
[0296] where w1, w2, and w3 are the weight coefficients of MSE, RMSE, and MAPE respectively, and w1 + w2 + w3 = 1. Through the comprehensive score, a balanced consideration can be provided for different evaluation metrics.
[0297] To verify the superiority of the proposed DWT-LSTM model, it needs to be compared with other traditional prediction methods. Common benchmark models include:
[0298] ARIMA model (Autoregressive Integrated Moving Average model)
[0299] SVR model (Support Vector Regression model)
[0300] Neural network model (such as BP neural network)
[0301] Comparing evaluation metrics such as MSE, RMSE, and MAPE of these benchmark models can further prove the advantages of the proposed method in power load forecasting. For each benchmark model, calculate its prediction error and compare it with the DWT-LSTM model. The specific formulas are as follows:
[0302]
[0303] where Error Base and Error DWT-LSTM represent the error values of the benchmark model and the DWT-LSTM model respectively (which can be MSE, RMSE, or MAPE). The higher the improvement rate, the better the prediction performance of the DWT-LSTM model compared to the benchmark model.
[0304] In addition to the above evaluation metrics, the stability and robustness of the model are also important indicators for evaluating model performance. To test the stability of the model, the cross-validation method can be used for evaluation:
[0305] K-fold cross-validation: Divide the dataset into K subsets. Each time, use K - 1 of these subsets for training, and the remaining one subset for testing. By averaging the results of each training and testing, the stability of the model can be effectively evaluated.
[0306] Time series cross-validation: Different from K-fold cross-validation, time series cross-validation divides the data according to the time order to ensure that the model can handle the time-dependent relationships in time series data.
[0307] By testing the model on different subsets, the stability and robustness of the model can be evaluated, ensuring its effectiveness in practical applications.
[0308] To further analyze the evaluation results, the prediction errors are usually visualized through charts. These visualization charts include:
[0309] Error distribution chart: Displays the error distribution at each prediction moment, evaluating the concentration degree and distribution characteristics of the model errors.
[0310] Residual chart: Displays the distribution of prediction residuals. The residual represents the difference between the predicted value and the actual value. The analysis of residuals helps to detect the deviation of the model in certain specific situations.
[0311] Prediction Error Comparison Chart: Compare the results of the DWT-LSTM model and the benchmark model on multiple evaluation metrics, and clearly show the advantages and disadvantages of different models through bar charts or line charts.
[0312] Through the visualization of these charts, the model performance can be intuitively displayed and provide a basis for further improvement of the model.
[0313] Visualize and analyze the prediction results of the model. Compare the original signal with the predicted signal, observe the performance of the model and the accuracy of the prediction, and ensure the effectiveness of the model in power load prediction.
[0314] To more intuitively evaluate the effect of the power load prediction model based on the improved DWT-LSTM, it is first necessary to visualize the prediction results of the model. Visualization not only helps to identify the patterns of prediction errors but also provides inspiration for further optimization of the model. Common visualization methods include:
[0315] True Value vs. Predicted Value Comparison Chart:
[0316] In time series prediction, showing the predicted value and the actual value in the same chart helps to intuitively see the relationship between the prediction trend of the model and the actual data. The horizontal axis of this chart is time (such as hours, days, etc.), and the vertical axis is the power load value. The true value and the predicted value are plotted in the chart simultaneously to observe their closeness.
[0317] The formula for the comparison chart is:
[0318]
[0319] where is the predicted value at the t-th moment, and y t is the actual value at the t-th moment. Through the comparison chart, the error between the prediction and the actual value can be intuitively displayed.
[0320] Residual Analysis Chart:
[0321] The residual is the difference between the predicted value and the actual value. The residual chart is used to analyze the error distribution of the model. If the residual chart shows a random distribution, it means that the model can fit the data well; if there are obvious regular fluctuations, it indicates that there are certain biases in the model. The formula for the residual chart is:
[0322]
[0323] By analyzing the residual chart, the fitting situation of the model can be better judged.
[0324] Error Distribution Chart:
[0325] The error distribution diagram shows the distribution of errors (i.e., the differences between predicted values and actual values) at various times during the entire prediction process in different intervals. Usually, a histogram is used to display the frequency distribution of errors to help analyze the overall bias of the model. The formula for the frequency distribution of errors is:
[0326]
[0327] where δ is the Dirac function, is the error, and P(∈) is the frequency distribution of the error.
[0328] Through the analysis of the visualization chart, next, the prediction effect of the model is quantitatively evaluated, the deviation between the predicted value and the actual value is compared, and further analysis is carried out on different evaluation indicators.
[0329] Analysis of prediction errors:
[0330] By comparing the predicted values and actual values at various times, the statistical quantity of the overall error is calculated. For example, by calculating the mean and standard deviation of the errors, the overall performance of the model can be judged:
[0331] Mean:
[0332] Standard deviation:
[0333] where, is the predicted value, y t is the actual value, is the mean of the errors, and σ ∈ is the standard deviation of the errors.
[0334] Mean error (Bias): Analyze the systematic bias of the prediction error to judge whether the model has long-term overestimation or underestimation.
[0335]
[0336] If Bias is zero, it means the model prediction is relatively accurate; if it is positive or negative, it means the model tends to overestimate or underestimate the actual value.
[0337] Analysis of the temporal characteristics of errors:
[0338] Power load data has obvious temporal characteristics. The analysis of the temporal characteristics of errors helps to reveal the performance of the model at different time periods. For example, at certain times (such as when the load fluctuates greatly), the prediction error of the model may increase. The analysis of temporal characteristics can analyze the stability of the model at different time periods by calculating the average error of each time period (such as hourly, daily, or weekly):
[0339]
[0340] Among them, \(T\) i represents all time points in the \(i\)-th time period, and \(n\) i is the number of samples in this time period.
[0341] Robustness analysis of model performance under different conditions:
[0342] To verify the robustness of the model under different data conditions, the performance of the model under different input variable conditions can be analyzed. For example, predictions can be made based on datasets with different weather conditions (temperature, humidity, etc.), and the performance of the model under extreme weather conditions can be analyzed. Specifically, by changing certain features in the input data (such as when the temperature changes significantly), the changes in the model output can be observed to determine its robustness.
[0343] Comparative analysis - Comparison with the benchmark model:
[0344] For different evaluation metrics (such as MSE, RMSE, MAPE, etc.), the improved DWT - LSTM model is compared with other commonly used power load forecasting models (such as traditional LSTM, SVR, ARIMA, etc.) to evaluate the relative advantages of the proposed method. The aforementioned performance evaluation formulas can be used, such as:
[0345]
[0346] Through this comparison, the performance improvement of the DWT - LSTM model compared to other benchmark models can be quantified.
[0347] By combining the frequency characteristics of the signal with the time - series modeling ability of LSTM, the DWT - LSTM algorithm can effectively capture the dynamic changes of power load. This method can handle power load data with complex seasonality and volatility.
[0348] The embodiments described above are relatively detailed and specific, expressing the preferred embodiments of the present invention. They are only used to illustrate the technical ideas and characteristics of the present invention. The purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. However, it is not limited to the present invention only. The patent scope of the present invention cannot be limited only by this embodiment. That is, any equivalent changes or modifications made in accordance with the spirit disclosed by the present invention, for researchers or technicians in the field, within the structure of the present invention, local improvements within the system and changes and transformations between subsystems are still within the patent scope of the present invention.
Claims
1. A power load forecasting method based on improved DWT-LSTM, characterized in that, It includes the following steps: Step S1: Collect power load data and its related input variables, and perform data cleaning; Step S2: Use the DWT method to decompose the cleaned input variables, decompose the signal into approximation coefficients and detail coefficients. The approximation coefficients represent the low-frequency components of the signal, and the detail coefficients capture the high-frequency changes of the signal; Step S3: Combine the input variables after DWT decomposition to construct a long short-term memory network (LSTM) model. The LSTM model structure design includes an input layer, an LSTM layer, and a fully connected layer; Step S4: Use the coefficients of each frequency band obtained by DWT decomposition as the input of the LSTM model, use the training set to train the LSTM model, and adopt an optimization algorithm to adjust the model weights to minimize the prediction error; Step S5: After the model training is completed, use the test set to predict the power load. According to the prediction results output by the LSTM model, merge it with the frequency levels decomposed by DWT and convert it into continuous data.
2. The power load forecasting method based on improved DWT-LSTM according to claim 1, characterized in that In the step S1, the relevant input variables include temperature, humidity and weather conditions. Let the power load data be a time series {y1, y2, …, y N}, where y t represents the power load at the t-th moment; temperature, humidity and weather conditions are represented by , where m is the number of input variables. Define the data set D as the observations of N time steps. The data set D can be expressed as: D = {(y1, x1), (y2, x2), …, (y N , x N )} The collected original data is divided into a training set and a test set. The dataset D is divided into the training set D train and the test set D test as follows: Among them, N train is the serial number of the last group of data in the training set; for the training set D train and the test set D test perform data cleaning, and the data cleaning process includes: missing value processing, outlier detection and removal, and data standardization.
3. A power load forecasting method based on improved DWT-LSTM according to claim 1, characterized in that, In step S2, when using the DWT method to decompose the cleaned input variables, use a low-pass filter h[n] to filter the signal to obtain the low-frequency component - approximation coefficients: Use a high-pass filter g[n] to filter the signal to obtain the high-frequency component - detail coefficients: where a t and d t represent the low-frequency component and the high-frequency component at the t-th moment respectively. The low-frequency component is the approximation coefficient, and the high-frequency component is the detail coefficient. x n is the n values of the input sequence.
4. A power load forecasting method based on improved DWT-LSTM according to claim 3, characterized in that Filter and downsample the low-frequency component obtained from the first decomposition. Let the low-frequency component obtained from the first decomposition be Continue to perform a secondary decomposition on it: Perform a second low-pass filtering on the first low-frequency part to obtain a new low-frequency component: Perform a second high-pass filtering on the first-layer low-frequency part to obtain a new high-frequency component: wherein is the low-frequency component obtained after the first decomposition of the nth input variable; the decomposition can be continued, and the signal is decomposed into low-frequency and high-frequency components at multiple different frequency levels; Approximation coefficient a t Represents the overall trend of the signal, and the detail coefficient d t Captures the detailed changes of the signal.
5. The power load forecasting method based on improved DWT-LSTM according to claim 4, characterized in that For the LSTM model in step S3, given time t, the calculation process of the LSTM model is as follows: Forgotten Gate: Controls the memory C from the previous moment t-1 which information in t needs to be discarded, and the output f of the forgotten gate is calculated through the activation function Sigmoid: f t = σ(W f · [h t-1 , x t + b f ) Among them, W f is the weight matrix of the forget gate, b f is the bias term, σ is the Sigmoid activation function, h t-1 is the hidden state of the previous time step, x t is the input of the current time step; Input gate: Controls the degree of update of the input information at the current moment. The output of the input gate is i t Calculated through the Sigmoid function: i t = σ(W i · [h t-1 , x t + b i ) Simultaneously generate a candidate value to update the memory: Memory cell update: Update the state of the memory cell according to the outputs of the forget gate and the input gate: Among them, C t is the memory cell state at the current time step; Output gate: Controls the output h at the current time step t and determines which information is passed to the next time step. The output o of the output gate t is calculated through the Sigmoid function: o t = σ(W o · [h t-1 , x t + b o ) Hidden state h t Calculated by the following formula: h t = o t ·tanh(C t ) where h t is the output at the current moment, and C t is the memory state at the current moment.
6. The power load forecasting method based on improved DWT-LSTM according to claim 5, wherein, The structure of the LSTM model includes the following parts: Input layer: Receive the data processed by DWT. Assume that the input feature dimension is m, that is, the input data at each time step is an m-dimensional vector, then the size of the input layer is m; LSTM layer: It consists of multiple LSTM cells. Each LSTM cell performs calculations according to the calculation process. The output of the LSTM layer is a sequence of hidden states h that contains time series information. t , with a length of T; Fully connected layer: After the LSTM layer, a fully connected layer is added to map the output of the LSTM to the final predicted value, and this layer outputs a scalar Indicating the predicted power load: Among them, W y is the weight matrix of the fully connected layer, and b y is the bias term, and h t is the output of the LSTM layer.
7. A power load forecasting method based on improved DWT-LSTM according to claim 6, characterized in that, The LSTM model uses the mean squared error (MSE) as the loss function for optimization. The expression of the loss function is: Among them, is the predicted value of the model, and y t is the true power load value, and T is the number of samples; In the optimization process, use the gradient descent method and the Adam optimizer to update the weights and biases of the model: where θ t is a parameter of the model, η is the learning rate, is the gradient of the loss function with respect to the model parameter.
8. A power load forecasting method based on improved DWT-LSTM according to claim 7, characterized in that, In step S4, use the coefficients of each frequency band obtained by DWT decomposition as the input of the LSTM model. The signal is decomposed into low-frequency and high-frequency components of multiple different frequency levels. Let the frequency band coefficients obtained after DWT decomposition be: Among them, is the low-frequency component (approximation coefficient) of the first layer, is the high-frequency component (detail coefficient) of the i-th layer, and J is the number of decomposition layers; the input X at each time step t is a multi-dimensional vector representing the coefficients of each frequency band at this time point; For each time step t, the input of the LSTM model is: Among them, X t is the input feature vector at time t, which contains signal features at multiple frequency levels; Divide the data into small batches for training. Assume there are N time steps of data, and the data dimension at each time step is m, that is, the number of frequency band coefficients. Then the data set is represented as an N×m matrix; Divide the data into B batches, each batch containing N b samples, and the input matrix X of each batch batch is of shape N b ×m; The specific form of batch data preparation is as follows: During the training process, the input data X batch will be fed into the LSTM network for processing; The LSTM model training process is based on the backpropagation method and the gradient descent method, and uses the mean squared error (MSE) as the loss function. This loss function is used to measure the difference between the model prediction value and the true value. The calculation formula is: Among them, is the predicted value of the model, and y t is the true power load value, and T is the total number of samples; The goal of LSTM model training is to minimize the loss function L, and update the model parameters through the gradient descent method: Among them, θ t are the parameters of the LSTM model (such as the weight matrix and bias term), η is the learning rate, is the gradient of the loss function with respect to the model parameters.
9. A power load forecasting method based on improved DWT-LSTM according to claim 8, characterized in that The Adam optimizer is used during the LSTM model training process. The update formula of the Adam optimizer is: where m t and v t are the first and second moment estimates of the gradient respectively, β1 and β2 are decay factors, η is the learning rate, and ∈ is a constant to prevent division by zero errors.
10. A power load forecasting method based on improved DWT-LSTM according to claim 9, characterized in that, In step S5, after the model training is completed, predict the power load. The process is as follows: At the t-th moment, the input data is as follows: Among them, represents the first low-frequency (approximate) coefficient, is the high-frequency (detail) coefficient of the i-th layer. Through the forward propagation process of the LSTM model, the predicted power load value at each time point is obtained That is: Among them, f(X t , θ) represents the mapping function of the LSTM network, θ is the weights and biases in the LSTM model, and X t is the input feature at the t-th moment, and is the predicted power load value output by the model; Perform the inverse DWT. Let the frequency band coefficients in the prediction process be and where is the approximation coefficient, is the detail coefficient, and the inverse transform formula is: where J is the number of layers of DWT decomposition, is the low-frequency approximation coefficient, is the high-frequency detail coefficient of the i-th layer, is the predicted power load value after inverse transformation. The high-frequency detail coefficients and low-frequency approximation coefficients of each layer are summed to restore the original signal of the time series; Perform time series recovery of the prediction results and inverse transformation for multi-step prediction.
Citation Information
Patent Citations
DWT-LSTM Power Load Forecasting Method Based on Attention Mechanism
CN114219139B
Cited By
Multivariable fusion power load prediction method and system
CN120566430A
A multivariable fusion power load forecasting method and system
CN120566430B
Risk prediction method and system for large-scale server cluster operation and maintenance
CN121833357A