A Probability Prediction Method for Electric Power Consumption Based on Neural Network
By adopting a neural network model based on convolutional architecture and self-attention mechanism in power consumption prediction, the problem of difficulty in capturing time series characteristics and inability to quantify prediction results in the prior art is solved, and high-precision power consumption prediction and multi-user data processing are achieved.
Patent Information
- Application Number
- CN202210499887.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-09
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-05-09
AI Technical Summary
The prior art is difficult to effectively capture time series characteristics in power consumption prediction, cannot process electricity consumption data of multiple users at the same time, and the prediction results cannot be quantified.
A neural network model based on convolutional architecture and self-attention mechanism is adopted. Through the time convolutional network, gated residual network, embedded layer and multi-head attention mechanism module, a model that can process power consumption data of multiple users at the same time is built to achieve efficient capture of short-term and long-term modes of the time series.
This method can improve the accuracy of power consumption prediction, capture short-term and long-term patterns in the time series, quantify the uncertainty of prediction results, and process the electricity consumption data of multiple users at the same time, improving work efficiency.
Smart Images

Figure CN114819372B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power consumption probability prediction, and more specifically, to a power consumption probability prediction method based on a neural network. Background Art
[0002] In recent years, with the rapid development of deep learning, many methods for time series prediction based on neural networks have emerged one after another: the DeepAR model based on the long short-term memory network (LSTM), the Temporal Convolution Networks (TCN) based on the convolutional neural network (CNN), and the Transformer model based on the attention mechanism. These existing deep learning-based methods have achieved good prediction accuracy in time series prediction tasks.
[0003] There are also some problems with the above-mentioned deep learning-based time series prediction architectures. The recurrent neural network (RNN) has problems of gradient disappearance and gradient explosion, and the LSTM cannot model overly long sequences. For example, using the LSTM model can only clearly distinguish the information of about 50 nearby positions. In addition, the recursive structure of the recurrent neural network causes the model to be unable to be trained in parallel during training, resulting in low training efficiency. The Transformer architecture uses the attention mechanism to process sequence data, allowing the model to access any part of the historical data regardless of the distance, which makes it easier for the Transformer architecture to capture long-term repeating patterns in the data. However, the Transformer model consists of multiple layers of attention to form the encoder and decoder, and the computational complexity of attention is the square of the time series length L. The stacking of multiple layers of attention makes the computational complexity of the Transformer architecture much higher than that of other models. In addition, most energy time series data have short-term dependencies. Since the Transformer architecture allows the model to access any part of the historical data, it is not sensitive to capturing the short-term dependencies of time series and requires a large amount of data to be trained for a long time to discover the short-term dependencies in the time series. The convolutional network can efficiently capture short-term time patterns, but it is insufficient in capturing long-term time patterns.
[0004] In the prior art, a power load prediction method, system and storage medium based on deep learning are disclosed. Specifically, the method is as follows: collecting the power load data, meteorological data and air quality data of users within a preset historical time period, and dividing the collected data into a training set and a test set; determining a deep learning model for power load prediction; inputting the test set into the deep learning model for power load prediction to obtain the power load prediction data of the user within the third time interval. This solution uses deep learning for power load prediction, and in the deep learning process, not only power load data is considered, but also meteorological data and air quality data are considered, which can provide the accuracy of power load prediction and can be considered as the prior art closest to the present application. The defect of this solution is that it can only capture the short-term patterns of time series and cannot quantify the uncertainty of the prediction results.
[0005] Therefore, in combination with the above requirements and the defects of the prior art, the present application proposes a power consumption probability prediction method based on a neural network. Summary of the Invention
[0006] The present invention provides a power consumption probability prediction method based on a neural network to overcome the defects of the existing power consumption prediction technology, such as weak ability to capture time series features, inability to process the power consumption data of multiple users simultaneously, and inability to quantify the prediction results.
[0007] The primary object of the present invention is to solve the above technical problems, and the technical solution of the present invention is as follows:
[0008] A power consumption probability prediction method based on a neural network includes the following steps:
[0009] S1. Collect historical data of power consumption, divide it into a training set and a test set, and perform normalization processing on the variables in the historical data.
[0010] S2. Construct a neural network model based on a convolutional architecture and a self-attention mechanism.
[0011] S3. Use the processed training set data to train the neural network model, and select the model with the best prediction accuracy using the test set as the trained neural network model.
[0012] S4. Select recent data of power consumption and perform preprocessing, input the preprocessed recent data into the trained neural network model, and perform inverse normalization processing on the output value of the model to obtain the probability prediction result.
[0013] Further, the historical data of power consumption described in step S1 is collected in a time series manner, and the time series includes the time information corresponding to each time point in the time series, the power consumption of electricity customers within the same time interval, and the user number information of different electricity customers.
[0014] Among them, the time information includes year, month, week, hour, minute, and second.
[0015] Further, the step S1 includes:
[0016] S1-1: Collect historical data of power consumption, divide the first n% of the data into a training set, and the latter (1 - n)% of the data into a test set;
[0017] S1-2: Divide the historical data in the training set and the test set into several samples of a preset length. The preset length is (T + H). Use the data of the first T time points of each sample as the historical time series data input into the neural network model, and use the data of the latter H time points as the true value of the prediction result;
[0018] S1-3: Perform normalization processing on the time information and power consumption included in the historical data in the training set.
[0019] Among them, the ratio of dividing the training set and the test set is 4:1, that is, the first 80% of the historical data is divided into the training set, and the latter 20% of the data is divided into the test set.
[0020] Further, the specific content of performing normalization processing on the time information is to convert the time information into months, weeks, and hours, and perform normalization processing on months, weeks, and hours respectively to obtain time information variables.
[0021] Among them, the power consumption of the user from time t1 to time t2 is expressed as The time information variable from time t1 to time t2 is expressed as It includes three dimensions; the user number information is expressed as s.
[0022] Among them, the specific situation of processing historical data is to store the power consumption of users within the same time interval in a text or database. The information corresponding to each piece of data should include the time point corresponding to data collection and the number information used to distinguish different users.
[0023] Further, the neural network model in step S2 has a structure including: Temporal Convolutional Network (TCN), Gated Residual Network (GRN), embedding layer, and multi-head attention mechanism module. The step S2 is specifically as follows:
[0024] S2-1: The power consumption z from the starting time to time T in the historical data 1:T ∈R1×T and the time information x from the starting moment to moment T 1:T ∈R 3×T After concatenating in the row dimension, the input data X is obtained. The input data X is input into the Temporal Convolutional Network (TCN), and the output value is transposed to obtain the feature a t ∈R 1×25 ; the user identification information s ∈ R 1×1 is input into the embedding layer, and after processing, the vector c ∈ R is obtained 1×160 .
[0025] Among them, the embedding layer is a mapping from categorical variables to continuous vectors, and the parameters of the embedding layer can be updated by the gradient descent algorithm.
[0026] S2-2. The user identification information vector c ∈ R 1×160 and the feature a extracted by the Temporal Convolutional Network (TCN) t ∈R 1×25 are fused through the Gated Residual Network (GRN) to obtain the feature vector h containing the user identification information and electricity consumption information GRN .
[0027] S2-3. After the feature vector h GRN is mixed with the positional encoding vector e pos [1:T], the vectors K and V are output through the fully connected layer (FC) respectively; the time information x from moment T+1 to moment T+H T+1:T+H is transposed, and after passing through the fully connected layer (FC), it is combined with the positional encoding vector e pos [T+1:T+H] to output the vector Q; the vectors K, V, and Q are input into the multi-head attention mechanism module to obtain the output value h attn ∈R H×160 .
[0028] S2-4. The output value h attn ∈R H×160 After passing through the normalization layer (Norm), it is activated by the linear activation function ReLU, and finally the quantile prediction result is output through the fully connected layer (FC) where q ∈ {0.1, 0.5, 0.9}; is used as the result of point prediction, and and are used as the upper and lower bounds of the probability prediction under 80% confidence.
[0029] Among them, combining the advantages of the convolutional architecture and the self-attention mechanism, the neural network model can more efficiently capture the short-term and long-term time patterns of time series. At the same time, due to the processing of the user identification information, it can model the electricity consumption of all users in a region simultaneously, avoiding cumbersome operations and improving work efficiency.
[0030] Among them, the input data of the temporal convolutional network X∈R 4×T Represents a 4-row and T-column matrix consisting of the power consumption z from the start time to time T 1:T ∈R 1×T and the time variable information x from the start time to time T 1:T ∈R 3×T Concatenated along the row dimension.
[0031] Among them, fractional regression is used to achieve probabilistic prediction and quantify the uncertainty of the prediction results.
[0032] Furthermore, in the neural network model, except for the following parameters, the dimensions of other variables are set to 160: the dimensions of input and output data are adaptively changed; and the relevant parameters in the temporal convolutional network TCN.
[0033] Furthermore, the step S3 is specifically as follows:
[0034] S3-1, input the data of the first T time points of the processed training set into the neural network model, and output the quantile prediction results;
[0035] S3-2, compare the inverse normalized quantile prediction result with the true value of the prediction result, that is, the data at H time points after the training set, calculate the model loss through the loss function, and update the parameters of the neural network model through the gradient descent algorithm according to the model loss;
[0036] S3-3. Use the test set data to calculate the evaluation index of the model prediction accuracy and complete one round of training;
[0037] S3-4. Preset the number of training rounds. After completing the preset number of training rounds, select the model with the highest evaluation index and save it.
[0038] Among them, in the model training, the data set used is able to obtain the true value of the model prediction result, that is, the actual electricity consumption from time T+1 to time T+H.
[0039] Among them, during the model training process, the neural network model can explore the correlation between different users and further improve the accuracy of the prediction.
[0040] Among them, the training process of the neural network is to use the true value and the output value of the model to construct a loss function, then use the loss to find the partial derivative of the model parameters, and use the result of the partial derivative to complete the update of the model parameters.
[0041] Furthermore, the loss function is a fractional loss function, and the specific function is as follows:
[0042]
[0043]
[0044] wherein represents the true value of the prediction target, Ω represents the training sample domain containing M samples, H represents the size of the prediction window, and Q ∈ {0.1, 0.5, 0.9} represents the set of quantiles taken, represents the q-quantile prediction result of the model at time point t.
[0045] Among them, the partial derivative of the parameter is obtained through the loss function to update the parameters of the model, and the new parameter = the old parameter - (learning rate) * partial derivative value.
[0046] Furthermore, the evaluation index for calculating the prediction accuracy of the model has the following specific function:
[0047]
[0048]
[0049] wherein represents the validation sample domain containing N samples; y t represents the true value of the prediction result at time t; represents the q-quantile prediction result output by the model; H represents the electricity consumption predicted for the next H time points.
[0050] Among them, this evaluation index selects the commonly used 0.5-risk and 0.9-risk for quantile prediction, as well as RMSE for point prediction results.
[0051] Furthermore, the specific content of step S4 is as follows:
[0052] S4-1. Select the data of the previous T time points before the current time point for preprocessing;
[0053] S4-2. Input the preprocessed data into the trained neural network model to obtain the model output;
[0054] S4-3. Inverse normalize the model output to restore it to the original time scale to obtain the point prediction result and the probability prediction result. The probability prediction result is the value range of the prediction result under 80% confidence.
[0055] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0056] The present invention provides a method for predicting the probability of power consumption based on a neural network. By processing data and constructing a neural network model based on a convolutional architecture and a self-attention mechanism, it is possible to simultaneously model the power consumption data of different users in the power grid, discover the correlation relationships between different users, and further improve the accuracy of prediction. The trained neural network model can capture short-term and long-term patterns in the time series and achieve high-precision prediction of the time series. The prediction results of the model include conventional point prediction results and more realistic probability prediction results. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 It is a flowchart of a method for predicting the probability of power consumption based on a neural network according to the present invention.
[0058] Figure 2 It is a schematic diagram of the neural network model of a method for predicting the probability of power consumption based on a neural network according to the present invention.
[0059] Figure 3 It is a schematic diagram of the structure of the time convolutional network part in the neural network model of the present invention.
[0060] Figure 4 It is a schematic diagram of the structure of the gated residual network part in the neural network model of the present invention.
[0061] Figure 5 It is a visualization diagram of the prediction result of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other.
[0063] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0064] Embodiment 1
[0065] As Figure 1 - Figure 2 shown, the present invention provides a method for predicting the probability of power consumption based on a neural network. Specifically, it includes the following steps:
[0066] S1. Collect historical data of power consumption, divide it into a training set and a test set, and perform normalization processing on the variables in the historical data.
[0067] S2. Construct a neural network model based on a convolutional architecture and a self-attention mechanism.
[0068] S3. Use the processed training set data to train the neural network model, and select the model with the best prediction accuracy using the test set as the trained neural network model.
[0069] S4. Select recent data of power consumption and perform preprocessing, input the preprocessed recent data into the trained neural network model, and perform inverse normalization on the output value of the model to obtain the probability prediction result.
[0070] Further, in step S1, the historical data of power consumption is collected in a time series manner, and the time series includes the time information corresponding to each time point in the time series, the power consumption of electricity customers within the same time interval, and the user number information of different electricity customers.
[0071] Among them, the time information includes year, month, week, hour, minute, and second.
[0072] Further, the step S1 includes:
[0073] S1-1. Collect historical data of power consumption, divide the first n% of the data into the training set, and the latter (1 - n)% of the data into the test set;
[0074] S1-2. Divide the historical data in the training set and the test set into several samples of a preset length, where the preset length is (T + H). Use the data of the first T time points of each sample as the historical time series data input into the neural network model, and use the data of the latter H time points as the true value of the prediction result;
[0075] S1-3. Perform normalization on the time information and power consumption included in the historical data in the training set.
[0076] Among them, the ratio of dividing the training set and the test set is 4:1, that is, the first 80% of the historical data is divided into the training set, and the latter 20% of the data is divided into the test set.
[0077] Further, the specific content of performing normalization on the time information is to convert the time information into months, weeks, and hours, and perform normalization on months, weeks, and hours respectively to obtain time information variables.
[0078] In a specific embodiment, the time information "Saturday, March 26, 2022, 00:00" is converted into the following three variables. Month: 3; Week 6; Hour: 0, and then normalization is performed on these three variables respectively.
[0079] Among them, the power consumption of the user from time t1 to time t2 is expressed as The time information variable from time t1 to time t2 is denoted as It includes three dimensions; the user's identification information is denoted as s.
[0080] Among them, the specific situation of processing historical data is that the electricity consumption of users within the same time interval is stored in a text or database, and the information corresponding to each piece of data should include the time point corresponding to data collection and the identification information used to distinguish different users.
[0081] Furthermore, the neural network model in step S2 has a structure including: Temporal Convolutional Network (TCN), Gated Residual Network (GRN), embedding layer, and multi-head attention mechanism module. Step S2 is specifically as follows:
[0082] S2-1: Concatenate the electricity consumption z 1:T ∈R 1×T from the starting time to time T in the historical data and the time information x 1:T ∈R 3×T in the row dimension to obtain the input data X. Input X into the Temporal Convolutional Network (TCN), and after transposing the output value, obtain the feature a t ∈R 1×25 ; Input the user's identification information s ∈ R 1×1 into the embedding layer, and after processing, obtain the vector c ∈ R 1×160 .
[0083] Among them, the embedding layer is a mapping from categorical variables to continuous vectors, and the parameters of the embedding layer can be updated through the gradient descent algorithm.
[0084] S2-2: Fuse the user identification information vector c ∈ R 1×160 and the feature a t ∈R 1×25 extracted by the Temporal Convolutional Network (TCN) through the Gated Residual Network (GRN) to obtain the feature vector h GRN .
[0085] S2-3: After mixing the feature vector h GRN with the positional encoding vector e pos [1:T], output the vectors K and V respectively through the fully connected layer FC; Transpose the time information x T+1:T+H from time T + 1 to time T + H, and after passing through the fully connected layer FC, combine it with the positional encoding vector e pos [T + 1:T + H] to output the vector Q; Input the vectors K, V, and Q into the multi-head attention mechanism module to obtain the output value h attn ∈R H×160 .
[0086] S2-4. The output value h attn ∈R H×160 After passing through the normalization layer Norm, it is activated by the linear activation function ReLU, and finally the quantile prediction result is output through the fully connected layer FC where q ∈ {0.1, 0.5, 0.9}; is used as the result of point prediction, and and are used as the upper and lower bounds of the probability prediction under 80% confidence.
[0087] Among them, combining the advantages of the convolutional architecture and the self-attention mechanism, the neural network model can more efficiently capture the short-term and long-term time patterns of time series. Since the user number information is processed, it can model the power consumption of all users in a region at the same time, without training different models for different users, but can use the data of all users to train the same model, avoiding cumbersome operations and improving work efficiency.
[0088] Among them, the input data of the temporal convolutional network X ∈ R 4×T represents a matrix of 4 rows and T columns, which is concatenated in the row dimension by the electricity consumption z 1:T ∈ R 1×T from the starting time to time T and the time variable information x 1:T ∈ R 3×T from the starting time to time T.
[0089] Among them, probability prediction is realized by quantile regression, which quantifies the uncertainty of the prediction result.
[0090] Furthermore, for the neural network model, except for the following parameters, the dimensions of the remaining variables are all set to 160: the dimensions of the input and output data adaptively change; the relevant parameters in the temporal convolutional network TCN.
[0091] Furthermore, the specific steps of S3 are as follows:
[0092] S3-1. Input the data of the first T time points of the processed training set into the neural network model to output the quantile prediction result;
[0093] S3-2. After inverse-normalizing the quantile prediction result, compare it with the true value of the prediction result, that is, the data of the last H time points of the training set, calculate the model loss through the loss function, and update the parameters of the neural network model according to the model loss through the gradient descent algorithm;
[0094] S3-3. Use the test set data to calculate the evaluation index of the model prediction accuracy to complete one round of training;
[0095] S3-4. The preset number of training rounds. After completing the preset number of rounds of training, select the model with the highest evaluation index in the highest round and save it.
[0096] Among them, in the model training, the dataset used can obtain the true values of the model prediction results, that is, the true electricity consumption from time T+1 to time T+H.
[0097] Among them, in the process of model training, the neural network model can dig out the correlation relationships between different users, further improving the prediction accuracy.
[0098] Among them, the training process of the neural network is to construct a loss function using the true value and the output value of the model, then take the partial derivative of the model parameters using the loss, and use the result of the partial derivative to complete the update of the model parameters.
[0099] Furthermore, the loss function is a quantile loss function, and the specific function is as follows:
[0100]
[0101]
[0102] Where represents the true value of the prediction target, Ω represents the training sample domain containing M samples, H represents the size of the prediction window, Q∈{0.1, 0.5, 0.9} represents the set of quantiles taken, represents the q-quantile prediction result of the model at time point t.
[0103] Among them, the partial derivative of the parameters can be obtained through the loss function to update the model parameters, and the new parameters = old parameters - (learning rate) * partial derivative value.
[0104] Furthermore, the evaluation index for calculating the prediction accuracy of the model, its specific function is:
[0105]
[0106]
[0107] Where represents the validation sample domain containing N samples; y t represents the true value of the prediction result at time t; represents the q-quantile prediction result output by the model; H represents the electricity consumption predicted for the next H time points.
[0108] Among them, this evaluation index selects the commonly used 0.5-risk and 0.9-risk for quantile prediction, as well as RMSE for point prediction results.
[0109] Further, step S4 is specifically as follows:
[0110] S4-1. Select the data of the previous T time points before the current time point for preprocessing;
[0111] S4-2. Input the preprocessed data into the trained neural network model to obtain the model output;
[0112] S4-3. Inverse normalize the model output to restore it to the original time scale, obtaining the point prediction result and the probability prediction result. The probability prediction result is the value range of the prediction result under 80% confidence.
[0113] Embodiment 2
[0114] Based on the above Embodiment 1, combined with Figure 2 - Figure 4 the specific content of the neural network model is elaborated in detail in this embodiment.
[0115] For the Temporal Convolutional Network (TCN), combined with Figure 3 , in this embodiment, the convolutional kernel size of the dilated causal convolution is set to 7, and the number of convolutional kernel channels is 25. That is, the output dimension size after convolutional processing is 25. Here, N is set to 4, that is, a total of 8 layers of dilated causal convolution are used. The dilation factor d of the i-th layer of dilated convolution is 2 i .
[0116] In a specific embodiment, the input data X ∈ R 4×T , which represents a matrix data of 4 rows and T columns here. The number of convolutional kernel channels is 25. After the calculation of the dilated causal convolution, the output matrix h ∈ R 25×T is obtained. Then, a layer normalization operation Norm is performed, and finally, the activation function ReLU is used for activation to obtain h'. Here, the dimension of the matrix is not changed.
[0117] Among them, the Dropout operation means that during the model training process, a certain proportion of convolutional parameters are randomly set to 0. In this embodiment, the set proportion is 0.2, that is, 20% of the convolutional parameters are set to 0, which is used to improve the generalization performance of the model.
[0118] There is also an optional 1*1 convolution operation in the temporal convolutional network. This operation is used to add the input data to the output after the convolutional operation. There is a problem of dimension matching during the addition. In this embodiment, the input data X ∈ R 4×T is a matrix of 4 rows and T columns. The output matrix after convolution, normalization, and activation is a matrix of 25 rows and T columns. In this case of dimension mismatch, the input data needs to be processed by 1*1 convolution to make it a matrix of 25 rows and T columns.
[0119] In summary, the input data X ∈ R 4×T, after being processed by the temporal convolutional network, it is transposed for subsequent calculations and outputs a. t ∈R 1×25 , representing the output result of the temporal convolutional network at the t-th position.
[0120] For the gated residual network GRN, combined with Figure 4 , the purpose of this network is to fuse the user's identification information with the features extracted by the temporal convolutional network to obtain a feature vector containing the user's identification information and power consumption information. The function of the gated residual network GRN is specifically as follows:
[0121] GRN(a t , c) = LayerNorm(a t + GLU(η1))
[0122] η1 = η2 × W3 + b3
[0123] η2 = ELU(a t × W4 + c + W5 + b4)
[0124] Among them, the feature a extracted by the temporal convolutional network t ∈R 1×25 , and the user identification information c ∈ R after being processed by the embedding layer 1 ×160 , the parameter matrix W3 ∈ R 160×160 , W4 ∈ R 25×160 , W5 ∈ R 160×160 , b3 ∈ R 1×160 , b4 ∈ R 1×160 .
[0125] Among them, the linear gated unit GLU and the exponential linear unit activation function ELU are specifically:
[0126]
[0127] x < 0: ELU(x) = e x - 1
[0128] x ≥ 0: ELU(x) = x
[0129] Among them represents the element-wise multiplication of the matrices, γ ∈ R 1×160 , W1 ∈ R 160×160 , W2 ∈ R 160×160 , b1 ∈ R 1 ×160 , b2 ∈ R 1×160 .
[0130] The gated residual network GRN is for all positions of a t ∈R 1×25Performing the above operations, the final output is represented as h GRN ∈R T×160 。
[0131] For the multi-head attention mechanism module, combined with Figure 2 , the specific contents of the input data K, V, and Q are as follows:
[0132] K = FC key (h GRN + e pos [1:T]), K ∈ R T×160
[0133] V = FC value (h GRN + e pos [1:T]), V ∈ R T×160
[0134] Q = FC query (x T+1:T+H ′) + e pos [T + 1:T + H], Q ∈ R H×160
[0135] where x T+1:T+H ∈ R H×3 , the meaning of x T+1:T+H is the time information from time T + 1 to time T + H, which can represent the true value of the prediction result and needs to be transposed before being input into the model. e pos is the vector obtained by the position encoder. Finally, the output after passing through the multi-head attention mechanism module is denoted as h attn ∈ R H×160 。
[0136] The final output needs to perform a layer normalization operation on the output value of the multi-head attention mechanism module, and after activation, output the quantile prediction result through a fully connected layer. Specifically:
[0137]
[0138]
[0139] where represents the result of the q-th quantile prediction, represents the q-th quantile prediction result at the t-th future time point.
[0140] Example 3
[0141] Based on the above-mentioned Embodiment 1, in this embodiment, the prediction accuracy of the model is tested based on the publicly available dataset ElectricityLoadDiagrams20112014. The prediction task is set to predict the electricity consumption data for the next 24 hours based on the electricity consumption data for the past 168 hours.
[0142] The cleaned dataset contains the hourly electricity consumption data of 321 users from 2012 to 2014. In this simulation, the data between January 1, 2012 and June 7, 2014 is used as the training set, and the data from June 8, 2014 to December 31, 2014 is used as the test set to verify the prediction accuracy of the model.
[0143] To more intuitively reflect the prediction effect of the model, a sample is randomly selected and its prediction results are visualized, as Figure 5 shown. The solid curve represents the true value, and the dashed curve represents the prediction result of the model. The gray range area represents between and
[0144] which represents the range of the prediction result under 80% confidence, that is, the probabilistic prediction result. The left side of the dashed straight line represents the historical electricity consumption input into the model, and the right side is the future electricity consumption to be predicted.
[0145] The icons describing the structural position relationship in the drawings are only for illustrative purposes and should not be construed as a limitation of this patent.
[0146] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limiting the embodiments of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the embodiments here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A probability prediction method for power consumption based on a neural network, characterized in that, It includes the following steps: S1. Collect historical data of power consumption, divide it into a training set and a test set, and normalize the variables in the historical data; S2. Construct a neural network model based on a convolutional architecture and a self-attention mechanism; the neural network model has a structure including: a temporal convolutional network (TCN), a gated residual network (GRN), an embedding layer, and a multi-head attention mechanism module. The specific steps of S2 are as follows: S2-1. Concatenate the electricity consumption z from the starting time to time T in the historical data 1:T ∈R 1×T and the time information x from the starting time to time T 1:T ∈R 3×T on the row dimension to obtain the input data X. Input X into the Temporal Convolutional Network (TCN), and after transposing the output value, obtain the feature a t ∈R 1×25 ; Input the user's identification information s ∈ R 1×1 into the embedding layer, and after processing, obtain the vector c ∈ R 1×160 ; S2-2. The user ID information vector c ∈ R 1×160 and the feature a extracted by the Temporal Convolutional Network (TCN) t ∈ R 1×25 are fused through the Gated Residual Network (GRN) to obtain the feature vector h containing user ID information and power consumption information GRN ; S2-3. The feature vector h GRN The mixed position encoding vector e pos After [1:T], the vectors K and V are output respectively through the fully connected layer FC; the time information x from the (T + 1)-th moment to the (T + H)-th moment T+1:T+H is transposed, combined with the position encoding vector e after passing through the fully connected layer FC pos [T+1:T+H], and the vector Q is output; the vectors K, V, and Q are input into the multi-head attention mechanism module to obtain the output value h attn ∈R H×160 ; S2-4, the output value h attn ∈R H×160 After passing through the normalization layer Norm, it is activated by the linear activation function ReLU, and finally the quantile prediction result is output through the fully connected layer FC where q ∈ {0.1, 0.5, 0.9}; as the result of point prediction, and as the upper and lower bounds of probability prediction under 80% confidence; S3. Use the processed training set data to train the neural network model, and select the model with the highest prediction accuracy using the test set as the trained neural network model; S4. Select recent data of power consumption and perform preprocessing, input the preprocessed recent data into the trained neural network model, and perform inverse normalization on the output value of the model to obtain a probability prediction result.
2. The probability prediction method for power consumption based on a neural network according to claim 1, characterized in that, In step S1, the historical data of power consumption is collected in a time series manner. The time series includes the time information corresponding to each time point in the time series, the power consumption of electricity customers within the same time interval, and the user ID information of different electricity customers. The time information includes year, month, week, hour, minute, and second.
3. The probability prediction method for power consumption based on a neural network according to claim 2, characterized in that, Step S1 includes: S1-1. Collect historical data of power consumption, divide the first n% of the data into a training set, and the subsequent (100 - n)% of the data into a test set; S1-2. Divide the historical data in the training set and the test set into several samples of a preset length. The preset length is (T + H). Use the data of the first T time points in each sample as the historical time series data input to the neural network model, and use the data of the subsequent H time points as the true value of the prediction result; S1-3. Normalize the time information and power consumption included in the historical data in the training set.
4. The probability prediction method for power consumption based on a neural network according to claim 3, characterized in that, The specific content of normalizing the time information is to convert the time information into months, weeks, and hours, and perform normalization on months, weeks, and hours respectively to obtain time information variables; The electricity consumption of the user from time t1 to time t2 is expressed as The time information variable from time t1 to time t2 is expressed as The number information of the user is expressed as s.
5. The probability prediction method for power consumption based on a neural network according to claim 4, characterized in that, For the neural network model, except for the following parameters, the dimensions of the remaining variables are all set to 160: the dimensions of the input and output data adaptively change; the relevant parameters in the temporal convolutional network (TCN).
6. The probability prediction method for power consumption based on a neural network according to claim 1, characterized in that, The specific steps of S3 are as follows: S3-1. Input the data of the first T time points in the processed training set into the neural network model to output a quantile prediction result; S3-2. Compare the inverse-normalized quantile prediction result with the true value of the prediction result, that is, the data of the subsequent H time points in the training set, calculate the model loss through a loss function, and update the parameters of the neural network model according to the model loss through the gradient descent algorithm; S3-3. Use the test set data to calculate the evaluation index of the model prediction accuracy to complete one round of training; S3-4. Preset the number of training rounds, and save the model of the round with the highest evaluation index after completing the preset number of rounds of training.
7. The probability prediction method for power consumption based on a neural network according to claim 6, characterized in that, The loss function is a quantile loss function, and the specific function is as follows: where represents the true value of the prediction target, Ω represents the training sample domain containing M samples, H represents the size of the prediction window, and Θ = {0.1, 0.5, 0.9} represents the set of quantiles taken, represents the q - quantile prediction result of the model at time point t.
8. The probability prediction method for power consumption based on a neural network according to claim 7, characterized in that, The specific function of the evaluation index used to calculate the model prediction accuracy is: Among them, represents a validation sample domain containing N samples; y t represents the true value of the prediction result at time t.
9. The power consumption probability prediction method based on neural network according to claim 1, characterized in that, The specific steps of S4 are as follows: S4-1. Select the data of the first T time points before the current time point for preprocessing; S4-2. Input the preprocessed data into the trained neural network model to obtain the model output; S4-3. Inverse normalize the model output to restore it to the original time scale, obtaining the point prediction result and the probability prediction result, where the probability prediction result is the value range of the prediction result under 80% confidence level.
Citation Information
Patent Citations
Power consumption prediction method of convolutional neural network based on multi-head attention
CN114118568A