User load baseline load calculation method based on deep belief network prediction
Through the prediction method based on the deep belief network, the user load data is preprocessed and grouped, which solves the problem of the reduction in the accuracy of user load baseline load calculation in the prior art, and achieves higher prediction accuracy and accuracy of response effect evaluation.
Patent Information
- Application Number
- CN202510071328.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
AI Technical Summary
The existing user load baseline load calculation method has decreased accuracy under the situation of load growth and demand response resources continuing to increase, resulting in inaccurate evaluation of response effect.
The prediction method based on the Deep Belief Network (DBN) is used to preprocess the user load data, homologous grouping is performed through clustering and similarity calculation, and the data is input into the trained DBN to achieve accurate prediction of the user baseline load.
Improve the accuracy of user baseline load, especially when users participate in demand response for a long time, reduce prediction errors and provide more accurate response effect evaluation.
Smart Images

Figure CN119988871A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to an electric power system and automation thereof, and in particular to a method for calculating a user load baseline load based on deep belief network prediction. Background Art
[0002] Customer Baseline Load (CBL) is an important basis for evaluating the effect of user demand response. CBL is the power load that users should have reached without participating in demand response. Since the user has executed the response instruction to change the original power consumption mode during the response period, its load measurement value can only represent the power consumption behavior after the response, and the response amount provided by it needs to be determined by the difference between CBL and its load measurement value. Common user load baseline load calculation methods can be divided into historical average method, control group method, and regression method. Among them, the historical average method is the most widely used method in the current domestic demand response practice plan because of its simple method and easy execution. However, with the growth of load, the domestic demand for demand response resources continues to increase, and power users, especially industrial users, need to provide demand response for several consecutive days to weeks during the period of tight power supply and demand. This results in a large time span between the reference day set obtained by taking the non-response day in the existing plan and the response day. The accuracy of CBL estimated based on the average method and regression method has dropped significantly, which will cause inaccurate evaluation of the response effect.
[0003] With the popularization of smart meters, a large amount of user load data has been accumulated. Data-driven methods have improved the accuracy of short-term load forecasting and are expected to be applied to similar user baseline load calculation problems. How to accurately calculate the adjustable user load baseline load based on user historical data through data-driven methods remains to be further studied. The present invention proposes a method for calculating an adjustable user load baseline load based on deep belief network prediction to address the problems of insufficient prediction accuracy and large errors caused by traditional baseline prediction methods after users have participated in demand response for a long time, and to provide a reference for the accurate calculation of CBL for industrial users. Summary of the invention
[0004] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a method for calculating user load baseline load based on deep belief network prediction.
[0005] The purpose of the present invention is achieved through the following technical solution: A method for calculating user load baseline load based on deep belief network prediction, comprising the following steps:
[0006] Step S1, preprocessing the user load data, including detecting abnormal data, and filling or deleting the abnormal data to ensure the accuracy and completeness of the data;
[0007] Step S2: Based on the pre-processed user load data, clustering and similarity calculation are performed to perform homologous grouping processing on different user historical data sets, and typical power consumption pattern characteristics of users are extracted, so as to perform targeted baseline load forecasting for users with different power consumption patterns;
[0008] Step S3: input the data sets of different groups into the trained deep belief network DBN, and use the nonlinear mapping ability and strong generalization performance of DBN to achieve accurate prediction of user CBL.
[0009] Specifically, the abnormal data detection includes empty data detection, over-limit data detection, error data detection and duplicate data detection; the abnormal data is detected by the following formula:
[0010]
[0011] In the formula, x j,k is the user load at time k on day j, kW; n is the total number of observation days, days; is the average user load, kW; is the user load variance, kW 2 ; ε is a constant factor, which is 1.
[0012] Specifically, the method for filling abnormal data is:
[0013] Step S11, setting KNN model parameter values and establishing a KNN model;
[0014] Step S12, using the 95 point load values of the correct data set as input and the i-th point of the correct data set as output to train the KNN model;
[0015] Step S13: taking the first 95 point load values in the data to be corrected as input, and using the trained KNN model to predict the i-th point value in the data to be corrected;
[0016] Step S14: Replace the original value with the predicted value to generate a new 96-point load curve.
[0017] Specifically, the specific steps of step 2 are:
[0018] S21. The historical electricity consumption sample set D of the park users is divided into k non-overlapping clusters {Cf|f=1,2,…,k} by K-Means clustering algorithm, where And D = ∪ f=1 k C f ;
[0019] S22. The Calinski-Harabas index is used to evaluate the clustering effect. If the similarity between the typical load curves of users is high, it is determined to be a group with homologous characteristics.
[0020] Specifically, the K-Means clustering algorithm is used for a given data set x containing N M-dimensional data points. i ={x i1 ,x i2 ,...,x in}, where x d is an M-dimensional array, which is divided into K data classes Ck. Each data class Ck corresponds to a cluster center Uk. The distance from each data sample in the class to Uk is calculated. Through repeated iterations, the total sum of square distances of all data classes is minimized.
[0021] Let the sample set D = {x1, x2, ..., x m} represents the electricity consumption of a user for m days, where x i , i≤m, represents a sample, representing the power consumption change of the user in a day; each sample x i ={x i1 ,x i2 ,...,x in} is an n-dimensional feature vector; it is assumed that data is collected from the user's electricity meter every 15 minutes, with a total of 96 collection points per day. 96 discrete points are used to approximate the changes in the user's electricity consumption within a day.
[0022] Specifically, the Calinski-Harabas index is calculated by the following formula:
[0023]
[0024] Where n is the number of samples in the training set, k is the i-th sample in the training set, k is the number of given clusters, trB(k) is the trace of the between-class deviation matrix, trW(k) is the trace of the within-class deviation matrix, x is the mean of all sample data, c is the mean of all sample data, and trB(k) is the trace of the within-class deviation matrix. j is the center of the jth cluster; w j,i is the affiliation relationship between the ith data point and the jth cluster; X j is the jth cluster.
[0025] Specifically, the deep belief network is formed by aggregating multiple layers of restricted Boltzmann machines. For a given state (v, h), the energy function of the restricted Boltzmann machine is defined as:
[0026]
[0027] Where: is the setting parameter of the restricted Boltzmann machine; vi , α i are the state and bias of the ith visual unit respectively; w ij is the connection weight between the i-th visible unit and the j-th hidden unit; n and m are the visual unit v i and hidden units h j the number of
[0028] The joint probability distribution of any set of states (v, h) in a restricted Boltzmann machine is as follows:
[0029]
[0030] The probability distribution of its visual unit is shown as follows:
[0031]
[0032] In the formula, is the normalization factor;
[0033] When a specific hidden layer / visible layer node state is given, the state space between nodes in each layer is independently distributed:
[0034]
[0035] In the formula, P(h j |v)、P(v i |h) are the activation probabilities of neurons j(i) in the hidden layer / visual layer when the states of the visible layer / hidden layer nodes are given respectively:
[0036]
[0037] In the formula, is the sigmoid activation function, which is used to limit the output amplitude of the neuron.
[0038] Specifically, the training objective of the deep belief network is to maximize the log-likelihood function L(θ) of the input training sample set to obtain the optimal neural network model parameter θ′:
[0039]
[0040] The contrast divergence algorithm is used to train the sample set. First, the visual layer node v is initialized, and then Gibbs sampling is performed iteratively k times. Finally, P(h|v (k) ),v (k) To approximate the partial derivative of the log-likelihood function L(θ) corresponding to the parameter, the parameters are updated based on the following formula to complete the training of a single RBM model:
[0041]
[0042] α=∈α+η(v (0) -v (k) )
[0043] ρ=∈ρ+η[∈(h=1|v (0) )-∈(h=1|v (k) )]
[0044] Where ∈ is the disturbance variable and η is the learning rate.
[0045] The present invention has the following advantages:
[0046] The present invention firstly performs data preprocessing to preprocess missing values and outliers; secondly, different user historical data sets are grouped through clustering and similarity calculation to extract the typical power consumption pattern characteristics of users; finally, different groups of data sets are input into the trained DBN neural network to achieve accurate calculation of CBL for specific users, and perform prediction and post-evaluation calculation. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a schematic diagram of the calculation method flow of the present invention;
[0048] Figure 2 It is a schematic diagram of the user load data preprocessing process of the present invention;
[0049] Figure 3 It is a schematic diagram of the resource grouping process of the present invention;
[0050] Figure 4 Schematic diagram of the basic structure of the restricted Boltzmann machine of the present invention;
[0051] Figure 5 It is a schematic diagram of the DBN training process of the present invention;
[0052] Figure 6 The CBL calculation results of each scheme of user 2 in Example 1 of the present invention;
[0053] Figure 7 The CBL calculation results of each scheme of user 7 in Example 1 of the present invention;
[0054] Figure 8 The CBL calculation results of each scheme of user 8 in Example 1 of the present invention;
[0055] Fig. 9 The CBL calculation results of each scheme of user 9 in Example 1 of the present invention;
[0056] Fig.10 The CBL calculation results of each scheme of user 10 in Example 1 of the present invention;
[0057] Fig.11 The CBL calculation results of each scheme of user 11 in Example 1 of the present invention;
[0058] Fig.12 The CBL calculation results of each scheme of user 2 in Example 2 of the present invention;
[0059] Fig.13 The CBL calculation results of each scheme of user 7 in Example 2 of the present invention;
[0060] Fig.14 The CBL calculation results of each scheme of user 10 in Example 2 of the present invention;
[0061] Fig.15 The CBL calculation results of each scheme of user 12 in Example 2 of the present invention;
[0062] Fig.16 The CBL calculation results of each scheme of user 14 in Example 2 of the present invention are shown. DETAILED DESCRIPTION
[0063] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the embodiments described are only part of the embodiments of the present invention, rather than all of the embodiments. The components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.
[0064] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.
[0065] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "including a..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0066] The present invention is further described below in conjunction with the accompanying drawings, but the protection scope of the present invention is not limited to the following description.
[0067] like Figures 1 to 16 As shown, a method for calculating user load baseline load based on deep belief network prediction includes the following steps: Step S1, preprocessing user load data, including detecting abnormal data, and filling or deleting abnormal data to ensure the accuracy and completeness of the data;
[0068] Step S2: Based on the pre-processed user load data, clustering and similarity calculation are performed to perform homologous grouping processing on different user historical data sets, and typical power consumption pattern characteristics of users are extracted, so as to perform targeted baseline load forecasting for users with different power consumption patterns;
[0069] Step S3: input the data sets of different groups into the trained deep belief network DBN, and use the nonlinear mapping ability and strong generalization performance of DBN to achieve accurate prediction of user CBL.
[0070] Furthermore, the abnormal data detection includes empty data detection, over-limit data detection, error data detection and repeated data detection; the abnormal load data is detected using smoothing technology, and the abnormal data is detected by the following formula. If the following conditions are met, the data is abnormal data;
[0071]
[0072] In the formula, x j,k is the user load at time k on day j, kW; n is the total number of observation days, days; is the average user load, kW; is the user load variance, kW 2 ; ε is a constant factor, which is 1.
[0073] Furthermore, assuming that the i-th data point in the user's 96-point data is missing, the method for filling the abnormal data is as follows:
[0074] Step S11, setting the KNN model parameter value, i.e., k value, to establish the KNN model;
[0075] Step S12, using the 95 point load values of the correct data set as input and the i-th point of the correct data set as output to train the KNN model;
[0076] Step S13: taking the first 95 point load values in the data to be corrected as input, and using the trained KNN model to predict the i-th point value in the data to be corrected;
[0077] Step S14: Replace the original value with the predicted value to generate a new 96-point load curve.
[0078] Establishing homologous characteristic groups by extracting typical load curves can improve efficiency and save computing power. The electricity consumption characteristics of users in the same industry are relatively consistent, so they can be divided into homologous characteristic groups, and the historical data of users in homologous characteristic groups are used as the basis for CBL calculation data. According to the clustered historical electricity consumption typical load curve types, by calculating the similarity between the typical load curves of each user, the homologous characteristic group of each user is determined. The specific steps of step S2 are:
[0079] S21, K-Means clustering algorithm for a given data set x containing N M-dimensional data points i ={x i1 ,x i2 ,...,x in}, where x d is an M-dimensional array, which is divided into K data classes Ck. Each data class Ck corresponds to a cluster center Uk. The distance from each data sample in the class to Uk is calculated. Through repeated iterations, the total sum of square distances of all data classes is minimized.
[0080] Let the sample set D = {x1, x2, ..., x m} represents the electricity consumption of a user for m days, where x i , i≤m, represents a sample, representing the power consumption change of the user in a day; each sample x i ={x i1 ,x i2 ,...,x in} is an n-dimensional feature vector; it is assumed that the user's electricity meter is collected once every 15 minutes, with a total of 96 collection points per day. 96 discrete points are used to approximate the user's power consumption changes within a day. The historical power consumption sample set D of the park users is divided into k non-overlapping clusters {Cf|f=1,2,…,k} through the K-Means clustering algorithm, where And D = ∪ f=1 k C f ;
[0081] S22. The clustering effect is evaluated by the Calinski-Harabas index. If the similarity between the typical load curves of the users is high, it is determined that they are homologous characteristic groups. The Calinski-Harabas index is calculated by the following formula; CH measures the compactness by calculating the sum of squares of the distances between the samples within the class and the cluster center of the class (i.e., the intra-class deviation matrix), and measures the separation by calculating the inter-class deviation matrix. The CH value is the ratio of the separation to the compactness. The larger the CH value, the better the clustering effect:
[0082]
[0083]
[0084] Where n is the number of samples in the training set, k is the i-th sample in the training set, k is the number of given clusters, trB(k) is the trace of the between-class deviation matrix, trW(k) is the trace of the within-class deviation matrix, x is the mean of all sample data, c is the mean of all sample data, and trB(k) is the trace of the within-class deviation matrix. j is the center of the jth cluster; w j,i is the affiliation relationship between the ith data point and the jth cluster; X j is the jth cluster.
[0085] Furthermore, the deep belief network is formed by aggregating multiple layers of restricted Boltzmann machines. The restricted Boltzmann machine (RBM) is an energy-based probability model, a two-layer recursive neural network consisting of n visible units and m hidden units, wherein the neural units of the input layer and the neural units of the hidden layer are both binary variables, the units within the RBM layer are unconnected, and the units between the layers are fully connected. The bottom layer (input layer) receives external input data, and the input data is further converted to the hidden layer based on the restricted Boltzmann machine, wherein the data output of the lower layer restricted Boltzmann machine is the data input of the higher layer restricted Boltzmann machine, and the homologous data group is used as the input set of the deep belief network algorithm to predict the CBL of a typical user on the demand response day. For a given state (v, h), the energy function of the restricted Boltzmann machine is defined as:
[0086]
[0087] Where: θ={w=(w ij ) n×m ,α=(α i ) 1×n ,ρ=(ρ j ) 1×m} is the setting parameter of the restricted Boltzmann machine; v i , α i are the state and bias of the ith visual unit respectively; w ij is the connection weight between the i-th visible unit and the j-th hidden unit; n and m are the visual unit v i and hidden units h j the number of
[0088] The joint probability distribution of any set of states (v, h) in a restricted Boltzmann machine is as follows:
[0089]
[0090] The probability distribution of its visual unit is shown as follows:
[0091]
[0092] In the formula, is the normalization factor;
[0093] Since RBM has a binary structure, when a specific hidden layer / visible layer node state is given, the state space between nodes in each layer is independently distributed:
[0094]
[0095] In the formula, P(h j |v)、P(v i |h) are the activation probabilities of neurons j(i) in the hidden layer / visual layer when the states of the visible layer / hidden layer nodes are given respectively:
[0096]
[0097] In the formula, is the sigmoid activation function, which is used to limit the output amplitude of the neuron.
[0098] Furthermore, RBM is the best fit for a given training sample S = {s1, s2, ..., s N}, its training goal is to maximize the log-likelihood function L(θ) of the input training sample set to obtain the optimal neural network model parameter θ′:
[0099]
[0100] The contrast divergence algorithm is used to train the sample set. First, the visual layer node v is initialized, and then Gibbs sampling is performed iteratively k times. Finally, P(h|v (k) ),v (k) To approximate the partial derivative of the log-likelihood function L(θ) corresponding to the parameter, the parameters are updated based on the following formula to complete the training of a single RBM model:
[0101]
[0102] α=∈α+η(v (0) -v (k) )
[0103] ρ=∈ρ+η[∈(h=1|v (0) )-∈(h=1|v (k) )]
[0104] Where ∈ is the disturbance variable and η is the learning rate.
[0105] The entire DBN training process mainly includes two stages: unsupervised pre-training and reverse fine-tuning.
[0106] The unsupervised pre-training stage is a bottom-up learning process. The unsupervised greedy learning mechanism based on the restricted Boltzmann machine is trained from the bottom layer to the top layer, and then the low-level multi-dimensional features are gradually aggregated into high-level features through continuous transmission.
[0107] The reverse fine-tuning stage is a top-down adjustment process. It obtains the error value between the previous classification result and the actual result by solving it. It uses algorithms such as gradient descent or conjugate gradient descent to adjust parameters layer by layer from the top layer to the bottom according to the error coefficient, and iteratively updates the internal parameters θ of the DBN model to further reduce the training error.
[0108] The first stage of unsupervised pre-training is to find the optimal solution range of the model in a larger range and complete the construction of the main structure of the model. In the second stage, the reverse parameter fine-tuning stage is to locate the optimal solution in a small range and realize internal fine-tuning of the optimal solution of the model.
[0109] In the model training phase, grid search and simple time series cross validation were used to try different hyperparameters, including training rounds, batch size, optimizer type, etc., to adjust and evaluate the model performance. In the validation phase, the true value and the predicted value were calculated and compared through model prediction and reverse conversion, iterative prediction and other steps to better verify the generalization ability of the model. Grid search can systematically combine parameters to find the best hyperparameter combination; the cross validation method takes into account the dependency and time sequence between data points, which can effectively avoid data leakage and is conducive to practical applications.
[0110] After the prediction is evaluated, the deviation between the predicted CBL and the true CBL is the deviation between the estimated CBL and the actual measured load. The CBL estimation performance can be verified by comparing the estimated CBL with the actual measured load.
[0111] Single estimation index,CBL estimation method mainly has two single evaluation indexes, namely, precision and accuracy.
[0112] The accuracy of CBL estimation can be expressed by the average value of the mean absolute error (MAE) of the baseline, which is expressed as follows:
[0113]
[0114] In the formula and Represent the estimated and actual CBL values of user i on day d, respectively; |δ| is the number of time nodes in the demand response time period; D is the number of days of the selected class demand response day. MAE reflects the average estimation performance of the estimation method at each sampling point. The smaller the MAE, the higher the accuracy of CBL estimation at a single sampling point.
[0115] The accuracy of CBL estimation can be measured by the average value of the average error (Bias) between the estimated and actual baseline, as shown in the following formula:
[0116]
[0117] Bias reflects the degree of deviation between the estimated total load value and the actual total load value during the demand response period. Because the total estimated load value during the demand response period is directly related to the compensation settlement for participating in the demand response, the performance of the Bias indicator is more important than MAE. Compared with MAE, which is all positive, Bias has positive and negative properties. If Bias is positive, it means that the estimated value is too high; if Bias is negative, it means that the estimated value is too low. The smaller the absolute value of Bias, the smaller the deviation of CBL estimation during the entire demand response period, and the better the estimation performance of CBL.
[0118] The comprehensive evaluation index, MAE and Bias, can evaluate the performance of a given CBL estimation method from two different aspects. In order to evaluate the overall performance of different CBL estimation methods, a normalized index, the Overall Performance Index (OPI), is proposed. Assuming that the set of CBL estimation methods to be compared is represented by W = {w|w = 1, 2, …, W}, the OPI index of the w-th method can be expressed as follows.
[0119]
[0120] Where M w represents the MAE value corresponding to the wth method; M max With M min is the corresponding maximum and minimum MAE values among all the methods involved in the comparison; |B w | represents the |Bias| value corresponding to the wth method; |B| max With |B| min The corresponding maximum and minimum values of |Bias| among all the methods involved in the comparison. The smaller the OPI value, the better the overall estimation performance of the method.
[0121] Example: To compare the effects of the traditional CBL calculation method and the CBL calculation method proposed in the present invention, two hypothetical demand response examples are designed: Example 1: single-day response; Example 2: continuous 5-day response.
[0122] Based on the actual historical data of 16 users in 2023, the single-day response date is set in the evening peak period of September 14 in the data set; the five-day response date is set in the evening peak period of September 14-18 in the data set. The above dates are all normal load days without demand response. Six CBL calculation methods are designed as follows for method comparison and verification:
[0123] Method 1: Historical average method (MEAN1) selects 10 days before the response day, excluding the days when no power outages and users participated in demand response, and then removes the two days with the largest and smallest daily maximum loads of power users from the above 10 days. The remaining 8 days are used as the load curve of the response period corresponding to the typical day as the baseline (i.e., the user-side baseline load calculation method used in the northwest region);
[0124] Method 2: Improved historical average method (MEAN2) Select the eight normal working days before the response day to form a baseline reference day set, and calculate the average load P of each reference day during the demand response period avi and the average load P during the demand response period of the eight reference days av , if any P avi <0.75*P av , then remove the day from the reference day set, and recursively select another day forward until 8 reference days that meet the requirements are selected;
[0125] Method 3: Improved historical average method with adjustment factor (MEAN2*) Multiply the CBL of MEAN2 by the adjustment factor, which is the ratio of the actual load to the predicted load 2 hours before the response;
[0126] Method 4: Control group method (CG), which selects the user cluster data with the highest load similarity in the same period for weighted average calculation;
[0127] Method 5: Long Short-Term Memory Neural Network (LSTM) prediction method, which uses the LSTM model to train and calculate a single user;
[0128] Method 5: The method proposed by the present invention.
[0129] Example results,In order to eliminate the influence of accidental factors, the program is run 100 times, and each time the test group users and the control group users are randomly selected from all users. The average values of MAE, Bias and OPI indicators of the seven methods under all running times are calculated, and the estimated indicator results of each method in each scenario are listed in Tables 1 and 2. It can be seen that the proposed method shows the best MAE, Bias, and OPI performance, and its Bias performance is also competitive after verification by various users.
[0130] Overall, the results of the control group method are better than those of the average method, considering that the average method is difficult to effectively deal with the instability of the load pattern of a single user. As the number of aggregated users increases, the temporal autocorrelation of the load curve will be further improved.
[0131] Table 1 Algorithm index calculation results (single-day response)
[0132]
[0133]
[0134]
[0135] Table 2 Algorithm indicator calculation results (five consecutive days of response)
[0136]
[0137]
[0138] Example 1: Single-day response
[0139] The calculation results of the CBL of some users under the six schemes in Example 1 are as follows: Figure 6 As shown in Table 1. The three historical average methods have similar effects, and the deviations of CBL calculation are all high; the overall accuracy of the control group method is good, but there is a large deviation in the period of 16:00-18:00; the CBL curve obtained by DBN is similar to the historical average method, but the curve is too smooth; and the method proposed in the present invention has higher calculation accuracy and smaller deviation at each moment in the response period.
[0140] Taking user 1 as an example, since user 1 intends to adopt a special production mode on the response day, the three averaging methods all underestimate the actual power consumption curve of the user during the response period to varying degrees. Since they only use historical data for estimation, they cannot account for the sudden increase in load due to unexpected factors. The performance of the control group is better than the average method, but it is still worse than the method proposed in this chapter. The reason for this phenomenon may be that the control group only uses historical load curves of multiple users similar to this user, but the accuracy cannot be guaranteed by accurate fitting. The method proposed in this chapter is based on the CBL evaluation method driven by the multi-dimensional production load data of this user, and has better prediction performance.
[0141] Example 2: Response for five consecutive days
[0142] Example 2: The calculation results of CBL of some enterprises under six calculation schemes when participating in the response for 5 consecutive days are as follows Figure 7As shown in Table 2. In this section, the period of DR event occurrence is set as 18:00--22:00 (4 hours) when the system load is high. As can be seen from the figure, under the response of 5 consecutive days, the historical average method can only maintain the same period prediction as the result of the first day, and it is difficult to ensure the accuracy of subsequent calculations; the calculation results of the control group method are gradually inaccurate; the method proposed in the present invention can still maintain a high accuracy under the condition of long-term continuous participation in response.
[0143] The above is only a preferred embodiment of the present invention and does not limit the present invention in any form. Any technician familiar with the art can make many possible changes and modifications to the technical solution of the present invention by using the above-mentioned technical content without departing from the scope of the technical solution of the present invention, or modify it into an equivalent embodiment of equivalent changes. Therefore, any changes, modifications, equivalent changes and modifications made to the above embodiments based on the technology of the present invention without departing from the content of the technical solution of the present invention belong to the protection scope of the present technical solution.
Claims
1. A method for calculating user load baseline based on deep belief network prediction, characterized in that: The following steps are involved: Step S1, preprocessing the user load data, including detecting abnormal data, and filling or deleting the abnormal data to ensure the accuracy and completeness of the data; Step S2: Based on the pre-processed user load data, clustering and similarity calculation are performed to perform homologous grouping processing on different user historical data sets, and the typical power consumption pattern characteristics of users are extracted to facilitate targeted baseline load prediction for users with different power consumption patterns; Step S3: input the data sets of different groups into the trained deep belief network DBN, and use the nonlinear mapping ability and strong generalization performance of DBN to achieve accurate prediction of user CBL.
2. The method for calculating user load baseline based on deep belief network prediction according to claim 1, characterized in that: The abnormal data detection includes empty data detection, over-limit data detection, error data detection and duplicate data detection; the abnormal data is detected by the following formula: In the formula, x j,k is the user load at time k on day j, kW; n is the total number of observation days, days; is the average user load, kW; is the user load variance, kW 2 ; ε is a constant factor, which is 1.
3. The method for calculating user load baseline based on deep belief network prediction according to claim 2, characterized in that: The method to fill in abnormal data is: Step S11, setting KNN model parameter values and establishing a KNN model; Step S12, using the 95 point load values of the correct data set as input and the i-th point of the correct data set as output to train the KNN model; Step S13: taking the first 95 point load values in the data to be corrected as input, and using the trained KNN model to predict the i-th point value in the data to be corrected; Step S14: Replace the original value with the predicted value to generate a new 96-point load curve.
4. The method for calculating user load baseline based on deep belief network prediction according to claim 1, characterized in that: The specific steps of step S2 are: S21. The historical electricity consumption sample set D of the park users is divided into k non-overlapping clusters {Cf|f=1,2,…,k} by K-Means clustering algorithm, where f'≠f, and D=∪ f=1 k C f ; S22. The Calinski-Harabas index is used to evaluate the clustering effect. If the similarity between the typical load curves of users is high, it is determined to be a group with homologous characteristics.
5. The method for calculating user load baseline based on deep belief network prediction according to claim 4, characterized in that: The K-Means clustering algorithm is given a data set x containing N M-dimensional data points. i ={x i1 ,x i2 ,...,x in }, where x d is an M-dimensional array, which is divided into K data classes Ck. Each data class Ck corresponds to a cluster center Uk. The distance from each data sample in the class to Uk is calculated. Through repeated iterations, the total sum of square distances of all data classes is minimized. Let the sample set D = {x1, x2, ..., x m } represents the electricity consumption of a user for m days, where x i , i≤m, represents a sample, representing the power consumption change of the user in a day; each sample x i ={x i1 ,x i2 ,...,x in } is an n-dimensional feature vector; it is assumed that data is collected from the user's electricity meter every 15 minutes, with a total of 96 collection points per day, and 96 discrete points are used to approximately represent the changes in the user's electricity consumption within a day.
6. The method for calculating user load baseline based on deep belief network prediction according to claim 5, characterized in that: The Calinski-Harabas index is calculated by the following formula: Where n is the number of samples in the training set, k is the i-th sample in the training set, k is the number of given clusters, trB(k) is the trace of the between-class deviation matrix; trW(k) is the trace of the within-class deviation matrix; x is the mean of all sample data; c j is the center of the jth cluster; w j,i is the affiliation relationship between the ith data point and the jth cluster; X j is the jth cluster.
7. The method for calculating user load baseline based on deep belief network prediction according to claim 5, characterized in that: The deep belief network is formed by aggregating multiple layers of restricted Boltzmann machines. For a given state (v, h), the energy function of the restricted Boltzmann machine is defined as: Where: θ={w=(w ij ) n×m ,α=(α i ) 1×n ,ρ=(ρ j ) 1×m } is the setting parameter of the restricted Boltzmann machine; v i , α i are the state and bias of the i-th visual unit respectively; w ij is the connection weight between the i-th visible unit and the j-th hidden unit; n and m are the visual unit v i and hidden units h j the number of The joint probability distribution of any set of states (v, h) in a restricted Boltzmann machine is as follows: The probability distribution of its visual unit is shown as follows: In the formula, is the normalization factor; When a specific hidden layer / visible layer node state is given, the state space between nodes in each layer is independently distributed: In the formula, P(h j |v)、P(v i |h) are the activation probabilities of neurons j(i) in the hidden layer / visual layer when the visible layer / hidden layer node states are given respectively: In the formula, is the sigmoid activation function, which is used to limit the output amplitude of the neuron.
8. The method for calculating user load baseline based on deep belief network prediction according to claim 7, characterized in that: The training goal of the deep belief network is to maximize the log-likelihood function L(θ) of the input training sample set to obtain the optimal neural network model parameter θ′: The contrast divergence algorithm is used to train the sample set. First, the visual layer node v is initialized, and then Gibbs sampling is performed iteratively k times. Finally, P(hv (k) ),v (k) To approximate the partial derivative of the log-likelihood function L(θ) corresponding to the parameter, the parameters are updated based on the following formula to complete the training of a single RBM model: α=∈α+η(v (0) -v (k) ) ρ=∈ρ+η[∈(h=1|v (0) )-∈(h=1|v (k) )] Where ∈ is the disturbance variable and η is the learning rate.
Citation Information
Cited By
Virtual power plant baseline calculation method introducing user confidence
CN120832460A
A virtual power plant baseline calculation method introducing user confidence
CN120832460B