Transformer Area Load Forecasting Method and System
Through the Prophet model and the unsupervised learning algorithm of the automatic encoder combined with the LightGBM model, the problem of poor load prediction accuracy in the station area is solved, more accurate load prediction is achieved, and the scheduling plan formulation of the power system is supported.
Patent Information
- Application Number
- CN202210654086.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-10
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-06-10
AI Technical Summary
The load prediction accuracy of the middle-end station area in the prior art is poor, especially due to the complex power consumption environment in the station area, the influence of bad data and a variety of random factors, and the traditional methods are difficult to deal with nonlinear relationships and deep features.
The Prophet model is used for feature analysis, combined with the automatic encoder unsupervised learning algorithm for nonlinear dimensionality reduction, and the LightGBM model is used for table area load prediction.
It improves the accuracy of load prediction in the station area, and can estimate the load changes in the power supply area in advance, providing a reference for equipment capacity increase.
Smart Images

Figure CN114925931B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distribution automation, and specifically, to a method and system for predicting the load of a distribution transformer area. Background Art
[0002] Accurate load prediction helps units at all levels to adjust dispatching plans and work arrangements, and is the basis for the safe operation of the power system. Currently, most prediction work focuses on the prediction of the load of the distribution network system, and there is little research on the prediction of the load of the distribution transformer area. The power consumption environment in the distribution transformer area is complex, there are more bad data, and the randomness is very strong, which will bring certain difficulties to the load prediction of the distribution transformer area. However, since the load data of the distribution transformer area is affected by various complex factors such as temperature, humidity, weather, date type, and season, all influencing factors need to be considered when predicting the load of the distribution transformer area. However, there is a certain blindness and limitation in manually selecting influencing factors, and the deep features of the influencing factors cannot be found. In addition, some unconventional factors also affect the accuracy of the load prediction of the distribution transformer area to a certain extent, but most studies do not involve the influence of unconventional factors on the load prediction of the distribution transformer area.
[0003] Finally, the relationship between the load prediction of the distribution transformer area and the influencing factors is more of a non-linear relationship. Traditional dimensionality reduction and feature extraction methods such as PCA cannot learn non-linear features, and most load prediction research work does not involve their deep features either. Summary of the Invention
[0004] The present invention provides a method and system for predicting the load of a distribution transformer area, which solves the problem of poor accuracy in the load prediction of the distribution transformer area in the related art.
[0005] The technical solution of the present invention is as follows:
[0006] In a first aspect, a method for predicting the load of a distribution transformer area includes:
[0007] Obtain historical data, where the historical data includes multiple historical records, and each historical record includes a corresponding load of the distribution transformer area, date, time, and influencing factors; the influencing factors include temperature and humidity;
[0008] Use the Prophet model to perform feature analysis and extraction on the historical data to obtain a training set feature data set A and a test set feature data set B;
[0009] Use the unsupervised learning algorithm of the autoencoder to perform automated feature extraction and non-linear dimensionality reduction processing on the training set feature data set A to obtain a feature data set Z; use the unsupervised learning algorithm of the autoencoder to perform automated feature extraction and non-linear dimensionality reduction processing on the training set feature data set B to obtain a feature data set Z';
[0010] Train a LightGBM model according to the feature dataset Z and the feature dataset Z'; the LightGBM model is used for predicting the load of the substation area.
[0011] In a second aspect, a substation area load prediction system includes:
[0012] A first acquisition unit for acquiring historical data, where the historical data includes multiple historical records, and each historical record includes a corresponding substation area load, date, time, and influencing factors; the influencing factors include temperature and humidity;
[0013] A first feature extraction unit for performing feature analysis and extraction on the historical data using a Prophet model to obtain a training set feature dataset A and a test set feature dataset B;
[0014] A second feature extraction unit uses an autoencoder unsupervised learning algorithm to perform automated feature extraction and nonlinear dimensionality reduction processing on the training set feature dataset A to obtain a feature dataset Z; uses an autoencoder unsupervised learning algorithm to perform automated feature extraction and nonlinear dimensionality reduction processing on the training set feature dataset B to obtain a feature dataset Z';
[0015] A first processing unit for training a LightGBM model according to the feature dataset Z and the feature dataset Z'; the LightGBM model is used for predicting the load of the substation area.
[0016] In a third aspect, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the substation area load prediction method are implemented.
[0017] The working principle and beneficial effects of the present invention are as follows:
[0018] A substation area load prediction method based on Prophet-AE-LightGBM proposed by the present invention, based on the complex nonlinear relationship between various influencing factors and the substation area load, this method extracts the nonlinear deep features of various influencing factors and the substation area load, improves the accuracy of substation area load prediction, and provides data support for formulating relevant scheduling plans and power consumption plans. Accurate substation area load prediction can estimate in advance the change of the load in the power supply area of the substation area, and provide an effective reference for the capacity expansion of equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0020] Figure 1 It is a flowchart of the method of the present invention;
[0021] Figure 2Schematic diagram of the autoencoder structure in the present invention; Specific implementation manner
[0022] Next, in combination with the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0023] 1. Data acquisition
[0024] Obtain the data at 24:00 of the daily historical load of a certain substation area in a certain region in the past 1 year, and then conduct statistical analysis to initially filter out bad data. Bad data refers to data with measurement problems or null values at 24:00. At the same time, obtain the data at 24:00 of temperature and the data at 24:00 of humidity.
[0025] 2. Data preprocessing
[0026] (1) For outliers, the following method is adopted, that is, the data value at a certain moment is compared with the "base value" of this data. The "base value" of the substation area load data is the maximum capacity of the substation area, and the "base value" of temperature and humidity is the average value in the past 7 days. If the deviation exceeds the threshold, the horizontal processing method is used for processing. The discrimination formula is as follows:
[0027]
[0028] y(t,d) is the data of the d-th type of data at time t, and y 基 is the "base value" of this type of data. If it exceeds 20% of the threshold, the average value method is used for processing. The formula of the average value method is as follows:
[0029]
[0030] (2) For missing values, the average method is also used for supplementation.
[0031] 3. Prophet feature analysis and extraction
[0032] Above, we analyzed the influence of conventional factors on the load forecasting of the substation area. However, in practical applications, some unconventional influencing factors will have a significant impact on the load forecasting of the substation area. For example, the loads of residential and commercial substations may be affected by social events; the load of industrial substations will also be affected by industrial models and production plans. Most load forecasting does not involve these unconventional factors when considering influencing factors. However, in the load forecasting of the substation area, these unconventional factors will indeed have a certain impact on the accuracy of the load forecasting of the substation area. To address this issue, this patent uses the Prophet algorithm to perform feature analysis and extraction on unconventional factors to measure the influence of these unconventional factors. At the same time, feature analysis and extraction work on conventional factors is also carried out.
[0033] Some unconventional factors may cause the load of residential substations in a certain area to increase, such as social events like large-scale competitions and large-scale conferences. On the contrary, some unconventional factors may cause the industrial load in a certain area to decrease or even stop production, and the load of industrial substations will show a downward trend, such as production plan changes, production cuts caused by changes in industrial models, etc. The same is true for the load of commercial substations. The Prophet algorithm can help us solve this problem. The specific steps are as follows:
[0034] (1) Preliminary research and statistics to determine unconventional factors and the duration of influence
[0035] First, conduct detailed research and statistical work. According to the different types of substation area loads, determine the degree of influence of a certain unconventional factor on the substation area load type. Obtain the start and end times of a certain unconventional factor in this region within historical dates. Combine the actual situation for specific analysis to determine the unconventional factors and the duration of influence. For convenience, this patent regards unconventional factors as special events.
[0036] (2) Dataset division
[0037] Then, divide the 24:00 data of the historical load of a certain distribution transformer substation area and temperature and humidity data obtained in the past year into a training dataset and a test dataset according to a ratio of 8:2;
[0038] (3) Fitting the load of the training set
[0039] Use the additive model of Prophet to extract features from the load of the training dataset. First, input the load data of the training dataset and the corresponding temperature and humidity data into Prophet. Then, add monthly, weekly, daily, and holiday effects to the model. Next, add special events to the model. Finally, use the additive model and fitting function of Prophet to obtain the fitted load of the training dataset. The formula for the additive model of Prophet is as follows:
[0040] u 拟 =u趋势 +u 月 +u 周 +u 天 +u 节假日 +u 回归 +u 特殊
[0041] u 拟 The load data fitted by Prophet is divided into 7 parts: u 趋势 Represents the overall trend change term in the fitted load data, u 月 Represents the monthly trend change term in the fitted load data, u 周 Represents the weekly trend change term in the fitted load data, u 天 Represents the daily trend change term in the fitted load data, u 节假日 Is the impact of holidays on the fitted load data, u 回归 Is the impact of changes in additional regression variables such as temperature and humidity on the fitted load data, u 特殊 Is the impact of special events on the fitted load. Then the above 8 items of data obtained are used as the training feature data set A.
[0042] (4) Training set feature extraction
[0043] Then, the feature extraction of the test data set is carried out using the additive model of Prophet constructed above. The temperature, humidity data and special events included in the test data set are input into the above additive model of Prophet, and its prediction function is used to predict the load in the test data set. The prediction results of the test data set also include the above 8 items of data, and finally they are used as the test feature data set B.
[0044] 4. Autoencoder feature extraction
[0045] Using the autoencoder unsupervised learning algorithm, automated feature extraction and nonlinear dimensionality reduction processing are carried out on the training set feature data set A and the test set feature data set B. The autoencoder is essentially an unsupervised feature learning method, and its specific structure consists of an encoder and a decoder, as Figure 2 shown.
[0046] (1) Feature data normalization
[0047] Taking A as an example, first normalize the feature data set, and the formula is as follows:
[0048]
[0049] Among them, maxAi and minAi represent the maximum and minimum values of the i-th feature data respectively. a ti*Represents the normalized data of the i-th feature at time t. The normalized data set is regarded as A*, which is specifically represented as the following matrix:
[0050]
[0051] The row vectors of A* represent the normalized data of different features at the same time, and the column vectors represent the normalized data of the same feature at different times. m is the number of historical times, and v is the number of features.
[0052] (2) Feature data extraction
[0053] The normalized feature data A* is used as the input data of the autoencoder. The input data and output data of the autoencoder are almost the same, that is to say, the autoencoder can perform data reproduction operations. First, the autoencoder uses a function f to map the input data to the hidden layer. This process is encoding, and the mathematical expression is:
[0054] Z = f(A * )
[0055] The output Z of the hidden layer is called the hidden feature variable. Use its hidden variable to reconstruct A~. At this time, A~ has the same structure as A*. This process is decoding, and the mathematical expression is:
[0056]
[0057] Both the function f and the function g are Sigmoid functions, indicating that the data is non-linearly mapped to 0-1. The principle of the autoencoder is to obtain the representation of its hidden layer by converting the input data; then it is reconstructed by the hidden layer to restore the new input data. When the feature data is processed by encoding and decoding, if the error between the reconstructed output data and the original feature data is within the limited range, it can be considered that the encoding process is an effective expression of the original feature data. The loss function for constructing the autoencoder is usually defined using the mean squared error:
[0058]
[0059] Z is the feature data set Z of the training set that we finally need. Z has the same number of rows as A*, and the number of columns is less than A*, that is, the feature dimension is less than A*. Similarly, use AE to perform the above operations on the feature data set B of the test set to obtain the feature data set Z' of the test set.
[0060] 5. LightGBM model prediction
[0061] (1) Data division
[0062] Divide the data at 24:00 of the historical load of a certain distribution transformer substation area in the past year obtained according to the ratio of the feature dataset Z to the feature dataset Z' (also in the ratio of 8:2), and divide it into a training load dataset and a test load dataset; then use the feature data as the input value of the LightGBM model and the load data as the output value of the LightGBM model, so that the input value and the output value correspond one by one.
[0063] (2) Model training
[0064] Input the feature dataset Z of the training set into the LightGBM model, select the LightGBM default parameters, and calculate the RMSE of the error between the predicted load and the actual load of the training set.
[0065]
[0066] pre h is the actual value at time h, is the predicted value at time h, and p is the total number of prediction times.
[0067] (3) Optuna parameter tuning
[0068] Use the RMSE error of the training set and the Optuna parameter tuning algorithm to tune the parameters, continuously search for the optimal model parameters to minimize the RMSE error of the training set until the optimal LightGBM model is found.
[0069] (4) Predict the load and output the results
[0070] Use the optimal LightGBM model to predict the daily load of the substation area, input the feature dataset Z' of the test set into the optimal LightGBM model, and finally output the predicted results of the test set substation area load.
[0071] 6. Prediction result evaluation
[0072] Mean Absolute Error (MAE), the formula is
[0073]
[0074] Mean Square Error (MSE), the formula is
[0075]
[0076] pre h is the actual value of the substation area at time h, is the predicted value of the power grid area at time h, and p is the total number of prediction times. The prediction results are evaluated by the above method and compared with other methods to verify the prediction effect of this method.
[0077] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for predicting the load in a power distribution area, characterized in that, Including: Obtain historical data, where the historical data includes multiple historical records, and each historical record includes corresponding substation area load, date, time, and influencing factors; the influencing factors include temperature and humidity; Use the Prophet model to perform feature analysis and extraction on the historical data to obtain a training set feature data set A and a test set feature data set B; Utilize the unsupervised learning algorithm of the autoencoder to perform automated feature extraction and non-linear dimensionality reduction processing on the training set feature data set A to obtain a feature data set Z; utilize the unsupervised learning algorithm of the autoencoder to perform automated feature extraction and non-linear dimensionality reduction processing on the test set feature data set B to obtain a feature data set Z'; Train a LightGBM model based on the feature data set Z and the feature data set Z'; the LightGBM model is used for substation area load prediction; For any historical record, before using the Prophet model to perform feature analysis and extraction on the historical data, it further includes: When any influencing factor value is null, perform missing value processing: y(t, d) is the data of the d-th type of data at time t, y(t - 1, d) is the data of the d-th type of data at time t - 1, and y(t + 1, d) is the data of the d-th type of data at time t + 1; Compare the value of any influencing factor with the base value of this influencing factor. If the difference is more than 20%, correct the value of this influencing factor to: ; The influencing factors further include: unconventional factors and their durations; the unconventional factors include large-scale events, large-scale conferences, and industrial production plans; the duration is the difference between the end time and the start time of the unconventional factor; The steps for obtaining the unconventional factors include: According to the different types of substation area loads, determine the influence degree of a certain unconventional factor on the substation area load type; Obtain the start and end times of a certain unconventional factor in the local area within the historical date; The use of the Prophet model to perform feature analysis and extraction on the historical data to obtain a training set feature data set A and a test set feature data set B specifically includes: Divide the historical data into a training data set and a test data set according to a ratio of 8:2; Input the load data of the training data set and the corresponding temperature and humidity into Prophet, then add monthly, weekly, daily, and holiday assignments to the model, then add unconventional factors to the model, and finally use the additive model and fitting function of Prophet to obtain the fitted load of the training data set; the additive model formula of Prophet is as follows: The load data fitted by Prophet is divided into 7 parts: It represents the overall trend term in the fitted load data, It represents the monthly trend term in the fitted load data, It represents the weekly trend term in the fitted load data, It represents the daily trend term in the fitted load data, It is the impact of holidays on the fitted load data, It is the impact of changes in temperature and humidity on the fitted load data, It is the impact of unconventional factors on the load. Then, the above 8 items of data are used as the training feature dataset A; Input the temperature, humidity data, and unconventional factors included in the test data set into the additive model, use its prediction function to predict the load in the test data set, the prediction results of the test data set also include the above 8 items of data, and finally use it as the test set feature data set B; The use of the unsupervised learning algorithm of the autoencoder to perform automated feature extraction and non-linear dimensionality reduction processing on the training set feature data set A to obtain a feature data set Z specifically includes: Normalize the feature data set A to obtain a feature data set A*; The normalized feature data set A* is used as the input data of the autoencoder, and the autoencoder uses an f function to map the input data to the hidden layer, specifically: ; the f function takes the Relu function; Modify the parameters of the f function until the loss function meets the set range; where the output Z of the hidden layer is the hidden feature variable, and the hidden feature variable is used to reconstruct , specifically: ; g takes the Sigmoid function.
2. The substation area load forecasting system is characterized in that, Including: The first acquisition unit is used to acquire historical data, where the historical data includes multiple historical records, and each historical record includes corresponding substation area load, date, time, and influencing factors; the influencing factors include temperature and humidity; The first feature extraction unit is used to perform feature analysis and extraction on the historical data using the Prophet model to obtain a training set feature data set A and a test set feature data set B; The second feature extraction unit uses the unsupervised learning algorithm of the autoencoder to perform automated feature extraction and non-linear dimensionality reduction processing on the training set feature data set A to obtain a feature data set Z; uses the unsupervised learning algorithm of the autoencoder to perform automated feature extraction and non-linear dimensionality reduction processing on the test set feature data set B to obtain a feature data set Z'; The first processing unit is used to train a LightGBM model according to the feature data set Z and the feature data set Z'; the LightGBM model is used for substation area load prediction; The substation area load prediction system further includes a preprocessing unit, and the preprocessing unit is used for: For any historical record, before performing feature analysis and extraction on the historical data using the Prophet model, it further includes: When any influencing factor value is null, perform missing value processing: y(t, d) is the data of the d-th type of data at time t, y(t - 1, d) is the data of the d-th type of data at time t - 1, and y(t + 1, d) is the data of the d-th type of data at time t + 1; Compare the value of any influencing factor with the base value of the influencing factor. If the difference is more than 20%, correct the value of the influencing factor to: ; The influencing factors further include: unconventional factors and their durations; the unconventional factors include large-scale events, large-scale conferences, and industrial production plans; the duration is the difference between the end time and the start time of the unconventional factor; The steps for obtaining the unconventional factors include: According to the different types of substation area loads, determine the influence degree of a certain unconventional factor on the substation area load type; Obtain the start and end times of a certain unconventional factor in the local area within the historical date; Performing feature analysis and extraction on the historical data using the Prophet model to obtain a training set feature data set A and a test set feature data set B specifically includes: The substation area load prediction system further includes: The second processing unit is used to divide the historical data into a training data set and a test data set according to a ratio of 8:2; Input the load data of the training data set and the corresponding temperature and humidity into Prophet, then add monthly, weekly, daily, and holiday assignments to the model, then add unconventional factors to the model, and finally use the additive model and fitting function of Prophet to obtain the fitted load of the training data set; The additive model formula of Prophet is as follows: The load data fitted by Prophet is divided into 7 parts: Represents the overall trend term in the fitted load data, Represents the monthly trend term in the fitted load data, Represents the weekly trend term in the fitted load data, Represents the daily trend term in the fitted load data, Is the impact of holidays on the fitted load data, Is the impact of changes in temperature and humidity on the fitted load data, Is the impact of unconventional factors on the fitted load. Then, the above 8 items of data obtained are used as the training feature data set A; Input the temperature, humidity data and unconventional factors included in the test data set into the additive model, use its prediction function to predict the load in the test data set, and the prediction results of the test data set also include the above 8 items of data, and finally use it as the test set feature data set B; The second feature extraction unit is specifically configured to: Normalize the feature data set A to obtain the feature data set A*; Use the normalized feature dataset A* as the input data of the autoencoder. The autoencoder uses an f function to map the input data to the hidden layer. Specifically: ; The f function takes the Relu function; Modify the parameters of the f function until the loss function meets the set range; where the output Z of the hidden layer is the hidden feature variable, and the hidden feature variable is used to reconstruct , specifically: ; g takes the Sigmoid function.
3. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the substation area load forecasting method described in claim 1 are implemented.
Citation Information
Patent Citations
Electric power load forecasting method, device, equipment and storage medium
CN109034490A
Day-ahead power load prediction method and device based on double-end automatic coding
CN111144643A
Short-term power consumption load prediction method
CN112215426A
SF6 equipment gas pressure prediction method based on Prophet-LSTM model
CN114065667A