Distribution transformer weight and overload prediction method and system based on daily granularity data, medium and equipment
Through the distribution-change heavy overload prediction method based on daily granular data, multiple models are built for load prediction and classification, which solves the problems of low prediction hit rate, high false alarm rate and insufficient response time in the prior art, and achieves more accurate and effective heavy overload prediction.
Patent Information
- Application Number
- CN202411812095.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-05-27
AI Technical Summary
In the real-time technology, the prediction of low hit rate, high false alarm rate and inability to effectively deal with short-term high loads in the power distribution transformer heavy overload prediction, and the ultra-short-term load prediction based on 15 min granular data cannot provide sufficient response time.
The matching variable heavy overload prediction method based on daily granularity data is adopted. By collecting daily granularity data in the past year, a classification model, the first regression model and the second regression model are constructed to predict the maximum load and average load, and a classification prediction is carried out in combination with the classification model.
It improves the accuracy of prediction, effectively solves the problem of poor model prediction results caused by sample imbalance, and provides longer response time so that users can take effective measures to avoid heavy overload.
Smart Images

Figure CN120046124A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distribution transformer load prediction, and in particular to a distribution transformer heavy overload prediction method, system, medium and equipment based on daily granularity data. Background Art
[0002] Distribution transformers (abbreviated as distribution transformers) are important nodes for power transmission to users, and their load conditions are crucial to the quality and safety of power supply in the region. If the distribution transformer is overloaded for a long time, it will shorten the service life of the equipment, increase the probability of line failure, cause the voltage of users at the end of the low-voltage line to drop, and affect the normal operation of the power equipment of users in the region.
[0003] The massive amount of equipment data has brought new challenges to the prediction and identification of heavy overload. At present, there are few studies directly targeting heavy overload prediction, and most studies still focus on the regression prediction of power load. These studies generally ignore the complex relationship between power load and heavy overload of distribution transformers. Due to the volatility of power load, high load is only a necessary condition for heavy overload to occur, and short-term high load will not trigger a heavy overload alarm of distribution transformers.
[0004] The existing research methods for heavy overload prediction mainly include those based on random forest models and neural networks. The research method based on random forest models has limited effect in dealing with sample imbalance problems. The model has a low prediction hit rate and a high false alarm rate in practical applications. The method based on neural networks is designed with a two-level filtering system. Its data filtering rules are too simple and are based only on the load data of the same period last year. It has major defects. At the same time, its research is based on 15-minute granularity data for power load prediction. This ultra-short-term load prediction leaves users with less response time and often fails to take effective measures to avoid the occurrence of heavy overload problems. Summary of the invention
[0005] Based on this, it is necessary to propose a distribution transformer overload prediction method, system, medium and equipment based on daily granularity data to address the above problems.
[0006] A distribution transformer heavy overload prediction method based on daily granularity data, the method comprising:
[0007] Collect daily granular data of each distribution transformer in the past year.
[0008] A classification model, a first regression model and a second regression model are constructed according to the daily granularity data.
[0009] The daily maximum load of the same period last year and the daily maximum load of this year on the current date in the daily granularity data are input into the first regression model to obtain the predicted value of the maximum load of the distribution transformer at time point t+N, where t is the time point and N is the number of days.
[0010] The daily average load of the same period last year and the daily average load of this year are matched according to the predicted value of the maximum load of the distribution transformer, and the matched daily average load of the same period last year and the daily average load of this year are input into the second regression model to obtain the predicted value of the average load at time point t+N.
[0011] The predicted value of the maximum load and the predicted value of the average load at the time point t+N replace the daily maximum load value and the daily load average of the corresponding date in the daily granularity data, and input them into the classification model to obtain the heavy overload classification prediction result at the time point t+N.
[0012] Wherein, constructing a classification model, a first regression model and a second regression model according to the daily granularity data specifically includes:
[0013] The daily granularity data is divided into an initial training set and an initial test set according to the time series.
[0014] A first preset number of daily granularity data in the initial training set is extracted by random sampling as a classification model training set, and a second preset number of daily granularity data in the initial training set is extracted as a classification model test set.
[0015] A classification feature set is extracted from the classification model training set.
[0016] A classification model is constructed using the classification feature set, and the classification model is evaluated using the classification model test set.
[0017] The daily granular data in the initial training set is divided into a regression model training set and a regression model test set according to the time series.
[0018] The regression model characteristics are determined according to the power load data in the regression model training set, wherein the power load data includes a daily maximum load and a daily average load, a first regression model and a second regression model are constructed through the regression model characteristics, and the first regression model and the second regression model are evaluated through the regression model test set.
[0019] The step of extracting the classification feature set from the classification model training set specifically includes:
[0020] The classification model features are determined according to the distribution transformer geographical location information, distribution transformer attribute information, time and load data in the classification model training set.
[0021] A categorical variable is set, wherein a value set of the categorical variable includes a plurality of different categorical values in lexicographic order.
[0022] A mapping function is defined to map each classification component to a unique natural number as the first classification feature of the non-numeric feature.
[0023] The numerical feature in the classification model feature is used as the second classification feature, and a classification feature set is constructed according to the first classification feature and the second classification feature.
[0024] The step of constructing a classification model by using the classification feature set and evaluating the classification model by using the classification model test set specifically includes:
[0025] The classification model objective function is defined based on the logistic regression loss function.
[0026] The latest tree is generated by the classification model objective function, the label corresponding to each classification feature in the classification feature set is used as the prediction target, the classification features in the classification feature set are used as input samples, and the classification model is constructed by adding the latest tree, and the labels include non-overloaded, overloaded and overloaded.
[0027] The step of determining the regression model characteristics according to the power load data in the regression model training set, wherein the power load data includes a daily maximum load and a daily average load, constructing a first regression model and a second regression model through the regression model characteristics, and evaluating the first regression model and the second regression model through the regression model test set, specifically includes:
[0028] The regression model characteristics are determined based on the power load data in the first preset time period of the same period last year and the power load data in the second preset time period of this year in the regression model training set, wherein the power load data includes a daily maximum load and a daily average load.
[0029] The regression model objective function is defined based on the squared error loss function.
[0030] The latest tree is generated through the regression model objective function, the maximum load at a historical time point is used as the prediction target, the daily maximum load in the regression model characteristics is used as the input sample, and the first regression model is constructed by adding the latest tree.
[0031] The latest tree is generated by the regression model objective function, the average load at the historical time point is used as the prediction target, the daily average load in the regression model feature is used as the input sample, and the second regression model is constructed by adding the latest tree.
[0032] The daily average load of the same period last year and the daily average load of this year are matched according to the predicted value of the maximum load of the distribution transformer, and the matched daily average load of the same period last year and the daily average load of this year are input into the second regression model to obtain the predicted value of the average load at the time point t+N, which specifically includes:
[0033] According to the comparison between the predicted value of the maximum load of the distribution transformer and the load matching threshold, the distribution transformers are screened to obtain a distribution transformer list.
[0034] The daily average load of the same period last year and the daily average load of this year in the daily granularity data are matched with the distribution transformer list to obtain the matched daily average load of the same period last year and the daily average load of this year.
[0035] The matched daily average load of the same period last year and the daily average load of this year are input into the second regression model to obtain the predicted value of the average load at the time point t+N.
[0036] The step of screening the distribution transformers to obtain a distribution transformer list according to the comparison between the predicted value of the maximum load of the distribution transformer and the load matching threshold specifically includes:
[0037] It is determined whether the predicted value of the maximum load of the distribution transformer is greater than or equal to a load matching threshold.
[0038] If the predicted value of the maximum load of the distribution transformer is greater than or equal to the load matching threshold, the maximum load of the distribution transformer is retained and a distribution transformer list is obtained.
[0039] If the predicted value of the maximum load of the distribution transformer is less than the load matching threshold, the maximum load of the distribution transformer is not retained.
[0040] Wherein, replacing the daily maximum load value and the daily load average of the corresponding date in the daily granularity data with the predicted value of the maximum load and the predicted value of the average load at the time point t+N, inputting them into the classification model, and obtaining the heavy overload classification prediction result at the time point t+N, further comprising:
[0041] The heavy overload classification prediction result at the time point t+N is compared with the daily granularity data of the corresponding date in the initial test set to determine the hit rate and the false alarm rate.
[0042] Based on the hit rate and false alarm rate, hyperparameter tuning is performed on the classification voting weight and load matching threshold of the classification model.
[0043] A distribution transformer overload prediction system based on daily granularity data, the system comprising:
[0044] The collection module is used to collect daily granular data of distribution transformers in the past year.
[0045] A model building module is used to build a classification model, a first regression model and a second regression model according to the daily granularity data.
[0046] The maximum load prediction module is used to input the daily maximum load of the same period last year and the daily maximum load of this year in the daily granularity data into the first regression model to obtain the predicted value of the maximum load of the distribution transformer at the time point t+N, where t is the time point and N is the number of days.
[0047] The average load prediction module is used to match the predicted value of the maximum load of the distribution transformer with the daily average load of the same period last year and the daily average load of this year, and input the matched daily average load of the same period last year and the daily average load of this year into the second regression model to obtain the predicted value of the average load at time point t+N.
[0048] The heavy overload classification prediction module is used to replace the daily maximum load value and the daily load average of the corresponding date in the daily granularity data with the predicted value of the maximum load and the predicted value of the average load at the t+N time point, input the classification model, and obtain the heavy overload classification prediction result at the t+N time point.
[0049] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the steps of the method described above.
[0050] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.
[0051] The embodiments of the present invention have the following beneficial effects:
[0052] The training data of the model of the present invention is derived from the daily granular data of each distribution transformer in the city in the past year. These data are the data that have been recorded in daily monitoring, and no additional complex data collection process is required. In addition, a secondary prediction method is adopted. The maximum load prediction value and the average load prediction value at the t+N time point are first predicted by the first regression model and the second regression model, and then the classification model is combined to perform heavy overload classification prediction, which further improves the accuracy of the prediction. At the same time, this solution effectively solves the problem of poor model prediction effect caused by sample imbalance by matching and filtering the daily granular data, so that the model can better cope with this common data distribution in practical applications, thereby improving the practicality of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0054] in:
[0055] Figure 1 A flow chart of another embodiment of a distribution transformer heavy overload prediction method based on daily granularity data provided by the present invention;
[0056] Figure 2 A schematic flow chart of another embodiment of a distribution transformer overload prediction method based on daily granularity data provided by the present invention;
[0057] Figure 3 A schematic diagram of the structure of an embodiment of a distribution transformer overload prediction system based on daily granularity data provided by the present invention;
[0058] Figure 4 A schematic diagram of the structure of an embodiment of the device provided by the present invention;
[0059] Figure 5 A schematic structural diagram of an embodiment of the medium provided by the present invention. DETAILED DESCRIPTION
[0060] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0061] The present invention provides an embodiment of a method for predicting a distribution transformer overload based on daily granularity data. A method for predicting a distribution transformer overload based on daily granularity data comprises:
[0062] S101: Collect daily granular data of each distribution transformer in the past year.
[0063] For example, the collected daily granularity data D = {x ij ,(i=1,2,…,n,j=1,2,…,m)}, where i is a daily report record of a distribution transformer on a certain day, and there are n records in total, and j is an attribute of a distribution transformer on a certain day, and there are m different attributes in total.
[0064] S102: Constructing a classification model, a first regression model and a second regression model according to the daily granularity data.
[0065] Exemplarily, according to the characteristics of time series prediction, the daily data D is divided into an initial training set and an initial test set according to the time series, wherein the training set is the data of the previous 11 months, and the test set is the data of the most recent month.
[0066] The initial training set is randomly sampled, 80% of the data is extracted as the classification model training set, and 20% of the data is extracted as the classification model test set.
[0067] The classification model features are determined based on the distribution transformer geographical location information, distribution transformer attribute information, time and load data in the classification model training set, where:
[0068] The geographical location information of the distribution transformer includes:
[0069] Counties and cities where distribution transformers are located: Factors such as economic development level, industrial structure, and electricity consumption habits in different counties and cities may lead to differences in power load demand. This information helps the model distinguish the electricity consumption characteristics of different regions.
[0070] Power supply station: The power supply range, power supply capacity and user types of the power supply station will also affect the distribution transformer load, providing the model with more detailed regional power consumption information.
[0071] 10kv line: The aging degree of the line, load capacity and the conditions of other connected equipment will affect the actual load conditions of the distribution transformer. As model input, it can further refine the description of the distribution transformer operating environment.
[0072] The distribution transformer attribute information includes:
[0073] Transformer cooling type: Transformers with different cooling types (such as oil-immersed, dry type, etc.) have different heat dissipation performance, which will affect the stability and reliability of the transformer when operating under high load and is one of the important factors in determining the risk of heavy overload.
[0074] Distribution transformer capacity: The distribution transformer capacity directly determines the maximum load it can carry and is one of the key indicators for measuring whether a distribution transformer may be severely overloaded.
[0075] Operating life of distribution transformer: With the increase of operating life, the aging, loss and performance degradation of distribution transformer equipment will gradually appear, affecting its load capacity. Therefore, the operating life is an important input feature of the model.
[0076] Time and load data include:
[0077] dayofyear: different time nodes in a year. Due to seasonal changes, holidays and other factors, power load demand often presents different patterns. This feature can help the model capture the power consumption change pattern in this time series.
[0078] The average maximum load of the day: reflects the maximum load pressure borne by the distribution transformer on that day, which is of great significance for judging whether it is close to or exceeds the load limit of the distribution transformer.
[0079] Average load of the day: reflects the average load level of the distribution transformer on that day. Combined with the average maximum load, it can more comprehensively describe the load status of the distribution transformer on that day.
[0080] Furthermore, we set the categorical variables, and the value set of the categorical variables is {c 1 ,c 2 ,…,c n}, where n is the number of different values of the categorical variable, c i The i-th category after all category values are sorted in lexicographic order.
[0081] Define a mapping function f:C→N, where N is a set of natural numbers, and use the mapping function to map each classification component to a unique natural number m i , and satisfy m 1 <m 2 <… <m n , as the first classification feature of non-numeric features.
[0082] The numerical feature in the classification model feature is used as the second classification feature, and a classification feature set is constructed according to the first classification feature and the second classification feature.
[0083] Furthermore, the classification model objective function is defined according to the logistic regression loss function; the latest tree is generated through the classification model objective function, and the label corresponding to each classification feature in the classification feature set is used as the prediction target, and the classification features in the classification feature set are used as input samples. The classification model is constructed by adding the latest tree, and the labels include non-overloaded, overloaded and overloaded, and the classification model is evaluated through the classification model test set.
[0084] According to the characteristics of time series prediction, the daily granularity data of the first 10 months in the initial training set are used as the regression model training set according to the time series, and the daily granularity data of the 11th month is used as the regression model test set.
[0085] The regression model features are determined based on the power load data in the regression model training set, where the power load data includes the daily maximum load and the daily average load. Specifically, x t represents this year's data, x′ tRepresents last year’s data, where t represents a certain time point, and t±k represents k days before and after time point t.
[0086] Data for the same period last year: Power load data for the 5 days before and after the same period last year, including the daily maximum load d′ max ={x′ i ,i={t-5,t-4,…,t+5}} and the daily average load
[0087] Recent load data: Power load data for the past five days of this year, including the daily maximum load max =
[0088] {x i ,i={t-5,t-4,…,t-1}} and the daily average load
[0089] Furthermore, regression model characteristics are determined based on power load data in the regression model training set, where the power load data includes daily maximum load and daily average load. A first regression model and a second regression model are constructed based on the regression model characteristics, and the first regression model and the second regression model are evaluated using the regression model test set.
[0090] S103: Input the daily maximum load of the same period last year and the daily maximum load of this year on the current date in the daily granularity data into the first regression model to obtain the predicted value of the maximum load of the distribution transformer at the time point t+N, where t is the time point and N is the number of days.
[0091] For example, the recent daily maximum load d max ={x i ,i={t-5,t-4,…,t-1}}Daily maximum load compared with the same period last year The daily maximum load of the same period last year and the daily maximum load of this year are input into the first model to obtain the maximum load forecast value of each distribution transformer at time t+2. Indicates that n is the number of all distribution transformers. The maximum load forecast value at time t+2 is the maximum load forecast value of the distribution transformer at time t tomorrow, based on the daily granularity data up to yesterday at time t today.
[0092] S104: Match the daily average load of the same period last year and the daily average load of this year according to the predicted value of the maximum load of the distribution transformer, input the matched daily average load of the same period last year and the daily average load of this year into the second regression model to obtain the predicted value of the average load at time point t+N.
[0093] Exemplarily, determine whether the predicted value of the maximum load of the distribution transformer is greater than or equal to the load matching threshold; if the predicted value of the maximum load of the distribution transformer is greater than or equal to the load matching threshold, the maximum load of the distribution is retained and the distribution transformer list is obtained; if the predicted value of the maximum load of the distribution transformer is less than the load matching threshold, the maximum load of the distribution is not retained.
[0094] Furthermore, the daily average load of the same period last year and the daily average load of this year in the daily granularity data are matched with the distribution transformer list to obtain the filtered daily average load Average daily load in the same period last year The matched data is input into the regression model R2 for load mean prediction, and the average load prediction value at time t+2 is obtained. express.
[0095] S105: Replace the daily maximum load value and daily load average of the corresponding date in the daily granularity data with the predicted value of the maximum load and the predicted value of the average load at the time point t+N, input them into the classification model, and obtain the heavy overload classification prediction result at the time point t+N.
[0096] For example, the maximum load prediction value at time t+2 is The average load forecast value at time t+2 Replace the daily maximum load value and daily load mean of the corresponding date of the classification feature set to obtain a new feature set. Input the new feature set data into the classification model to obtain the heavy overload classification prediction result y at time t+2 c ,y c ∈(non-heavy overload, heavy load, overload).
[0097] From the above description, it can be known that the training data of the model of the present invention is derived from the daily granular data of each distribution transformer in the city in the past year. These data are the data that have been recorded in daily monitoring, and no additional complex data collection process is required. In addition, a secondary prediction method is adopted, first through the first regression model, the second regression model first predicts the maximum load prediction value and the average load prediction value at the t+N time point, and then combines the classification model to perform heavy overload classification prediction, which further improves the accuracy of the prediction. At the same time, this solution effectively solves the problem of poor model prediction effect caused by sample imbalance by matching and filtering the daily granular data, so that the model can better cope with this common data distribution in practical applications, thereby improving the practicality of the model.
[0098] like Figure 1 As shown, Figure 1 A flow chart of another embodiment of a method for predicting distribution transformer overload based on daily granularity data provided by the present invention. A method for predicting distribution transformer overload based on daily granularity data, the method comprising:
[0099] S201: Collect daily granular data of each distribution transformer in the past year.
[0100] S202: Divide the daily granularity data into an initial training set and an initial test set according to the time series.
[0101] Exemplarily, according to the characteristics of time series prediction, the daily data D is divided into an initial training set and an initial test set according to the time series, wherein the initial training set is the data of the previous 11 months, and the initial test set is the data of the most recent month.
[0102] S203: extracting a first preset number of daily granularity data in the initial training set as a classification model training set by a random sampling method, and using a second preset number of daily granularity data in the initial training set as a classification model test set.
[0103] Exemplarily, the initial training set is randomly sampled to extract 80% of the data as the classification model training set and 20% of the data as the classification model test set.
[0104] S204: Extracting a classification feature set from the classification model training set.
[0105] Exemplarily, the classification model features are determined according to the distribution transformer geographical location information, distribution transformer attribute information, time and load data in the classification model training set; the classification variable is set, and the value set of the classification variable contains several different classification values in dictionary order; a mapping function is defined, and each classification component is mapped to a unique natural number through the mapping function as the first classification feature of the non-numeric feature; the numerical feature in the classification model feature is used as the second classification feature, and a classification feature set is constructed according to the first classification feature and the second classification feature.
[0106] S205: Define a classification model objective function according to the logistic regression loss function.
[0107] Exemplarily, the classification model objective function consists of two parts: the logistic regression loss function and the regularization term
[0108] The logistic regression loss function used by the classification model is:
[0109]
[0110] Among them, L 1 is the logistic regression loss function, y i1 is the true value of the label corresponding to the i-th input sample, is the label prediction value corresponding to the i-th input sample.
[0111] The regular term Ω is used to control the complexity of the model and prevent overfitting. Its form is:
[0112]
[0113] Among them, T is the number of leaf nodes, w j is the weight of the jth leaf node, λ is the leaf weight regularization parameter, γ is the leaf number regularization parameter, and z is the round.
[0114] S206: Generate the latest tree through the classification model objective function, take the label corresponding to each classification feature in the classification feature set as the prediction target, take the classification features in the classification feature set as input samples, and build the classification model by adding the latest tree. The labels include non-overloaded, overloaded and overloaded.
[0115] Exemplarily, the average load at historical time points is used as the prediction target, and the daily average load in the regression model characteristics is used as the input sample for modeling to determine the initial prediction value; based on the gain function, the left child node and the right child node when the gain function is minimized are determined according to the first-order derivative and the second-order derivative of the objective function of the regression model; the latest tree is generated through the left child node and the right child node; the latest tree is fit to the initial prediction value, and the regression model of the current round is determined by adding the latest tree; prediction is performed through the regression model of the current round to obtain the prediction value of the current round; the above steps are repeated until the preset number of trees is reached, and the second regression model is determined.
[0116] S207: Divide the daily granularity data in the initial training set into a regression model training set and a regression model test set according to the time series.
[0117] For example, according to the characteristics of time series prediction, the daily granularity data of the first 10 months in the initial training set are used as the regression model training set according to the time series, and the daily granularity data of the 11th month are used as the regression model test set.
[0118] The regression model features are determined based on the power load data in the regression model training set, where the power load data includes the daily maximum load and the daily average load. Specifically, x t represents this year's data, x′ t Represents last year’s data, where t represents a certain time point, and t±k represents k days before and after time point t.
[0119] Data for the same period last year: Power load data for the 5 days before and after the same period last year, including the daily maximum load d′ max ={x′ i ,i={t-5,t-4,…,t+5}} and the daily average load
[0120] Recent load data: Power load data for the past five days of this year, including the daily maximum load max =
[0121] {x i ,i={t-5,t-4,…,t-1}} and the daily average load
[0122] S208: Determine the regression model characteristics based on the power load data in the first preset time period of the same period last year and the power load data in the second preset time period of this year in the regression model training set, where the power load data includes the daily maximum load and the daily average load.
[0123] Exemplarily, the daily granularity data in the initial training set is divided into a regression model training set and a regression model test set according to the time series. The regression model features are determined based on the power load data of the same period last year within the 5 days before and after and this year within the past 5 days in the regression model training set, and the power load data includes the daily maximum load and the daily average load; the regression model objective function is defined according to the square error loss function; the latest tree is generated through the regression model objective function, and the maximum load at the historical time point is used as the prediction target, and the daily maximum load in the regression model feature is used as the input sample, and the first regression model is constructed by adding the latest tree; the latest tree is generated through the regression model objective function, and the average load at the historical time point is used as the prediction target, and the daily average load in the regression model feature is used as the input sample, and the second regression model is constructed by adding the latest tree.
[0124] S209: Define the regression model objective function according to the square error loss function.
[0125] Exemplarily, the regression model objective function consists of two parts: the squared error loss function and the regularization term
[0126] The squared error loss function used by the regression model is:
[0127]
[0128] Among them, L 2 is the square error loss function, y i2 is the true value of the load at the historical time point in the i-th input sample, is the load forecast value at the historical time point in the i-th input sample. The load includes: maximum load and average load.
[0129] The regular term Ω is used to control the complexity of the model and prevent overfitting. Its form is:
[0130]
[0131] Among them, T is the number of leaf nodes, w j is the weight of the jth leaf node, λ is the leaf weight regularization parameter, γ is the leaf number regularization parameter, and z is the round.
[0132] S210: Generate the latest tree through the regression model objective function, take the maximum load at the historical time point as the prediction target, take the daily maximum load in the regression model feature as the input sample, and build the first regression model by adding the latest tree.
[0133] Exemplarily, the maximum load at a historical time point is used as the prediction target, and the daily maximum load in the regression model feature is used as the input sample for modeling to determine the initial prediction value; based on the gain function, the left child node and the right child node when the gain function is minimum are determined according to the first-order derivative and the second-order derivative of the objective function of the regression model; the latest tree is generated through the left child node and the right child node; the latest tree is fit to the initial prediction value, and the regression model of the current round is determined by adding the latest tree; prediction is performed through the regression model of the current round to obtain the prediction value of the current round; the above steps are repeated until the preset number of trees is reached to determine the first regression model.
[0134] In addition, the regression model test set is used to evaluate and verify the regression model, and its MSE and MAE are calculated. Among them, y i2 is the true value of the load at the historical time point in the i-th input sample, is the load forecast value at the historical time point in the i-th input sample, and N represents the total number of samples; The MSE and MAE are used to understand the regression model fitting. S211: Generate the latest tree through the regression model objective function, take the average load at the historical time point as the prediction target, take the daily average load in the regression model feature as the input sample, and build the second regression model by adding the latest tree.
[0135] Exemplarily, the average load at historical time points is used as the prediction target, and the daily average load in the regression model characteristics is used as the input sample for modeling to determine the initial prediction value; based on the gain function, the left child node and the right child node when the gain function is minimized are determined according to the first-order derivative and the second-order derivative of the objective function of the regression model; the latest tree is generated through the left child node and the right child node; the latest tree is fit to the initial prediction value, and the regression model of the current round is determined by adding the latest tree; prediction is performed through the regression model of the current round to obtain the prediction value of the current round; the above steps are repeated until the preset number of trees is reached, and the second regression model is determined.
[0136] S212: Input the daily maximum load of the same period last year and the daily maximum load of this year on the current date in the daily granularity data into the first regression model to obtain the predicted value of the maximum load of the distribution transformer at the time point t+N, where t is the time point and N is the number of days.
[0137] For example, the recent daily maximum load d max ={x i ,i={t-5,t-4,…,t-1}}Daily maximum load compared with the same period last year The daily maximum load of the same period last year and the daily maximum load of this year are input into the first model to obtain the maximum load forecast value of each distribution transformer at time t+2. Indicates that n is the number of all distribution transformers. The maximum load forecast value at time t+2 is the maximum load forecast value of the distribution transformer at time t tomorrow, based on the daily granularity data up to yesterday at time t today.
[0138] S213: According to the comparison between the predicted value of the maximum load of the distribution transformer and the load matching threshold, the distribution transformers are screened to obtain a distribution transformer list.
[0139] For example, the distribution transformers are screened according to the load matching threshold K, and the screening condition is the predicted value of the maximum load Get the distribution transformer list S that meets the conditions. At this time, the maximum load prediction value corresponding to the list n s is the number of distribution transformers in list S.
[0140] S214: Match the daily average load of the same period last year and the daily average load of this year in the daily granularity data with the distribution transformer list to obtain the matched daily average load of the same period last year and the daily average load of this year.
[0141] For example, the daily average load of the same period last year and the daily average load of this year in the daily granularity data are matched with the distribution transformer list to obtain the filtered daily average load Average daily load in the same period last year The matched data is input into the regression model R2 for load mean prediction, and the average load prediction value at time t+2 is obtained. express.
[0142] S215: Input the matched daily average load of the same period last year and the daily average load of this year into the second regression model to obtain a predicted value of the average load at time point t+N.
[0143] For example, the maximum load prediction value at time t+2 is The average load forecast value at time t+2 Replace the daily maximum load value and daily load mean of the corresponding date of the classification feature set to obtain a new feature set. Input the new feature set data into the classification model to obtain the heavy overload classification prediction result y at time t+2 c ,y c ∈(non-heavy overload, heavy load, overload).
[0144] It can be seen from the above description that the present invention filters data by setting a load matching threshold, excludes distribution transformer data with load values less than the load matching threshold, reduces unnecessary calculations and misjudgments, and makes the prediction results more focused on distribution transformers that may be heavily overloaded, thereby improving the accuracy of the overall prediction. In addition, the classification model, the first regression model, and the second regression model of the present invention are all implemented using XGBoost. XGBoost is based on a gradient boosting framework and optimizes the model by iteratively constructing and combining weak classifiers or weak regressors, and has high prediction performance. It can effectively learn complex patterns in data and has good adaptability to complex nonlinear relationship problems such as distribution transformer heavy overload prediction.
[0145] like Figure 2 As shown, Figure 2 A flow chart of another embodiment of a method for predicting distribution transformer overload based on daily granularity data provided by the present invention. A method for predicting distribution transformer overload based on daily granularity data, the method comprising:
[0146] S301: Compare the heavy overload classification prediction result at time point t+N with the daily granularity data of the corresponding date in the initial test set to determine the hit rate and false alarm rate.
[0147] Exemplarily, the heavy overload classification prediction result at time point t+N is compared with the daily granularity data of the corresponding date in the initial test set, and the hit rate and false alarm rate of the current date are calculated to obtain 1 evaluation result.
[0148] The heavy overload predictions for all dates in the initial test set are obtained in turn using the classification model, the first regression model, and the second regression model. The overall hit rate and false alarm rate are calculated to evaluate the model performance and stability. Finally, hyperparameters are tuned based on the performance of the model indicators.
[0149] S302: Based on the hit rate and the false alarm rate, hyperparameter tuning is performed on the classification voting weight and the load matching threshold of the classification model.
[0150] Exemplarily, the model hyperparameters are the load matching threshold and the classification voting weight of the classification model. First, adjust the load matching threshold and observe the changes in the hit rate and false alarm rate. Then, adjust the voting weight and observe the changes in the hit rate and false alarm rate. By constantly trying different values, find the best hyperparameter combination that makes the hit rate high and the false alarm rate within an acceptable range. After hyperparameter tuning, the heavy overload hit rate of the model is greatly improved, while the false alarm rate can still be controlled within a small range.
[0151] like Figure 3 As shown, Figure 3 A schematic diagram of the structure of an embodiment of a distribution transformer overload prediction system based on daily granularity data provided by the present invention. A distribution transformer overload prediction system 10 based on daily granularity data, the system comprises:
[0152] The collection module 11 is used to collect daily granular data of the distribution transformer in the past year.
[0153] The model building module 12 is used to build a classification model, a first regression model and a second regression model according to the daily granularity data.
[0154] The maximum load prediction module 13 is used to input the daily maximum load of the same period last year and the daily maximum load of this year in the daily granularity data into the first regression model to obtain the predicted value of the maximum load of the distribution transformer at the time point t+N, where t is the time point and N is the number of days.
[0155] The average load prediction module 14 is used to match the predicted value of the maximum load of the distribution transformer with the daily average load of the same period last year and the daily average load of this year, and input the matched daily average load of the same period last year and the daily average load of this year into the second regression model to obtain the predicted value of the average load at the time point t+N.
[0156] The heavy overload classification prediction module 15 is used to replace the daily maximum load value and daily load average of the corresponding date in the daily granularity data with the predicted value of the maximum load and the predicted value of the average load at the time point t+N, input the classification model, and obtain the heavy overload classification prediction result at the time point t+N.
[0157] Exemplarily, in the acquisition module 11, the daily granular data of the distribution transformer in the past year are collected. In the model construction module 12, the daily granular data are divided into an initial training set and an initial test set according to the time series; the first preset number of daily granular data in the initial training set is extracted as the classification model training set by random sampling, and the second preset number of daily granular data in the initial training set is used as the classification model test set; the classification feature set in the classification model training set is extracted; the classification model is constructed by the classification feature set, and the classification model is evaluated by the classification model test set; the daily granular data in the initial training set is divided into a regression model training set and a regression model test set according to the time series; the regression model features are determined according to the power load data in the regression model training set, the power load data includes the daily maximum load and the daily average load, the first regression model and the second regression model are constructed by the regression model features, and the first regression model and the second regression model are evaluated by the regression model test set. In the maximum load prediction module 13, the daily maximum load of the same period last year and the daily maximum load of this year in the daily granularity data are input into the first regression model to obtain the predicted value of the maximum load of the distribution transformer at the time point t+N, where t is the time point and N is the number of days. In the average load prediction module 14, the distribution transformers are screened to obtain the distribution transformer list according to the comparison between the predicted value of the maximum load of the distribution transformer and the load matching threshold; the daily average load of the same period last year and the daily average load of this year in the daily granularity data are matched with the distribution transformer list to obtain the matched daily average load of the same period last year and the daily average load of this year; the matched daily average load of the same period last year and the daily average load of this year are input into the second regression model to obtain the predicted value of the average load at the time point t+N. In the heavy overload classification prediction module 15, the predicted value of the maximum load at the time point t+N and the predicted value of the average load replace the daily maximum load value and the daily load average of the corresponding date in the daily granularity data, and input into the classification model to obtain the heavy overload classification prediction result at the time point t+N.
[0158] like Figure 4 As shown, Figure 4 The device 20 includes a memory 21 and a processor 22. The memory 21 stores a computer program, and the processor 22 executes the computer program when working to implement the following. Figure 1 and Figure 2 The method shown.
[0159] The specific technical details of a distribution transformer overload prediction method based on daily granularity data implemented when the above-mentioned device 20 executes a computer program have been discussed in detail in the above-mentioned method steps, so they will not be repeated here.
[0160] like Figure 5 As shown, Figure 5The structure diagram of an embodiment of the medium provided by the present invention is shown in FIG. The medium 30 stores at least one computer program 31, and the computer program 31 is executed by the processor 22 to implement the following Figure 1 In one embodiment, the medium 30 may be a storage chip, a hard disk, a mobile hard disk, a USB flash drive, an optical disk, or other readable and writable storage tools, or a server, etc.
[0161] The above describes specific embodiments of the present specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily have to be performed in the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0162] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment, and non-volatile computer-readable storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0163] The apparatus, device, non-volatile computer-readable storage medium and method provided in the embodiments of this specification correspond to each other, and therefore, the apparatus, device, and non-volatile computer storage medium also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the corresponding apparatus, device, and non-volatile computer storage medium will not be repeated here.
[0164] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0165] For the convenience of description, the above device is described by being divided into various units according to their functions and described separately. Of course, when implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware. It should be understood by those skilled in the art that this specification embodiment can be provided as a method, system, or computer program product. Therefore, this specification embodiment can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification embodiment can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0166] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0167] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0168] The above disclosure is only the preferred embodiment of the present invention, which certainly cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.
Claims
1. A distribution transformer overload prediction method based on daily granularity data, characterized in that: The method comprises: Collect daily granular data of each distribution transformer in the past year; Constructing a classification model, a first regression model, and a second regression model according to the daily granularity data; Input the daily maximum load of the same period last year and the daily maximum load of this year on the current date in the daily granularity data into the first regression model to obtain the predicted value of the maximum load of the distribution transformer at time point t+N, where t is the time point and N is the number of days; According to the predicted value of the maximum load of the distribution transformer, the daily average load of the same period last year and the daily average load of this year are matched, and the matched daily average load of the same period last year and the daily average load of this year are input into the second regression model to obtain the predicted value of the average load at the time point t+N; The predicted value of the maximum load and the predicted value of the average load at the time point t+N replace the daily maximum load value and the daily load average of the corresponding date in the daily granularity data, and input them into the classification model to obtain the heavy overload classification prediction result at the time point t+N.
2. The method for predicting distribution transformer overload based on daily granularity data according to claim 1 is characterized in that: The step of constructing a classification model, a first regression model and a second regression model according to the daily granularity data specifically includes: Dividing the daily granularity data into an initial training set and an initial test set according to a time series; A first preset number of daily granularity data in the initial training set is extracted by random sampling as a classification model training set, and a second preset number of daily granularity data in the initial training set is extracted as a classification model test set; Extracting a classification feature set from the classification model training set; Building a classification model through the classification feature set, and evaluating the classification model through the classification model test set; Dividing the daily granularity data in the initial training set into a regression model training set and a regression model test set according to the time series; The regression model characteristics are determined according to the power load data in the regression model training set, wherein the power load data includes a daily maximum load and a daily average load, a first regression model and a second regression model are constructed through the regression model characteristics, and the first regression model and the second regression model are evaluated through the regression model test set.
3. The method for predicting distribution transformer overload based on daily granularity data according to claim 2 is characterized in that: The extracting of the classification feature set in the classification model training set specifically includes: Determine the classification model features according to the distribution transformer geographical location information, distribution transformer attribute information, time and load data in the classification model training set; Setting a categorical variable, wherein the value set of the categorical variable includes a plurality of different categorical values in lexicographic order; Define a mapping function, and use the mapping function to map each classification component to a unique natural number as the first classification feature of the non-numeric feature; The numerical feature in the classification model feature is used as the second classification feature, and a classification feature set is constructed according to the first classification feature and the second classification feature.
4. The method for predicting distribution transformer overload based on daily granularity data according to claim 3 is characterized in that: The step of constructing a classification model by using the classification feature set and evaluating the classification model by using the classification model test set specifically includes: Define the classification model objective function based on the logistic regression loss function; The latest tree is generated by the classification model objective function, the label corresponding to each classification feature in the classification feature set is used as the prediction target, the classification features in the classification feature set are used as input samples, and the classification model is constructed by adding the latest tree, and the labels include non-overloaded, overloaded and overloaded.
5. The method for predicting distribution transformer overload based on daily granularity data according to claim 2 is characterized in that: The step of determining the regression model features according to the power load data in the regression model training set, wherein the power load data includes a daily maximum load and a daily average load, constructing a first regression model and a second regression model through the regression model features, and evaluating the first regression model and the second regression model through the regression model test set specifically includes: Determine the regression model features according to the power load data in the first preset time period of the same period last year and the power load data in the second preset time period this year in the regression model training set, wherein the power load data includes a daily maximum load and a daily average load; Define the regression model objective function based on the squared error loss function; Generate the latest tree through the regression model objective function, take the maximum load at the historical time point as the prediction target, take the daily maximum load in the regression model feature as the input sample, and build the first regression model by adding the latest tree; The latest tree is generated by the regression model objective function, the average load at the historical time point is used as the prediction target, the daily average load in the regression model feature is used as the input sample, and the second regression model is constructed by adding the latest tree.
6. The method for predicting distribution transformer overload based on daily granularity data according to claim 2 is characterized in that: The daily average load of the same period last year and the daily average load of this year are matched according to the predicted value of the maximum load of the distribution transformer, and the matched daily average load of the same period last year and the daily average load of this year are input into the second regression model to obtain the predicted value of the average load at the time point t+N, which specifically includes: According to the comparison between the predicted value of the maximum load of the distribution transformer and the load matching threshold, the distribution transformers are screened to obtain a distribution transformer list; Matching the daily average load of the same period last year and the daily average load of this year in the daily granularity data with the distribution transformer list to obtain the matched daily average load of the same period last year and the daily average load of this year; The matched daily average load of the same period last year and the daily average load of this year are input into the second regression model to obtain the predicted value of the average load at the time point t+N.
7. A distribution transformer overload prediction method based on daily granularity data according to claim 6, characterized in that: The step of screening the distribution transformers to obtain a distribution transformer list according to the comparison between the predicted value of the maximum load of the distribution transformer and the load matching threshold specifically includes: Determining whether the predicted value of the maximum load of the distribution transformer is greater than or equal to a load matching threshold; If the predicted value of the maximum load of the distribution transformer is greater than or equal to the load matching threshold, the maximum load of the distribution transformer is retained and a distribution transformer list is obtained; If the predicted value of the maximum load of the distribution transformer is less than the load matching threshold, the maximum load of the distribution transformer is not retained.
8. The method for predicting distribution transformer overload based on daily granularity data according to claim 2 is characterized in that: After replacing the daily maximum load value and the daily load average value of the corresponding date in the daily granularity data with the predicted value of the maximum load and the predicted value of the average load at the time point t+N, and inputting them into the classification model to obtain the heavy overload classification prediction result at the time point t+N, the method further includes: Compare the heavy overload classification prediction result at the time point t+N with the daily granularity data of the corresponding date in the initial test set to determine the hit rate and false alarm rate; Based on the hit rate and false alarm rate, hyperparameter tuning is performed on the classification voting weight and load matching threshold of the classification model.
9. A distribution transformer overload prediction system based on daily granularity data, characterized in that: The system comprises: The collection module is used to collect daily granular data of distribution transformers in the past year; A model building module, used to build a classification model, a first regression model and a second regression model according to the daily granularity data; A maximum load prediction module, used for inputting the daily maximum load of the same period last year and the daily maximum load of this year in the daily granularity data into the first regression model to obtain a predicted value of the maximum load of the distribution transformer at a time point t+N, where t is a time point and N is a number of days; an average load prediction module, used to match the predicted value of the maximum load of the distribution transformer with the daily average load of the same period last year and the daily average load of this year, input the matched daily average load of the same period last year and the daily average load of this year into the second regression model, and obtain the predicted value of the average load at time point t+N; The heavy overload classification prediction module is used to replace the daily maximum load value and the daily load average of the corresponding date in the daily granularity data with the predicted value of the maximum load and the predicted value of the average load at the t+N time point, input the classification model, and obtain the heavy overload classification prediction result at the t+N time point.
10. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 8.
11. A computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 8.