A time series power parameter prediction method based on new energy segmentation
By adopting a segmented time series prediction method in the prediction of power parameters of new energy generation, combined with gray correlation analysis, DBSCAN algorithm, genetic algorithm and Transformer model, the prediction of intermittent and uncertainty of power generation in new energy generation is solved, and high-accurate power parameter prediction is achieved, supporting the stable operation and cost optimization of the power system.
Patent Information
- Application Number
- CN202411405193.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-10
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-10-10
AI Technical Summary
When predicting the power parameters of new energy generation, it is difficult for the existing technology to effectively deal with the intermittent and uncertainty of new energy generation, as well as complex factors such as load changes and market supply and demand relationships in the power market, resulting in insufficient prediction accuracy.
The time series power parameter prediction method based on the new energy segmentation is adopted, and the historical sample number set is screened through the gray correlation analysis method, the DBSCAN algorithm and the European distance screening cluster center number set, and the genetic algorithm processes the sample number set, and finally the Transformer time series model is constructed for prediction, and the intermittentity and uncertainty of new energy power generation are considered in combination with the correction period.
Accurate prediction of the power parameters of new energy power generation is achieved, and predicted values close to the real power parameters can be obtained under the influence of intermittent and uncertainty of new energy power generation, supporting the load management and scheduling of the power system, and reducing power production losses and costs.
Smart Images

Figure CN119359440B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power market transactions, and in particular to a time series power parameter prediction method based on a new energy output segmentation method. Background Art
[0002] With the gradual reform and development of the power market, the rules of the power market have become more and more complex, and the factors involved have become more and more extensive. In recent years, with the rapid growth of installed capacity of new energy, the participation of new energy in the power market has increased significantly. Especially in some provinces where new energy accounts for a relatively high proportion, the growth of new energy power generation has had a profound impact on the operation mode of the power market and the electricity price mechanism.
[0003] Renewable energy generation, especially wind and solar energy, has attracted widespread attention due to its clean and renewable characteristics. However, the power generation of these energy sources is limited by natural conditions and has strong intermittency and uncertainty. With the increase in the installed capacity of renewable energy, the impact of this uncertainty on the power market has become increasingly obvious.
[0004] On the one hand, the cost of renewable energy generation continues to decline, resulting in a reduction in the marginal clearing parameters of the electricity market, which in turn affects the economic benefits of traditional energy sources (such as coal-fired power and natural gas power generation). On the other hand, due to the volatility of renewable energy generation, the power system requires more flexible adjustment resources to ensure the balance between supply and demand, which not only increases the complexity of the system, but also poses new challenges to the pricing mechanism of the electricity market.
[0005] In some areas, the substantial growth in renewable energy generation may even lead to abnormal fluctuations in spot market power parameters, such as the occurrence of negative electricity prices. In addition, with the promotion of the peak-valley electricity price mechanism, many provinces have implemented time-of-use electricity price policies to encourage consumers to use electricity during the peak period of renewable energy generation, which further affects the trend of power parameters.
[0006] Therefore, for the power market in areas with large amounts of renewable energy generation, accurate prediction of power parameters becomes particularly important. This not only helps power market participants to better formulate trading strategies, but also helps power system operators optimize dispatch plans and ensure the safe and stable operation of the power system.
[0007] In the prior art, there are mainly linear regression or nonlinear regression models based on historical data and related economic indicators, which are mainly suitable for short-term power parameter prediction, but have limitations in dealing with nonlinear relationships and complex patterns. And through machine learning methods such as support vector machine (SVM), random forest (RF) and gradient boosting tree (GBT), these methods can handle complex nonlinear relationships, but have limited processing capabilities for long time series data. Although the above methods have improved the accuracy of power parameter prediction for renewable energy power generation to a certain extent, there are still some limitations. For example, renewable energy power generation is intermittent and uncertain, as well as key factors such as load changes and market supply and demand relationships in the power market, which all affect the accuracy of power parameter prediction. No relevant analysis has been seen in the prior art, which affects the accuracy of power parameter prediction. Summary of the invention
[0008] The present invention aims at solving the problems existing in the prior art and provides a
[0009] To achieve the above purpose, the technical solution adopted by the present invention is as follows: a time series power parameter prediction method based on new energy segmentation, comprising the following steps:
[0010] Obtain a target period and a target period feature number set; obtain a sampling period according to the target period; obtain a historical sample number set, the historical sample number set including a feature number set and a power parameter value set;
[0011] Dividing the historical sample number set according to the sampling period to obtain a plurality of sub-historical sample number sets, wherein the sub-historical sample number sets are historical sample number sets within the sampling period;
[0012] Based on the target period feature number set, the plurality of sub-historical sample number sets are screened according to a grey correlation analysis method to obtain a first sample number set;
[0013] Using the DBSCAN algorithm to process the plurality of sub-historical sample number sets to obtain a plurality of cluster centers and a plurality of cluster number sets, wherein the cluster centers correspond to the cluster number sets one by one; obtaining the cluster center number sets according to the cluster centers;
[0014] Based on the target time period feature number set, a plurality of cluster center number sets are selected according to the Euclidean distance formula to obtain the optimal cluster center number set;
[0015] Obtaining an optimal clustering number set according to the optimal clustering center number set and the clustering number set;
[0016] Obtaining a preliminary sample set by finding the intersection of the first sample set and the optimal clustering set;
[0017] Processing the preliminary sample set according to the genetic algorithm to obtain a training sample set;
[0018] Based on the Transformer time series, a power parameter prediction model is constructed and trained according to the training sample set;
[0019] Obtaining an initial prediction number set of power parameters within the target time period according to the target time period feature number set and the power parameter prediction model;
[0020] Get the preset cycle and new energy output value set;
[0021] Filter the historical sample number set according to the preset period to obtain a preset historical sample number set, wherein the preset historical sample number set is a historical sample number set within the preset period;
[0022] Using the 3sigma criterion, obtaining a correction period and a correction algorithm according to the new energy output value set and the preset historical sample set;
[0023] The initial prediction number set is modified based on the modification period and the modification algorithm to obtain a final prediction number set.
[0024] Furthermore, the step of obtaining a historical sample set includes:
[0025] Obtaining a target time, and obtaining a sampling time according to the target time; obtaining first historical data;
[0026] Filter the first historical data based on the sampling time to obtain an original historical data set;
[0027] Cleaning the original historical data set for missing values and removing outliers to obtain a first historical data set;
[0028] The first historical number set is subjected to standardization processing to obtain the historical sample number set.
[0029] Furthermore, the sub-historical sample number set includes a first feature number set, which is a collection of feature data at each sampling moment in a sampling period.
[0030] Furthermore, the step of obtaining the first sample number set is:
[0031] Get the resolution factor;
[0032] Using a grey correlation analysis method, obtaining a correlation degree according to the resolution coefficient, the first feature number set and the target time period feature number set;
[0033] Sorting the plurality of sub-historical sample sets according to the association degree;
[0034] Obtaining a first number, where the first number is the number of sub-historical sample number sets in the first sample number set;
[0035] The sub-historical sample number set with a high correlation value is selected according to the first number to obtain the first sample number set.
[0036] Furthermore, the step of obtaining the optimal clustering number set is:
[0037] Using an interval statistics method to process the plurality of sub-historical sample sets to obtain a clustering K value;
[0038] Clustering the plurality of sub-historical sample sets using the DBSCAN algorithm according to the clustering K value to obtain a plurality of cluster centers;
[0039] Obtaining the number of clusters, where the number of clusters is the number of sub-historical sample number sets in the cluster number set;
[0040] Obtaining the plurality of cluster number sets according to the plurality of cluster centers and the number of clusters;
[0041] Obtaining a cluster center number set according to the cluster center, wherein the cluster center number set includes a second feature number set, and the second feature number set is a set of feature data at each sampling time in the sampling period where the cluster center is located;
[0042] Using the Euclidean distance formula, obtaining the Euclidean distance of each cluster center according to the second feature number set and the target time period feature number set;
[0043] Sorting the plurality of cluster center number sets according to the Euclidean distance, and obtaining the optimal cluster center number set according to the minimum Euclidean distance value;
[0044] An optimal clustering number set is obtained according to the optimal clustering center number set and the cluster number set.
[0045] Furthermore, the preliminary sample set includes a third feature set; and the step of obtaining the training sample set is:
[0046] Performing feature expansion based on the third feature number set to obtain a fourth feature number set;
[0047] Using a genetic algorithm to screen the fourth feature number set to obtain a key feature number set;
[0048] The training sample set is obtained according to the key feature set and the preliminary selected sample set.
[0049] Furthermore, the steps of obtaining the power parameter prediction model are:
[0050] Based on the Transformer time series, an initial prediction model is constructed according to the training sample set;
[0051] Dividing the training sample set to obtain a training set and a verification set, wherein the verification set includes a verification power parameter value set;
[0052] Training the initial prediction model according to the training set to obtain a first prediction value set;
[0053] Based on the verification set, constructing a mean square error function according to the first prediction value set and the verification power parameter value set;
[0054] The initial prediction model with the smallest mean square error function value is selected to obtain the power parameter prediction model.
[0055] Furthermore, the steps of training the initial prediction model are:
[0056] Step S1: setting the input data of the initial prediction model;
[0057] Step S2: Set the maximum number of iterations and initialize the current number of iterations;
[0058] Step S3: running the initial prediction model to obtain a first prediction value;
[0059] Step S4: Let t = t + 1, and determine whether t ≤ t max Is it true? In the formula, t is the current iteration number, t max is the maximum number of iterations:
[0060] If established, repeat step S3;
[0061] If not, the first prediction value set is obtained according to the first prediction value.
[0062] Further, the preset historical sample number set includes a first power parameter value set, which is a data set of power parameters within the preset period;
[0063] The steps of obtaining the modified time period are:
[0064] Obtaining a first mean value set according to the first power parameter value set, wherein the first mean value set is a set of power parameter mean values at the same sampling time within the preset period;
[0065] Obtaining a second mean value set according to the new energy output value set, wherein the second mean value set is a set of new energy output mean values at the same sampling time within the preset period;
[0066] Using the least squares method, a typical value function of the power parameter is obtained according to the first mean value set, and a typical value function of the new energy output is obtained according to the second mean value set;
[0067] Obtain a typical curve of electric power parameters according to the typical value function of electric power parameters; obtain a typical curve of new energy output according to the typical value function of new energy output;
[0068] Using a box plot algorithm, a first new energy output value is obtained on the new energy output typical curve according to the lower quartile of the power parameter typical curve, and a second new energy output value is obtained on the new energy output typical curve according to the upper quartile of the power parameter typical curve;
[0069] The correction period is obtained by dividing the sampling period according to the first new energy output value and the second new energy output value, and the correction period includes a low system new energy output section, a flat system new energy output section and a high system new energy output section.
[0070] Further, the step of correcting the initial prediction number set based on the correction period to obtain the final prediction number set is:
[0071] In the flat system new energy output section and the high system new energy output section, the initial prediction number set is corrected by using a deviation correction method to obtain the final prediction number set;
[0072] In the low system new energy output stage, the initial prediction number set is corrected by setting the minimum value method to obtain the final prediction number set.
[0073] Compared with the prior art, the present invention has the following beneficial effects:
[0074] The present invention screens historical sample sets by grey correlation analysis method, screens historical sample sets by DBSCAN algorithm plus Euclidean distance, and obtains a preliminary sample set by finding the intersection of the two, and obtains a training sample set by genetic algorithm. The set of feature data in the training sample set more realistically fits key factors such as load changes and market supply and demand relationships in the power market during the target period, thereby achieving accurate modeling for the target period. The power parameter prediction model obtained by training the training sample set is combined with the correction period, fully considering the intermittency and uncertainty of new energy power generation, and can obtain final predicted power parameter values close to real power parameters under the influence of fluctuation factors such as intermittency and uncertainty of new energy power generation. At the same time, the final predicted power parameter values can be used to reversely guide market supply and demand, provide reliable support for load management and scheduling of power systems, and reduce power production losses and costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 It is a flow chart of the time series power parameter prediction method based on new energy segmentation of the present invention;
[0076] Figure 2 Schematic diagram of part of the first historical data in an embodiment of the present invention Figure I ;
[0077] Figure 3 Schematic diagram of part of the first historical data in an embodiment of the present invention Figure II ;
[0078] Figure 4 A schematic diagram of feature expansion using feature engineering in an embodiment of the present invention;
[0079] Figure 5 This is a schematic diagram of some data in the key feature number set in an embodiment of the present invention;
[0080] Figure 6 Schematic diagram of comparison between the final predicted power parameter value and the actual power parameter value in an embodiment of the present invention. DETAILED DESCRIPTION
[0081] It is worth noting that the methods used in the present invention are all conventional methods unless otherwise specified; the raw materials and devices used are all conventional commercially available products, and their sources are not specifically limited unless otherwise specified.
[0082] In the existing technology, there are mainly linear regression or nonlinear regression models based on historical data and related economic indicators, which are mainly suitable for short-term power parameter prediction, but have limitations in dealing with nonlinear relationships and complex patterns. And through machine learning methods such as support vector machine (SVM), random forest (RF) and gradient boosting tree (GBT), these methods can handle complex nonlinear relationships, but have limited processing capabilities for long time series data. Although the above methods have improved the accuracy of power parameter prediction for renewable energy power generation to a certain extent, there are still some limitations. For example, they may not fully consider the intermittency and uncertainty of renewable energy power generation, as well as the impact of other key factors in the power market (such as load changes, market supply and demand, etc.).
[0083] A time series power parameter prediction method based on new energy segmentation includes the following steps:
[0084] Obtain the target period and the target period feature set; the target period is the period where the power parameter to be predicted is located, usually the target period is one day, that is, the target day, and the sampling period is obtained according to the target period, and the sampling period is the same as the target period;
[0085] A historical sample number set is obtained, the historical sample number set includes a feature number set and a power parameter value set, and the feature number set corresponds to the power parameter value set; specifically, the steps of obtaining the historical sample number set include:
[0086] Obtain the target time, usually collect historical data every 15 minutes, that is, the target time duration is 15 minutes, and obtain the sampling time according to the target time, and the sampling time duration is consistent and corresponding to the target time duration;
[0087] Obtain the first historical data; the first historical data is the data disclosed by the market, including characteristic data and corresponding power parameters. In the first historical data, the types of characteristic data include weather data and energy and load data. The energy and load data include the load data of the entire network, interconnection lines, wind power output, photovoltaic output, hydropower output, new energy output, and load rate, etc.; usually, the market disclosed data Y days (Y≥180) from the target date forward is selected as the first historical data;
[0088] The first historical data is filtered based on the sampling time to obtain an original historical data set; the characteristic data and the corresponding power parameters at a specific sampling time are filtered, see Table 1:
[0089] Table 1 Feature data of original historical data set
[0090]
[0091] Cleaning the original historical data set for missing values and removing outliers to obtain a first historical data set;
[0092] Taking the target period as one day and the target time as 15 minutes as an example, the number of sampling points per day is 96;
[0093] First, the missing values are filled using interpolation. If there are more than 95 consecutive target periods without data, no data will be added. The data of that day, i.e. the target period, will be deleted directly. The formula of the interpolation method is:
[0094]
[0095] In the formula, X d,t is the characteristic data value at sampling time t and sampling period d, that is, it represents the missing value in the formula, The sampling period is d and the sampling time is t k The characteristic data value of The sampling period is d and the sampling time is t k+1 The characteristic data value of t is the sampling time corresponding to the missing value, t k is the sampling time before sampling time t, t k+1 is the sampling time after the sampling time t;
[0096] Secondly, the upper and lower limits of the characteristic data are determined through the box plot. The characteristic data values outside the range of the lower and upper limits are outliers, which are removed.
[0097] The first historical data set is processed by standardization to obtain the historical sample data set; since the collected feature data are in different data dimension standards, in order to avoid the deviation caused by the difference between the two levels of feature data affecting the generalization ability of the model, the data is standardized to a fixed range, such as [0, 1] or [-1, 1], so as to speed up the model training speed and improve the prediction accuracy, that is, the collected feature data needs to be standardized; preferably, normalization (min-max normalization method) is used to standardize each type of feature data, and the normalization processing formula is:
[0098]
[0099] Among them, X' d,t is the normalized characteristic data value at sampling time t and sampling period d, X d,t is the characteristic data value with sampling period d and sampling time t, is the minimum value of the same type of feature data in the first historical data set, is the maximum value of the same type of feature data in the first historical data set, d is the sampling period, and D is the total number of the same type of feature data in the first historical data set.
[0100] According to the sampling period, the historical sample number set is divided into several sub-historical sample number sets, wherein the sub-historical sample number set is the historical sample number set within the sampling period, and the sub-historical sample number set includes a first feature number set, which is a collection of feature data at each sampling moment within the sampling period; usually, the sampling period is one day, and the sub-historical sample number set is the historical sample number set within one day, see Figure 2 , 3 ;
[0101] Based on the target period feature number set, the plurality of sub-historical sample number sets are screened according to the grey correlation analysis method to obtain a first sample number set; the steps of obtaining the first sample number set are:
[0102] Get the resolution factor;
[0103] The grey correlation analysis method is used to obtain the correlation degree according to the resolution coefficient, the first characteristic number set and the characteristic number set of the target period;
[0104] By analyzing the first historical data set, it is found that when the change trends of power parameters are similar, the influencing factors of power parameters, i.e., characteristic data, are also close. Therefore, the first historical data set is first classified by day (sampling period length), and then the grey correlation analysis method is used to screen out the sub-historical sample data set closest to the target period; the formula used by the grey correlation analysis method is:
[0105]
[0106] In the formula, is the correlation, X' d,t is the normalized characteristic data value at sampling time t and sampling period d, is the normalized characteristic data value of the sampling time t in the target period, t is the sampling time, T is the total number of sampling times in the sampling period, α is the resolution coefficient, and α is generally taken as 0.5;
[0107] Sort several sub-historical sample sets according to the relevance;
[0108] Obtain a first number N, where the first number N is the number of sub-historical sample number sets in the first sample number set;
[0109] The calculated correlations are sorted by size. The larger the correlation, the higher the similarity between the sequence and the reference sequence, that is, the higher the similarity between the characteristic data of the sampling period and the characteristic data of the target period. Finally, the first N sub-historical sample sets with high correlation are selected as the first sample set.
[0110] According to the first number N, a sub-historical sample number set with a high correlation value is selected to obtain a first sample number set.
[0111] The DBSCAN algorithm is used to process several sub-historical sample sets to obtain several cluster centers and several cluster number sets, and the cluster centers correspond to the cluster number sets one by one; the cluster center number set is obtained according to the cluster centers; the steps to obtain the optimal cluster number set are:
[0112] Use the interval statistics method to process several sub-historical sample sets to obtain the clustering K value;
[0113] According to the clustering K value, the DBSCAN algorithm is used to cluster several sub-historical sample sets to obtain several cluster centers;
[0114] Get the number of clusters N * , the number of clusters is the number of historical sample sets in the cluster set;
[0115] According to several cluster centers and the number of clusters N * Obtain several clustering data sets;
[0116] Obtaining a cluster center number set according to the cluster center, the cluster center number set includes a second feature number set, and the second feature number set is a set of feature data at each sampling time in the sampling period where the cluster center is located;
[0117] Based on the target time period feature set, several cluster center sets are selected according to the Euclidean distance formula to obtain the optimal cluster center set; that is, the Euclidean distance formula is:
[0118]
[0119] Among them, dist is the Euclidean distance, X' d,t is the normalized characteristic data value at sampling time t and sampling period d, X' di,t is the normalized characteristic data value at sampling time t in the target period, t is the sampling time, and T is the total number of sampling times in the sampling period;
[0120] Sort several cluster center sets according to the Euclidean distance, and obtain the optimal cluster center set according to the minimum Euclidean distance value;
[0121] Obtaining the optimal cluster number set according to the optimal cluster center number set and cluster number set;
[0122] Calculate the Euclidean distance between the target period feature set and each cluster center set, and sort the results. The cluster center set with the smallest Euclidean distance is the optimal cluster center set, that is, the feature data in the optimal cluster center set is closest to the feature data of the target period.
[0123] Obtain a preliminary sample set by finding the intersection of the first sample set and the optimal clustering set;
[0124] The training sample set is obtained by processing the preliminary sample set according to the genetic algorithm; specifically, the preliminary sample set includes the third feature set; the steps of obtaining the training sample set are:
[0125] Perform feature expansion based on the third feature number set to obtain a fourth feature number set;
[0126] By using feature engineering and data analysis methods, feature expansion is performed on the basis of feature data in the primary sample set; using the length of point X as the time window, the existing feature data in the third feature set is lagged, and each feature data is shifted backward by the time window length as the new feature data based on the time window, thereby obtaining the fourth feature set. The details are as follows:
[0127] Feature feature1 = [1,2,3,4,5,6,...], assuming that the time window length X is 3, then feature1 is shifted backward by X points in order, that is, data point 1 is moved to the position of data 4, data point 2 is moved to the position of data 5, data point 3 is moved to the position of data 6, and so on, and feature2 = [0,0,0,1,2,3,...] is obtained, where the first X data points of feature2 are filled with 0. The specific process is demonstrated as follows Figure 4 shown.
[0128] The fourth feature set is screened by a genetic algorithm to obtain a key feature set. The genetic algorithm is a prior art and will not be described in detail here. Some key feature data in the training sample set are as follows: Figure 5 As shown, Figure 5 The characteristic data of renewable energy output, wind power output, photovoltaic output, hydropower output, system load, interconnection line, etc. are listed respectively. Each characteristic data lists the characteristic data of multiple sampling days, i.e., multiple sampling periods. The modulus coordinate "time point" represents the sampling moment, and the ordinate "power" represents the value of each characteristic data.
[0129] A training sample set is obtained based on a key feature set and a preliminary sample set.
[0130] Based on the Transformer time series, a power parameter prediction model is constructed and trained according to the training sample set; specifically, the steps to obtain the power parameter prediction model are:
[0131] Based on the Transformer time series, an initial prediction model is constructed according to the training sample set;
[0132] The training sample set is divided into a training set and a validation set, wherein the validation set includes a validation power parameter value set; the training sample set is usually divided according to a certain ratio. In this embodiment, the training sample set is divided into a training set and a validation set according to a ratio of 8:2;
[0133] Training the initial prediction model according to the training set to obtain a first prediction value set;
[0134] Based on the verification set, a mean square error function is constructed according to the first prediction value set and the verification power parameter value set; the mean square error function is:
[0135]
[0136] Among them, MSE is the mean square error, T1 is the number of samples in the validation set, and t1 is the sample number in the validation set. To verify the power parameter values, the actual published power parameter values are collected. is the power parameter value obtained by the initial prediction model;
[0137] Select the initial prediction model with the smallest mean square error function value to obtain the power parameter prediction model;
[0138] According to the target period feature number set and the power parameter prediction model, an initial prediction number set of power parameters in the target period is obtained; the initial prediction number set is a set of initial prediction power parameters;
[0139] Get the preset cycle and new energy output value set;
[0140] Specifically, the second historical data is obtained, which is also the data disclosed by the market, including characteristic data and corresponding power parameters. Unlike the first historical data, the characteristic data in the second historical data only includes data related to the output of new energy. The second historical data is processed according to a preset period to obtain a new energy output value set. In this embodiment, the preset period is one week, and the new energy output values corresponding to 96 sampling moments every day (sampling period) in the past week are taken as the new energy output value set Xnew. ij (i=1,2,...,7,j=1,2,...,96), i is the serial number of the sampling period in the preset cycle, j is the serial number of the sampling time in the sampling period;
[0141] According to the preset period, the historical sample number set is screened to obtain the preset historical sample number set, which is the historical sample number set within the preset period; the preset historical sample number set includes a first power parameter value set, which is a data set of power parameters within the preset period; the first power parameter value set P ij (i=1,2,...,7,j=1,2,...,96);
[0142] The 3sigma criterion is used to obtain the correction period according to the new energy output value set and the preset historical sample set; the steps for obtaining the correction period are:
[0143] A first mean value set is obtained according to the first power parameter value set, where the first mean value set is a set of power parameter mean values at the same sampling time within a preset period; the first mean value set for:
[0144]
[0145] The second mean value set is obtained according to the new energy output value set, and the second mean value set is the set of the new energy output mean values at the same sampling time within the preset period; the second mean value set for:
[0146]
[0147] The least squares method is used to obtain the typical value function minf of the power parameter according to the first mean value set. P (x), and obtain the typical value function minf of new energy output according to the second mean value set Xnew (x);
[0148]
[0149]
[0150] Obtain a typical curve of power parameters according to a typical value function of power parameters; obtain a typical curve of new energy output according to a typical value function of new energy output;
[0151] Using the box plot algorithm, the first new energy output value is obtained on the new energy output typical curve according to the lower quartile of the power parameter typical curve, and the second new energy output value is obtained on the new energy output typical curve according to the upper quartile of the power parameter typical curve;
[0152] The sampling period is divided according to the first new energy output value and the second new energy output value to obtain a correction period, and the correction period includes a low system new energy output section, a flat system new energy output section and a high system new energy output section.
[0153] The initial prediction number set is corrected based on the correction period to obtain the final prediction number set, which is a set of final prediction power parameters; specifically:
[0154] The steps for correcting the initial prediction number set based on the correction period to obtain the final prediction number set are:
[0155] In the flat system new energy output section and the high system new energy output section, the deviation correction method is used to correct the initial prediction number set to obtain the final prediction number set; the formula of the deviation correction method is:
[0156]
[0157] Where M is the length of the window period, m is the sampling time sequence number from the target time forward, is the mean value of the power parameter value in the M window period forward from the target time, P m(日前) is the power parameter value at the mth sampling moment before the target moment, P m(修正) To predict the final power parameter value, P m(预测) is the initial predicted power parameter value, is the second mean, Xnew h Xnew is the new energy output value of the high system new energy output stage. p It is the new energy output value of the new energy output section of the flat system;
[0158] In the low system new energy output stage Xnew l , the initial prediction number set is corrected by setting it to the minimum value method to obtain the final prediction number set; the formula of the minimum value method is:
[0159]
[0160] Where P m(修正) To predict the final power parameter value, P m(预测) is the initial predicted power parameter value.
[0161] The historical sample set is screened by the grey correlation analysis method, and the DBSCAN algorithm plus the Euclidean distance is used to screen the historical sample set. The initial sample set is obtained by finding the intersection of the two, and the training sample set is obtained by the genetic algorithm. The set of feature data in the training sample set has a high correlation with the feature set of the target period, that is, the set of feature data in the training sample set more realistically fits the key factors such as load changes and market supply and demand relationships in the power market during the target period. The power parameter prediction model obtained by training the training sample set is combined with the correction period, fully considering the intermittent and uncertainties of new energy power generation, and the target period feature set is input into the power parameter prediction model and corrected through the correction period, so that the final predicted power parameter values close to the real power parameters can be obtained. The target time-final predicted power parameter values obtained after correction are as follows Figure 6 As shown, it can be seen that the values and trends of the final predicted power parameters are close to the actual power parameter values.
[0162] Furthermore, the steps for training the initial prediction model are:
[0163] Step S1: setting the input data of the initial prediction model;
[0164] Step S2: Set the maximum number of iterations and initialize the current number of iterations;
[0165] Step S3: running the initial prediction model to obtain a first prediction value;
[0166] Step S4: Let t = t + 1, and determine whether t ≤ t max Is it true? In the formula, t is the current iteration number, t max is the maximum number of iterations:
[0167] If established, repeat step S3;
[0168] If not, a first prediction value set is obtained according to the first prediction value.
[0169] The embodiment of the present invention also provides a device for implementing a time series power parameter prediction method based on new energy segmentation, comprising a data input unit, a first data processing unit, a modeling unit, a second data processing unit and a calculation unit;
[0170] Data input unit:
[0171] Used to obtain the target period and the target period feature number set;
[0172] Used to obtain the sampling period according to the target period;
[0173] Used to obtain historical sample data sets;
[0174] Used to obtain preset cycle and new energy output value data set;
[0175] First data processing unit:
[0176] Used to divide the historical sample number set according to the sampling period to obtain several sub-historical sample number sets;
[0177] Based on the target period feature set, a plurality of sub-historical sample sets are screened according to the grey correlation analysis method to obtain a first sample set;
[0178] It is used to process several sub-historical sample sets using DBSCAN algorithm to obtain several cluster centers and several cluster number sets, and the cluster centers correspond to the cluster number sets one by one; and the cluster center number sets are obtained according to the cluster centers;
[0179] It is used to select several cluster center number sets based on the target time period feature number set according to the Euclidean distance formula to obtain the optimal cluster center number set;
[0180] Used to obtain the optimal cluster number set based on the optimal cluster center number set and cluster number set;
[0181] Used to obtain a preliminary sample set based on the intersection of the first sample set and the optimal clustering set;
[0182] Modeling unit:
[0183] Used to obtain a preliminary sample set based on the intersection of the first sample set and the optimal clustering set;
[0184] Used to process the preliminary sample set according to the genetic algorithm to obtain the training sample set;
[0185] Used for Transformer-based time series, to build and train a power parameter prediction model based on the training sample set;
[0186] Second data processing unit:
[0187] Used to filter the historical sample number set according to the preset period to obtain the preset historical sample number set, the preset historical sample number set is the historical sample number set within the preset period;
[0188] Used to obtain the correction period according to the new energy output value set and the preset historical sample set using the 3sigma criterion;
[0189] Computational Unit:
[0190] Used to obtain an initial prediction number set of power parameters within a target period according to a target period feature number set and a power parameter prediction model;
[0191] Used to correct the initial prediction number set based on the correction period to obtain the final prediction number set.
[0192] Finally, it should be noted that the above content is only used to illustrate the technical solution of the present invention, rather than to limit the scope of protection of the present invention. Simple modifications or equivalent substitutions of the technical solution of the present invention by ordinary technicians in this field do not deviate from the essence and scope of the technical solution of the present invention.
Claims
1. A time series power parameter prediction method based on new energy segmentation, characterized in that: The following steps are involved: Obtaining a target period and a target period feature number set; obtaining a sampling period according to the target period; Acquire a historical sample number set, wherein the historical sample number set includes a feature number set and a power parameter value set; Dividing the historical sample number set according to the sampling period to obtain a plurality of sub-historical sample number sets, wherein the sub-historical sample number sets are historical sample number sets within the sampling period; Based on the target period feature number set, the plurality of sub-historical sample number sets are screened according to a grey correlation analysis method to obtain a first sample number set; Using the DBSCAN algorithm to process the plurality of sub-historical sample number sets to obtain a plurality of cluster centers and a plurality of cluster number sets, wherein the cluster centers correspond to the cluster number sets one by one; obtaining the cluster center number sets according to the cluster centers; Based on the target time period feature number set, a plurality of cluster center number sets are selected according to the Euclidean distance formula to obtain the optimal cluster center number set; Obtaining an optimal clustering number set according to the optimal clustering center number set and the clustering number set; Obtaining a preliminary sample set by finding the intersection of the first sample set and the optimal clustering set; Processing the preliminary sample set according to the genetic algorithm to obtain a training sample set; Based on the Transformer time series, a power parameter prediction model is constructed and trained according to the training sample set; Obtaining an initial prediction number set of power parameters within the target time period according to the target time period feature number set and the power parameter prediction model; Get the preset cycle and new energy output value set; Filter the historical sample number set according to the preset period to obtain a preset historical sample number set, wherein the preset historical sample number set is a historical sample number set within the preset period; Using the 3sigma criterion, obtaining a correction period according to the new energy output value set and the preset historical sample set; Modifying the initial prediction number set based on the modification period to obtain a final prediction number set; The preset historical sample number set includes a first power parameter value set, which is a data set of power parameters within the preset period; The steps of obtaining the modified time period are: Obtaining a first mean value set according to the first power parameter value set, wherein the first mean value set is a set of power parameter mean values at the same sampling time within the preset period; Obtaining a second mean value set according to the new energy output value set, wherein the second mean value set is a set of new energy output mean values at the same sampling time within the preset period; Using the least squares method, a typical value function of the power parameter is obtained according to the first mean value set, and a typical value function of the new energy output is obtained according to the second mean value set; Obtain a typical curve of electric power parameters according to the typical value function of electric power parameters; obtain a typical curve of new energy output according to the typical value function of new energy output; Using a box plot algorithm, a first new energy output value is obtained on the new energy output typical curve according to the lower quartile of the power parameter typical curve, and a second new energy output value is obtained on the new energy output typical curve according to the upper quartile of the power parameter typical curve; The correction period is obtained by dividing the sampling period according to the first new energy output value and the second new energy output value, and the correction period includes a low system new energy output section, a flat system new energy output section and a high system new energy output section.
2. The time series power parameter prediction method based on new energy segmentation according to claim 1 is characterized in that: The steps to obtain the historical sample data set include: Obtaining a target time, and obtaining a sampling time according to the target time; obtaining first historical data; Filter the first historical data based on the sampling time to obtain an original historical data set; Cleaning the original historical data set for missing values and removing outliers to obtain a first historical data set; The first historical number set is subjected to standardization processing to obtain the historical sample number set.
3. The time series power parameter prediction method based on new energy segmentation according to claim 2 is characterized in that: The sub-historical sample number set includes a first feature number set, which is a collection of feature data at each sampling moment in a sampling period.
4. The time series power parameter prediction method based on new energy segmentation according to claim 3 is characterized in that: The steps of obtaining the first sample number set are: Get the resolution factor; Using a grey correlation analysis method, obtaining a correlation degree according to the resolution coefficient, the first feature number set and the target time period feature number set; Sorting the plurality of sub-historical sample sets according to the association degree; Obtaining a first number, where the first number is the number of sub-historical sample number sets in the first sample number set; The sub-historical sample number set with a high correlation value is selected according to the first number to obtain the first sample number set.
5. The time series power parameter prediction method based on new energy segmentation according to claim 3 is characterized in that: The steps of obtaining the optimal clustering number set are: Using an interval statistics method to process the plurality of sub-historical sample sets to obtain a clustering K value; Clustering the plurality of sub-historical sample sets using the DBSCAN algorithm according to the clustering K value to obtain a plurality of cluster centers; Obtaining the number of clusters, where the number of clusters is the number of sub-historical sample number sets in the cluster number set; Obtaining the plurality of cluster number sets according to the plurality of cluster centers and the number of clusters; Obtaining a cluster center number set according to the cluster center, wherein the cluster center number set includes a second feature number set, and the second feature number set is a set of feature data at each sampling time in the sampling period where the cluster center is located; Using the Euclidean distance formula, obtaining the Euclidean distance of each cluster center according to the second feature number set and the target time period feature number set; Sorting the plurality of cluster center number sets according to the Euclidean distance, and obtaining the optimal cluster center number set according to the minimum Euclidean distance value; An optimal clustering number set is obtained according to the optimal clustering center number set and the cluster number set.
6. The time series power parameter prediction method based on new energy segmentation according to claim 3 is characterized in that: The preliminary sample set includes a third feature set; the steps of obtaining the training sample set are: Performing feature expansion based on the third feature number set to obtain a fourth feature number set; Using a genetic algorithm to screen the fourth feature number set to obtain a key feature number set; The training sample set is obtained according to the key feature set and the preliminary selected sample set.
7. The time series power parameter prediction method based on new energy segmentation according to claim 6 is characterized in that: The steps of obtaining the power parameter prediction model are: Based on the Transformer time series, an initial prediction model is constructed according to the training sample set; Dividing the training sample set to obtain a training set and a verification set, wherein the verification set includes a verification power parameter value set; Training the initial prediction model according to the training set to obtain a first prediction value set; Based on the verification set, constructing a mean square error function according to the first prediction value set and the verification power parameter value set; The initial prediction model with the smallest mean square error function value is selected to obtain the power parameter prediction model.
8. The time series power parameter prediction method based on new energy segmentation according to claim 7 is characterized in that: The steps of training the initial prediction model are: Step S1: setting the input data of the initial prediction model; Step S2: Set the maximum number of iterations and initialize the current number of iterations; Step S3: running the initial prediction model to obtain a first prediction value; Step S4: Let t = t + 1, and determine whether t ≤ t max Is it true? In the formula, t is the current iteration number, t max is the maximum number of iterations: If established, repeat step S3; If not, the first prediction value set is obtained according to the first prediction value.
9. The time series power parameter prediction method based on new energy segmentation according to claim 1 is characterized in that: The steps of correcting the initial prediction number set based on the correction period to obtain the final prediction number set are: In the flat system new energy output section and the high system new energy output section, the initial prediction number set is corrected by using a deviation correction method to obtain the final prediction number set; In the low system new energy output stage, the initial prediction number set is corrected by setting the minimum value method to obtain the final prediction number set.
Citation Information
Patent Citations
Multi-microgrid energy storage configuration optimization method, device and equipment and storage medium
CN115498623A
Short-term photovoltaic power generation power prediction method, system and equipment based on similar weather discovery and medium
CN118213987A