Short-term photovoltaic power intelligent interval prediction method
By constructing a GRU-Informer and SVR double-layer prediction model based on the Blending integrated learning framework, and combining the KDE non-parametric estimation method, the problems of overfitting and poor stability of existing photovoltaic power prediction methods are solved, and more accurate and stable short-term interval prediction of photovoltaic power is achieved.
Patent Information
- Application Number
- CN202510234979.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-17
AI Technical Summary
Existing photovoltaic power prediction methods are prone to overfitting risks, resulting in poor prediction stability.
A short-term photovoltaic power intelligent interval prediction method is adopted to collect weather data, process data, build a GRU-Informer and SVR double-layer prediction model based on the Blending integrated learning framework, and use the KDE non-parametric estimation method to perform point prediction error estimation and interval prediction.
It improves the stability and accuracy of short-term prediction of photovoltaic power, reduces the risk of overfitting, and achieves more accurate photovoltaic power interval prediction.
Smart Images

Figure CN120163441A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of new energy prediction, and specifically relates to a short-term photovoltaic power intelligent interval prediction method. Background Art
[0002] Under the background of the green transformation of the power market, the penetration rate of the photovoltaic industry in the power market is increasing day by day. With the continuous deepening of the reform of the power system, photovoltaic power generation has become increasingly popular and has become an important part of the new energy system. However, photovoltaic power generation has characteristics such as intermittency, volatility, and discontinuity, which increase the uncertainty of the scheduling and operation of the power system. Accurate interval prediction of photovoltaic power generation can reduce the operation risks and costs of the newly built power system, provide accurate forecast forward information for the power system to support grid scheduling and power system peak shaving, and ensure the safe and effective operation of the power system.
[0003] The defect of the existing technology is that the prediction model of the existing photovoltaic power prediction method is prone to overfitting risk, and due to the characteristics of intermittency, volatility, and discontinuity of photovoltaic power generation, the prediction stability is poor. Summary of the Invention
[0004] Aiming at the problems that the prediction model of the existing photovoltaic power prediction method is prone to overfitting risk and the prediction stability is poor, the present invention provides a short-term photovoltaic power intelligent interval prediction method.
[0005] To achieve the above technical purpose, the technical solution adopted by the present invention is as follows:
[0006] A short-term photovoltaic power intelligent interval prediction method includes the steps of:
[0007] S1. Collect weather data and establish an original data set;
[0008] S2. Perform data processing on the original data set, and the data processing includes outlier detection, sorting of weather feature importance, K-Means weather clustering, data mode decomposition, determination of feature input and output, and data set division;
[0009] S3. Construct a GRU-Informer and SVR double-layer prediction model based on the Blending ensemble learning framework;
[0010] S4. Estimate the point prediction error of photovoltaic power through the KDE non-parametric estimation method, and estimate the upper and lower limits of the photovoltaic power prediction interval at a given confidence level;
[0011] S5. Conduct comparative analysis and error analysis on the prediction results.
[0012] Further, the detailed steps of step S2 include:
[0013] S201. Processing of abnormal data, detecting and replacing power abnormal data;
[0014] S202. Standardizing the data set to eliminate the order of magnitude and dimension differences between each index data;
[0015] S203. Introducing recursive feature elimination with cross-validation (RFECV) to improve XGBoost, constructing an RFECV-XGBoost feature screening model, calculating the importance of weather features and performing feature selection, and selecting weather factors with high importance as the model input features;
[0016] S204. Using the K-Means clustering algorithm to divide weather data, and dividing daily data into three weather conditions: sunny, cloudy and sudden change;
[0017] S205. Photovoltaic power decomposition and noise reduction based on CEEMDAN, decomposing the non-stationary sequence into multiple modal components, and using each modal component to characterize the characteristics of photovoltaic data.
[0018] Furthermore, the detailed steps of step S3 include:
[0019] S301. For the weather data after feature screening, select the first 60% of the data as the training set, select 30% of the data as the validation set, and the remaining 10% of the data as the test set;
[0020] S302. Take GRU and Informer as the base learners of the first layer of the model, input the test set and the validation set for model training, output the results of the first layer and use them as the new training set of the second layer;
[0021] S303. Take SVR as the meta-learner of the second layer, construct a Blenging ensemble learning framework, input the new training set and the test set for model training, and obtain the final XGB-GRU-Informer ensemble model.
[0022] Furthermore, in step S202, to eliminate the order of magnitude and dimension differences between each index data, first perform standardization processing, through the formula Standardize the original index data to the interval [a, b], and [-1, 1] standardization is adopted this time; where represents the data of each dimension after standardization processing, x i represents the original data of each dimension, and σ(x) represent the mean and variance of each dimension data.
[0023] Furthermore, in step S203, for D = {(x1, y1), (x2, y2), …, (x n , yn )} Training samples, where x i is a feature vector, and y i and are the actual value and the predicted value respectively. Assume that each decision tree model is The XGBoost objective function is:
[0024]
[0025] In the formula: L is the loss function, used to measure the error between the predicted value and the true value y i ; T is the number of regression trees; Ω controls the complexity of the model; γ is the penalty coefficient; λ is the regularization term coefficient; ω is the feature vector composed of the predicted values output by all leaf nodes through the decision tree model. The Boosting algorithm uses an additive model, and the predicted value of the strong classifier is equal to the sum of the predicted value of the current tree and the predicted value of the previous tree. Therefore, the objective function is transformed into the objective function of the previous t - 1 times plus the t - th time. Perform a second - order Taylor expansion on the objective function, obtain a univariate quadratic function about the feature vector and take the derivative to obtain the extreme point of the objective function, which is the optimal solution. Define the gain of splitting through the optimal solution. According to the greedy criterion, traverse all possible splitting points of the feature to calculate the gain value, and select the feature with the largest gain value for splitting. After splitting, the decision tree model is constructed. The importance of a certain feature j in the k - th decision tree can be calculated by formula (4).
[0026]
[0027] In the formula: is the importance of feature j in the k - th decision tree; v t is the feature related to node t; is the square of the loss value after node t is split; I is the sign function, which takes 1 when v t = j, and takes 0 when v t ≠ j; T - 1 is the number of non - leaf nodes. Assume there are M decision trees, then the global importance of feature j under the XGBoost model is measured by the average of the importance of feature j in all trees, and its calculation formula is as follows:
[0028]
[0029] Repeatedly calculate the error of the feature subset through K - fold cross - validation, and take the feature subset with the lowest error score as the prediction input to achieve the optimization of the prediction feature set. The steps to improve XGBoost using RFECV are as follows:
[0030] 1) Divide the sample set D = {x1,..., x m , y} into k sample subsets D1 ∼ Dk ;
[0031] 2) Set the j-th validation set as D j , and the training set as S = {D i | i = 1, 2,..., k, and i ≠ j};
[0032] 3) Input the training set S into the XGBoost training model and simultaneously calculate the importance J of feature x; Input the validation set D j into the trained XGBoost model to calculate the root mean square error (RMSE);
[0033] 4) Delete the feature x corresponding to the minimum feature importance J j in the training set S and the validation set D min ; min ;
[0034] 5) Save the RMSE corresponding to the m feature subsets obtained by sequentially deleting the features with the minimum importance;
[0035] 6) Calculate the average RMSE corresponding to the m feature subsets in the k-fold cross-validation as the cross-validation error score;
[0036] 7) Select the feature subset corresponding to the minimum cross-validation error score in D as the optimal feature subset F.
[0037] Furthermore, in step S204, the samples are clustered into k clusters according to the similarity, and finally, the similarity within the clusters is the highest and the similarity between the clusters is the lowest. The detailed steps are as follows:
[0038] 1) Randomly select k sample data from the data samples and use them as the original clustering centers {μ1, μ2,..., μ k};
[0039] 2) Calculate the Euclidean distance from the remaining samples to each initial center, and select the nearest initial clustering center to form k clusters. The distance formula is:
[0040]
[0041] where x is the sample in the sample space and μ i is the centroid of cluster C i ;
[0042] 3) Recalculate the clustering center for each cluster. The formula for calculating the clustering center is:
[0043]
[0044] Finally, repeat steps 2) and 3) until the similarity condition is met or the maximum number of iterations is reached and terminated. The termination condition is:
[0045] |μ n+1 -μ n |≤ε (8)
[0046] where ε is the threshold condition.
[0047] Further, in step S205, the empirical mode decomposition (EMD) method is used to smooth the data signal, and Gaussian white noise is introduced into the model to eliminate the aliasing effect. The decomposition steps of CEEMD are as follows:
[0048] Add Gaussian white noise with opposite signs and the same amplitude to the original signal to generate two sets of mixed signals:
[0049]
[0050] In the formula: is the i-th mixed signal with positive and negative Gaussian white noise added; is the i-th Gaussian white noise with opposite signs and the same amplitude;
[0051] Perform EMD decomposition on the composite signal to obtain two sets of IMF components, positive and negative:
[0052]
[0053] In the formula: is the j-th IMF component after the decomposition of the i-th positive noise addition; is the j-th IMF component after the decomposition of the i-th negative noise addition; and is the residual component.
[0054] Repeat the above steps M times, adding different Gaussian white noise each time;
[0055] Calculate the mean of each set of positive and negative IMF components to obtain the j-th IMF component:
[0056]
[0057] Further, in step S302, set the short-term photovoltaic power prediction parameters and determine the base learner model. During the operation of the neural network, the unit hidden state information at time t - 1 enters the reset gate of this memory unit from the previous memory unit. The reset gate decides whether to retain or discard this hidden state information, and then it is input into the update gate together with the input information at time t. Next, the update gate decides the writing degree of the retained hidden state information. Activate the reset gate and the update gate through the tanh function, combine them with the input information at time t, output the unit hidden state information at time t + 1, and transfer it to the next memory unit.
[0058] The GRU structure is as follows Figure 2 as shown. The specific calculation formula is as follows:
[0059] z t = σ(W z ·[h t-1 , x t ) (14)
[0060] r t = σ(W r ·[h t-1 , x t ) (15)
[0061]
[0062] In the formula: x t is the input sequence; h t is the short-term memory information of the output sequence at time t; c t is the long-term memory information of the output sequence at time t; is the temporary information of the memory cell at time t; f t is the forgetting gate state at time t, which determines the amount of information retained in c t ; o t-1 and i t and i t represent the states of the input gate and the output gate respectively.
[0063] The hyperparameters of the GRU neural network are set as follows: the number of nodes in the input layer is 7, corresponding to the 7 weather features after screening, the prediction step is 24, the number of nodes in the output layer is 1, the number of hidden layers is 2, the number of hidden neurons is 128, the training batch is 48, the number of training times of the GRU neural network is set to 200, the learning rate is 0.0001, and the parameter value to prevent overfitting is set to 0.05.
[0064] Secondly, the Informer consists of an encoder and a decoder. In the encoder layer, the ProbSparse self-attention mechanism is used instead of the traditional self-attention mechanism, reducing the model network dimension. Secondly, a generative decoding method is proposed in the decoder layer to directly generate all prediction results, improving the model prediction speed. Then, the self-attention mechanism often adopts the key-value-query mode, assigning the corresponding value to the key by comparing the similarity between the key and the query, while the Informer allows the key to only focus on the top u important Queries, that is
[0065]
[0066] Wherein, A represents the calculation method of the probabilistic sparse self-attention mechanism; V is the matrix representing value; K T is the transpose of the matrix representing Key; d is the input dimension; is the sparse matrix containing only the first u Queries.
[0067] In addition, Informer uses the data distillation method to assign higher weights to the dominant features and generate the focused self-attention feature map of the previous layer at the j+1 layer, as shown in the formula.
[0068]
[0069] In the formula: () att represents the attention module; is the matrix of the j-th layer; Convld represents a one-dimensional convolution operation on the time series, and ELU is used as the activation function. Finally, through a MaxPooling layer with a stride of 2, the computational complexity of each layer is halved, and the model can retain the information of the long input time series and solve the problem of temporal information loss.
[0070] For the decoder, the input time series is divided into two parts, namely the known sequence before the predicted point and the predicted sequence that needs to mask the future weather data That is
[0071]
[0072] The decoder of Informer adopts generative prediction and can directly output the multi-step prediction results. Through a single forward inference, this method can predict all the outputs of the long sequence. The hyperparameters of the Informer model are set as follows: the number of nodes in the input layer is 7, corresponding to the 7 weather features after screening respectively, the number of nodes in the output layer is 1, the number of encoder layers is 2, the encoder sequence length is 96, the number of decoder layers is 1, the decoder sequence length is 48, the prediction step length is 24, the number of self-attention heads is 8, the loss function is MSE, the training batch is 48, the number of hidden neurons is 512, the number of training times of the Informer model is set to 20, the learning rate is 0.0001, and the parameter value to prevent overfitting is set to 0.05.
[0073] Further, in step S303, parameter settings are performed on the meta-learner support vector machine regression (SVR); the Gaussian kernel function is selected as the SVR kernel function, and the penalty coefficient is often set to 2; the data is decomposed into a training set, a validation set, and a test set, which are input into multiple base learners in the first layer. The base learners are trained using the training set and fitted to the validation set, and the training results are input into the second-layer meta-learner for training, fitted to the test set, and the prediction results are output.
[0074] Further, in step S4, the KDE non-parametric estimation method is used to estimate the point prediction error of the photovoltaic power, and the confidence interval is obtained. Let q = [q1, q2,..., q n be the prediction error of the test set for photovoltaic prediction, then the total density function The kernel density estimate at point m can be defined as:
[0075]
[0076] In the formula, n represents the number of samples, h represents the window width, q m represents the m-th sample value of q, K( ) represents the kernel function, and the Gaussian kernel function is selected as the kernel function;
[0077] For a given confidence level, the probability that the predicted value of the photovoltaic power of the test set samples falls within the prediction interval is the prediction interval confidence level (PINC):
[0078] PINC = (1 - α) * 100% (22)
[0079] Based on formula (22), at this confidence level, the prediction interval of the sample is
[0080]
[0081] In the formula, and represent the upper and lower limits of the prediction interval respectively.
[0082] Further, in step S5, in order to measure the effect of interval prediction, PICP (interval coverage probability) and PINAW (interval average width) are used as indicators to evaluate the interval prediction effect; among them, PICP shows the probability that the true value falls within the prediction range at a given confidence level. If PICP is much smaller than the confidence level, it indicates that the prediction is invalid. On the contrary, the higher PICP is, the greater the probability that the true value falls between the estimated upper and lower limits. The expression is as follows:
[0083]
[0084] The average width of the prediction interval (PINAW) reflects the common width between the upper and lower bounds of the prediction interval. When the predicted PICP is the same, the smaller the PINAW, the better the prediction effect, which is expressed as:
[0085]
[0086] and are the upper and lower bounds of the prediction interval respectively;
[0087] And set the prediction scenario models for sunny, cloudy and sudden change weather, compare the interval prediction performance of the models under the three prediction scenarios, and compare the interval coverage rate and interval width of the models at different confidence levels.
[0088] Compared with the prior art, the present invention has the following beneficial effects:
[0089] A short-term photovoltaic power prediction model of XGBoost-GRU-Informer based on the Blending integration algorithm is proposed. First, the construction principle and prediction process of the prediction model are introduced through the data processing and model construction sections, and then the prediction effect and accuracy of the constructed model are verified through the empirical analysis section to achieve accurate short-term prediction of photovoltaic power.
[0090] In the data processing section, aiming at the problem of numerous relevant weather characteristics of photovoltaic power generation, XGBoost is used to calculate the feature importance, eliminate the features with lower importance, reduce data redundancy, and improve the model efficiency; K-Means clustering is used to divide the weather data into three categories to reduce the error caused by the power law change caused by weather fluctuations.
[0091] In the model construction section, an integrated model based on the Blending algorithm is constructed. Combining the model characteristics of GRU, Informer and SVR, the base learner and the meta-learner are built respectively to perform two-stage prediction on the photovoltaic power generation, reduce the model calculation amount, avoid the risk of overfitting, and improve the stability of short-term photovoltaic power prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] Figure 1 is the overall flowchart of a short-term photovoltaic power intelligent interval prediction method in an embodiment of the present invention;
[0093] Figure 2 is the importance ranking diagram of the RFECV-XGBoost feature screening model in an embodiment of the present invention;
[0094] Figure 3 is the weather clustering diagram based on the K-Means clustering algorithm in an embodiment of the present invention;
[0095] Figure 4 This is a comparison chart of the prediction results of the XGBoost-GRU-Informer integrated model constructed in the embodiments of the present invention;
[0096] Figure 5 This is the CEEMDAN decomposition result chart of the photovoltaic power in the embodiments of the present invention;
[0097] Figure 6 This is a schematic diagram of the error probability distribution and kernel density estimation curve under three kinds of weather in the embodiments of the present invention;
[0098] Figure 7 This is a schematic diagram of the cumulative error probability distribution under three kinds of weather in the embodiments of the present invention;
[0099] Figure 8 This is a schematic diagram for comparing the interval prediction results at different confidence levels in the embodiments of the present invention. Detailed implementation manners
[0100] For the convenience of those skilled in the art to understand, the present invention will be further described below in conjunction with embodiments and drawings. The content mentioned in the implementation manners does not limit the present invention.
[0101] As Figure 1 shown, this embodiment provides a short-term photovoltaic power intelligent interval prediction method, including the steps:
[0102] S1. Collect weather data and establish an original data set;
[0103] S2. Perform data processing on the original data set. The data processing includes outlier detection, sorting of weather feature importance, K-Means weather clustering, data modal decomposition, determination of feature input and output, and data set division;
[0104] S3. Construct a GRU-Informer and SVR double-layer prediction model based on the Blending integrated learning framework;
[0105] S4. Estimate the point prediction error of the photovoltaic power by the KDE non-parametric estimation method, and estimate the upper and lower limits of the photovoltaic power prediction interval at a given confidence level;
[0106] S5. Conduct comparative analysis and error analysis on the prediction results.
[0107] S201. Processing of abnormal data, detecting and replacing power abnormal data;
[0108] S202. Standardize the data set to eliminate the differences in the order of magnitude and dimension between the data of each index; In step S202, to eliminate the differences in the order of magnitude and dimension between the data of each index, first perform standardization processing, through the formula Standardize the original index data to the interval [a, b]. In this case, [-1, 1] standardization is adopted. Among them represents the data of each dimension after standardization, and x i represents the original data of each dimension, and σ(x) represent the mean and variance of the data of each dimension.
[0109] S203. Introduce recursive feature elimination cross-validation (RFECV) to improve XGBoost, construct an RFECV-XGBoost feature screening model, calculate the importance of weather features and perform feature selection, and select weather factors with high importance as the input features of the model;
[0110] In step S203, for the training samples D = {(x1, y1), (x2, y2), …, (x n , y n )}, where x i is the feature vector, and y i and are the actual value and the predicted value respectively. Assume that each decision tree model is The XGBoost objective function is:
[0111]
[0112] In the formula: L is the loss function, used to measure the error between the predicted value and the true value y i ; T is the number of regression trees; Ω controls the complexity of the model; γ is the penalty coefficient; λ is the regularization term coefficient; ω is the feature vector composed of the predicted values output by all leaf nodes through the decision tree model. The Boosting algorithm uses an additive model, and the predicted value of the strong classifier is equal to the sum of the predicted value of the current tree and the predicted value of the previous tree. Therefore, the objective function is transformed into the objective function of the previous t - 1 times plus the t-th time. Perform a second-order Taylor expansion on the objective function, obtain a unary quadratic function about the feature vector and take the derivative to obtain the extreme point of the objective function, that is, the optimal solution. Define the gain of splitting through the optimal solution. According to the greedy criterion, traverse all possible splitting points of the features to calculate the gain value, and select the feature with the largest gain value for splitting. After the splitting is completed, the decision tree model is constructed. The importance of a certain feature j in the k-th decision tree can be calculated by formula (4).
[0113]
[0114] In the formula: is the importance of feature j in the k-th decision tree; v t is the feature related to node t; is the square of the loss value after the node t splits; I is the sign function, which takes 1 when v t = j and takes 0 when v t ≠ j; T - 1 is the number of non - leaf nodes. Suppose there are M decision trees, then the global importance of feature j in the XGBoost model is measured by the average of the importance of feature j in all trees, and its calculation formula is as follows:
[0115]
[0116] Repeatedly calculate the error of the feature subset through k - fold cross - validation, and take the feature subset with the lowest error score as the prediction input to achieve the optimization of the prediction feature set. The steps to improve XGBoost using RFECV are as follows:
[0117] 1) Divide the sample set D = {x1,..., x m , y} into k sample subsets D1 ∼ D k ;
[0118] 2) Set the j - th validation set as D j , and the training set as S = {D i |i = 1, 2,..., k, and i ≠ j};
[0119] 3) Input the training set S into the XGBoost training model and calculate the importance J of feature x at the same time; input the validation set D j into the trained XGBoost model to calculate the root mean square error (RMSE);
[0120] 4) Delete the feature x corresponding to the minimum feature importance J j in the training set S and the validation set D min ; min ;
[0121] 5) Save the RMSE corresponding to the m feature subsets obtained by successively deleting the features with the minimum importance;
[0122] 6) Calculate the average RMSE corresponding to the m feature subsets in the k - fold cross - validation as the cross - validation error score;
[0123] 7) Select the feature subset corresponding to the minimum cross - validation error score in D as the optimal feature subset F.
[0124] The summary of weather features is shown in Table 1, and the cross - validation scores of the feature subsets are as Figure 2 shown; the weather screening results are shown in the appendix Figure 3 ;
[0125] Table 1 Summary of weather features
[0126]
[0127] As Figure 3 shown, according to the cross-validation scores of different feature input subsets, it can be found that when the feature input subset is 4, the cross-validation error score is the lowest. Therefore, the top 4 in terms of importance scores are selected as the optimal feature input subset. The four features with the highest importance scores are: Direct Normal Irradiance (DNI), Global Horizontal Irradiance (GHI), Temperature (T), and Wind Direction (WD). Therefore, DNI, GHI, T, and WD are selected as the input variables for photovoltaic power point prediction.
[0128] S204. Use the K-Means clustering algorithm to partition the weather data, and divide the daily data into three weather conditions: sunny, cloudy, and abrupt change;
[0129] In step S204, the samples are clustered into k clusters according to similarity, and finally the similarity within the clusters is the highest and the similarity between the clusters is the lowest. The detailed steps are as follows:
[0130] 1) Randomly select k sample data from the data samples and use them as the original clustering centers {μ1, μ2, …, μ k};
[0131] 2) Calculate the Euclidean distance from the remaining samples to each initial center, and select the initial clustering center with the closest distance to form k clusters. The distance formula is:
[0132]
[0133] where x is the sample in the sample space and μ i is the centroid of cluster C i ;
[0134] 3) Recalculate the clustering center for each cluster. The formula for calculating the clustering center is:
[0135]
[0136] Finally, repeat steps 2) and 3) until the similarity condition is met or the maximum number of iterations is reached and terminated. The termination condition is:
[0137] |μ n+1 - μ n | ≤ ε (8)
[0138] where ε is the threshold condition.
[0139] The specific clustering diagram is as Figure 4 shown. The partitioning results of the weather data set are shown in Table 2.
[0140] Table 2 Partitioning Results of Weather Data Set
[0141]
[0142] S205. Photovoltaic power decomposition and noise reduction based on CEEMDAN, decomposing the non-stationary sequence into multiple modal components and using each modal component to characterize the characteristics of photovoltaic data.
[0143] In step S205, the empirical mode decomposition (EMD) method is used to smooth the data signal, and Gaussian white noise is introduced into the model to eliminate the aliasing effect. The decomposition steps of CEEMD are as follows:
[0144] Add Gaussian white noise with opposite signs and the same amplitude to the original signal to generate two sets of mixed signals:
[0145]
[0146] In the formula: is the i-th mixed signal with positive and negative Gaussian white noise added; is the i-th Gaussian white noise with opposite signs and the same amplitude;
[0147] Perform EMD decomposition on the composite signal to obtain two sets of IMF components, positive and negative:
[0148]
[0149] In the formula: is the j-th IMF component after the decomposition of the i-th positive noise addition; is the j-th IMF component after the decomposition of the i-th negative noise addition; and are the residual components.
[0150] Repeat the above steps M times, adding different Gaussian white noises each time;
[0151] Calculate the mean of each set of positive and negative IMF components to obtain the j-th IMF component:
[0152]
[0153] Taking the sudden change weather with large fluctuations in photovoltaic power as an example, the decomposition results are as shown in the appendix Figure 5 shown. The fluctuation frequencies of IMF1 - IMF3 are relatively high, which conforms to the randomness of photovoltaic power data; IMF4 - IMF5 have a certain periodicity, which coincides with the morning and evening fluctuation periods of photovoltaic power; the oscillation amplitudes of IMF6 - IMF7 are different, which has a large correlation with the photovoltaic power fluctuations in sudden change weather; the fluctuation frequencies of IMF8 - IMF11 gradually decrease, and the curves gradually become smooth, belonging to the long-term components.
[0154] The detailed steps of step S3 include:
[0155] S301. For the weather data after feature screening, select the first 60% of the data as the training set, 30% of the data as the validation set, and the remaining 10% of the data as the test set. The training data is shown in Table 3.
[0156] Table 3 Weather data set after screening
[0157]
[0158] S302. Take GRU and Informer as the base learners in the first layer of the model. Input the test set and the validation set for model training, and output the results of the first layer as the new training set for the second layer.
[0159] In step S302, set the parameters for short-term photovoltaic power prediction and determine the base learner model. During the operation of the neural network, the unit hidden state information at time t - 1 enters the reset gate of this memory unit from the previous memory unit. The reset gate decides whether to retain or discard this hidden state information, and then it is input into the update gate together with the input information at time t. Next, the update gate decides the degree of writing of the retained hidden state information. The reset gate and the update gate are activated through the tanh function, combined with the input information at time t, to output the unit hidden state information at time t + 1 and transmit it to the next memory unit.
[0160] The specific calculation formula of GRU is as follows:
[0161] z t =σ(W z ·[h t-1 ,x t ) (14)
[0162] r t =σ(W r ·[h t-1 ,x t ) (15)
[0163]
[0164] In the formula: x t is the input sequence; h t is the short-term memory information of the output sequence at time t; c t is the long-term memory information of the output sequence at time t; is the temporary information of the memory unit at time t,; f t is the forgetting gate state at time t, which determines the amount of information retained in the memory state c t ; o t-1 and i t represent the states of the input gate and the output gate respectively. t
[0165] The hyperparameters of the GRU neural network are set as follows: the number of nodes in the input layer is 7, corresponding to the 7 selected weather features respectively, the prediction step is 24, the number of nodes in the output layer is 1, the number of hidden layers is 2, the number of hidden neurons is 128, the training batch is 48, the number of training times of the GRU neural network is set to 200, the learning rate is 0.0001, and the parameter value to prevent overfitting is set to 0.05.
[0166] Secondly, Informer consists of an encoder and a decoder. In the encoder layer, the ProbSparse self-attention mechanism is used instead of the traditional self-attention mechanism, reducing the model network dimension. Secondly, in the decoder layer, a generative decoding method is proposed to directly generate all prediction results, improving the model prediction speed. Then, the self-attention mechanism often adopts the key-value-query mode, assigning the corresponding value to the key by comparing the similarity between the key and the query, while Informer makes the key only focus on the top u important Queries, that is
[0167]
[0168] In the formula, A represents the calculation method of the ProbSparse self-attention mechanism; V is the matrix representing value; K T is the transpose of the matrix representing Key; d is the input dimension; is the sparse matrix containing only the top u Queries.
[0169] In addition, Informer uses the data distillation method to assign higher weights to the dominant features and generates the focused self-attention feature map of the previous layer at the j+1 layer, as shown in the following formula:
[0170]
[0171] In the formula: () att represents the attention module; is the matrix of the j-th layer; Convld represents a one-dimensional convolution operation on the time series, and ELU is used as the activation function. Finally, through a MaxPooling layer with a stride of 2, the computational complexity of each layer is halved, and the model can retain the information of the long input time series, solving the problem of temporal information loss.
[0172] For the decoder, the input time series is divided into 2 parts, namely the known sequence before the predicted point and the prediction sequence that needs to mask the future weather data, that is
[0173]
[0174] The decoder of Informer adopts generative prediction and can directly output multi-step prediction results. Through a single forward inference, this method can predict all outputs of a long sequence. The hyperparameters of the Informer model are set as follows: the number of nodes in the input layer is 7, corresponding to the 7 selected weather features respectively; the number of nodes in the output layer is 1; the number of encoder layers is 2; the encoder sequence length is 96; the number of decoder layers is 1; the decoder sequence length is 48; the prediction step is 24; the number of heads of self-attention is 8; the loss function is MSE; the training batch is 48; the number of hidden neurons is 512; the number of training times of the Informer model is set to 20; the learning rate is 0.0001; the parameter value to prevent overfitting is set to 0.05.
[0175] S303: Use SVR as the meta-learner in the second layer to construct a Blenging ensemble learning framework. Input the new training set and test set for model training to obtain the final XGB-GRU-Informer ensemble model.
[0176] In step S303, set the parameters of the meta-learner support vector machine regression (SVR); select the Gaussian kernel function as the SVR kernel function, and the penalty coefficient is often set to 2; decompose the data into a training set, a validation set and a test set, and input them into multiple base learners in the first layer. Use the training set to train the base learners and fit the validation set, input the training results into the meta-learner in the second layer for training, fit the test set and output the prediction results.
[0177] In step S4, use the KDE non-parametric estimation method to estimate the point prediction error of the photovoltaic power, obtain the confidence interval, and set q = [q1, q2,..., q n as the prediction error of the test set for photovoltaic prediction, then the total density function The kernel density estimation at point m can be defined as:
[0178]
[0179] In the formula, n represents the number of samples, h represents the window width, q m represents the m-th sample value of q, and K( ) represents the kernel function. In this paper, the Gaussian kernel function is selected as the kernel function;
[0180] For a given confidence level, the probability that the predicted value of the photovoltaic power of the test set samples falls within the prediction interval is the prediction interval confidence level (PINC):
[0181] PINC = (1 - α) * 100% (22)
[0182] Based on formula (22), at this confidence level, the sample has a prediction interval of
[0183]
[0184] wherein and represent the upper and lower limits of the prediction interval respectively.
[0185] In step S5, in order to measure the effect of interval prediction, PICP (Prediction Interval Coverage Probability) and PINAW (Prediction Interval Normalized Average Width) are used as indicators to evaluate the effect of interval prediction; among them, PICP shows the probability that the true value falls within the prediction range at a given confidence level. If PICP is much smaller than the confidence level, it indicates that the prediction is invalid. On the contrary, the higher the PICP, the greater the probability that the true value falls between the estimated upper and lower limits. The expression is as follows:
[0186]
[0187] The average width of the prediction interval (PINAW) reflects the common width between the upper and lower limits of the prediction interval. When the predicted PICP is the same, the smaller the PINAW, the better the prediction effect, which is expressed as:
[0188]
[0189] and are the upper and lower bounds of the prediction interval respectively.
[0190] Interval estimations are respectively performed on the point prediction errors of four models, namely LSTM(#1), Informer(#2), CEEMD-GRU(#3), and the proposed CEEMD-GRU-Informer-SVR(#4) in this paper, and their error estimation curves are plotted using the Gaussian kernel function. As Figure 6 shown, it is the probability distribution of the prediction error and the kernel density estimation curve. It can be seen that the probability distributions of the point prediction errors for sunny and cloudy days are relatively average, while the prediction error distribution for sudden changes is relatively unstable. The Gaussian kernel function curve can better fit the probability distributions of the point prediction errors under the three weather conditions, indicating that the interval estimation effect is good.
[0191] Estimate the upper and lower limits of the calculation interval according to the error cumulative function, and output the short-term photovoltaic power interval prediction result. The error cumulative function is as Figure 7 shown.
[0192] Table 4 presents the interval coverage rate and interval width of the comparison models at different confidence levels under different weather conditions. It can be found that #4 has a relatively high interval coverage rate at the three confidence levels of 95%, 90%, and 85%. As shown in Table 4 and Figure 8It can be found that at the 95% confidence level, the interval coverage rate of the #4 model under sunny and cloudy days can reach 96.5%, and the coverage effect is good; the interval coverage rate of the #4 model under mutant weather is above 95%, indicating that it can be confirmed that the model in this paper has a high interval prediction accuracy and is less affected by weather fluctuations. Taking the interval prediction results of sunny weather as an example, the interval prediction effects of multiple models are compared at the 95% confidence level. Compared with other models, the model constructed in the present invention can achieve a higher coverage rate on the premise of a narrower prediction interval width, further verifying that the interval prediction of the short-term photovoltaic power of the model in the present invention is relatively accurate.
[0193] Table 4 Interval prediction error indicators at different confidence levels
[0194]
[0195] Compared with the prior art, the present invention has the following beneficial effects:
[0196] A short-term photovoltaic power prediction model of XGBoost-GRU-Informer based on the Blending integration algorithm is proposed. First, the construction principle and prediction process of the prediction model are introduced through the data processing and model construction sections, and then the prediction effect and accuracy of the constructed model are verified through the empirical analysis section to achieve accurate short-term prediction of photovoltaic power.
[0197] In the data processing section, aiming at the problem of numerous relevant weather characteristics of photovoltaic power generation, XGBoost is used to calculate the feature importance, eliminate the features with lower importance, reduce data redundancy, and improve the model efficiency; K-Means clustering is used to divide the weather data into three categories to reduce the error caused by the change of power law caused by weather fluctuations.
[0198] In the model construction section, an integrated model based on the Blending algorithm is constructed. Combining the model characteristics of GRU, Informer and SVR, the base learner and the meta-learner are built respectively to perform two-stage prediction on the photovoltaic power generation, reduce the model calculation amount, avoid the risk of overfitting, and improve the stability of short-term photovoltaic power generation prediction.
[0199] The above has introduced in detail a short-term photovoltaic power intelligent interval prediction method provided by this application. The description of the specific embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A short-term photovoltaic power intelligent interval prediction method, characterized in that: Includes steps: S1, collect weather data and establish original data set; S2, data processing is performed on the original data set, including outlier detection, weather feature importance ranking, K-Means weather clustering, data mode decomposition, feature input and output determination, and data set division; S3, build a GRU-Informer and SVR double-layer prediction model based on the Blending ensemble learning framework; S4. Estimate the point prediction error of PV power by using the KDE nonparametric estimation method, and estimate the upper and lower limits of the PV power prediction interval at a given confidence level; S5. Conduct comparative analysis and error analysis on the prediction results.
2. A short-term photovoltaic power intelligent interval prediction method according to claim 1, characterized in that: The detailed steps of step S2 include: S201, abnormal data processing, detecting and replacing abnormal power data; S202, standardize the data set to eliminate the magnitude and dimension differences between the indicator data; S203, introducing recursive feature elimination cross-validation to improve XGBoost, constructing the RFECV-XGBoost feature screening model, calculating the importance of weather features and performing feature selection, and selecting weather factors with high importance as model input features; S204, using K-Means clustering algorithm to divide the weather data, and divide the daily data into three types of weather: sunny, cloudy and sudden weather; S205. Photovoltaic power decomposition and denoising based on CEEMDAN, decomposing the non-stationary series into multiple modal components, and using each modal component to characterize the characteristics of photovoltaic data.
3. A short-term photovoltaic power intelligent interval prediction method according to claim 2, characterized in that: The detailed steps of step S3 include: S301, for the weather data after feature screening, select the first 60% of the data as a training set, select 30% of the data as a validation set, and the remaining 10% of the data as a test set; S302, using GRU and Informer as the base learners of the first layer of the model, inputting the test set and the validation set for model training, outputting the first layer results and using them as the new training set for the second layer; S303. Use SVR as the meta-learner of the second layer, build a Blenging ensemble learning framework, input new training sets and test sets for model training, and obtain the final XGB-GRU-Informer ensemble model.
4. A short-term photovoltaic power intelligent interval prediction method according to claim 3, characterized in that: In step S202, the order of magnitude and dimension differences between the indicator data are eliminated by first performing standardization. The original indicator data are standardized to the interval [a, b], and the [-1, 1] standardization is adopted this time; Represents the dimensional data after standardization, x i Represents the original dimensional data, and σ(x) represent the mean and variance of each dimensional data.
5. A short-term photovoltaic power intelligent interval prediction method according to claim 4, characterized in that: In step S203, for D = {(x1, y1), (x2, y2), ..., (x n ,y n )} training samples, where x i is the feature vector, y i and The actual value and the predicted value are respectively, assuming that each decision tree model is The XGBoost objective function is: Where: L is the loss function, which is used to measure the predicted value and the true value y i The error between them; T is the number of regression trees; Ω controls the complexity of the model; γ is the penalty coefficient; λ is the regular term coefficient; ω is the feature vector composed of the predicted values output by all leaf nodes through the decision tree model.
6. A short-term photovoltaic power intelligent interval prediction method according to claim 5, characterized in that: The steps to improve XGBoost using recursive feature elimination cross-validation are as follows: 1) Set the sample set D = {x1, ..., x m , y} is divided into k sample subsets D1~D k ; 2) Set the jth validation set to D j , the training set is S = {D i |i=1,2,...,k,and i≠j}; 3) The training set S is input into the XGBoost training model and the importance J of feature x is calculated at the same time; the validation set D j Input the trained XGBoost model to calculate the root mean square error (RMSE); 4) Delete the training set S and the validation set D j The minimum feature importance J min The corresponding feature x min ; 5) Save the RMSE corresponding to the m feature subsets obtained by successively deleting the features with the least importance; 6) Calculate the average RMSE corresponding to the m feature subsets in the k-fold cross validation as the cross validation error score; 7) Select the feature subset corresponding to the smallest cross-validation error score in D as the optimal feature subset F.
7. A short-term photovoltaic power intelligent interval prediction method according to claim 6, characterized in that: In step S204, samples are clustered into k clusters according to similarity, and finally the highest similarity in the cluster and the lowest similarity between clusters are achieved. The detailed steps are as follows: 1) Randomly select k sample data from the data sample and use them as the original cluster centers {μ1,μ2,…,μ k }; 2) Calculate the Euclidean distance from the remaining samples to each initial center, and select the initial clustering center with the closest distance to form k clusters. The distance formula is: Where x is the sample in the sample space, μ i Cluster C i The centroid of 3) Recalculate the cluster center for each cluster. The formula for calculating the cluster center is: Finally, repeat steps 2) and 3) until the similarity condition is met or the maximum number of iterations is reached and the termination condition is: |m n+1 -m n |≤ε (8) Where ε is the threshold condition.
8. A short-term photovoltaic power intelligent interval prediction method according to claim 7, characterized in that: In step S205, the data signal is stabilized using the empirical mode decomposition (EMD) method, and Gaussian white noise is introduced into the model to eliminate the aliasing effect. The decomposition steps of CEEMD are as follows: Add Gaussian white noise of opposite signs and the same amplitude to the original signal to generate two sets of mixed signals: Where: is the i-th mixed signal with positive and negative Gaussian white noise added; is the i-th Gaussian white noise with opposite signs and the same amplitude; Perform EMD decomposition on the composite signal to obtain two groups of positive and negative IMF components: Where: is the jth IMF component after adding positive noise decomposition to the i-th component; is the jth IMF component after adding negative noise decomposition to the i-th component; and is the residual component; Repeat the above steps M times, adding different Gaussian white noise each time; The jth IMF component is obtained by averaging the sets of positive and negative IMF components:
9. A short-term photovoltaic power intelligent interval prediction method according to claim 8, characterized in that: In step S302, short-term photovoltaic power prediction parameters are set and the base learner model is determined. During the operation of the neural network, the hidden state information of the unit at time t-1 enters the reset gate of the memory unit from the previous memory unit. The reset gate decides whether to retain or discard the hidden state information, and then inputs the hidden state information together with the input information at time t into the update gate. Next, the update gate determines the writing degree of the retained hidden state information, activates the reset gate and the update gate through the tanh function, combines with the input information at time t, outputs the hidden state information of the unit at time t+1, and passes it to the next memory unit.
10. A short-term photovoltaic power intelligent interval prediction method according to claim 9, characterized in that: In step S4, the KDE nonparametric estimation method is used to estimate the point prediction error of photovoltaic power and obtain the confidence interval. Let q = [q1, q2, ..., q n ] is the prediction error of the test set of photovoltaic prediction, then the total density function The kernel density estimate at point m can be defined as: In the formula, n represents the number of samples, h represents the window width, and q m represents the mth sample value of q, K() represents the kernel function, and this paper selects the Gaussian kernel function As a kernel function; For a given confidence level, the probability that the PV power prediction value of the test set sample falls within the prediction interval is the prediction interval confidence (PINC): PINC=(1-α)*100% (22) Based on formula (22), at this confidence level, the sample The prediction interval for In the formula, and Represent the upper and lower limits of the prediction interval respectively.