Non-parameter Probability Prediction Method for Ultra-short-term Photovoltaic Power Based on Fuzzy Sample Particles

Through the photovoltaic power prediction method of fuzzy sample granulation and adaptive expansion, the problem of insufficient uncertainty description in photovoltaic power prediction is solved, high-precision probability prediction is achieved, and the reliability of power grid scheduling is improved.

CN115907235BActive Publication Date: 2025-07-25JIANGSU OCEAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310018364.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-06
Publication Date
2025-07-25
Estimated Expiration
2043-01-06

AI Technical Summary

Technical Problem

The existing photovoltaic power prediction methods lack descriptions of uncertainty. Traditional deterministic prediction cannot provide objective grid scheduling information, and it is difficult to accurately construct prediction intervals based on probability prediction based on parameterized models.

Method used

The ultra-short-term non-parametric probability prediction method of photovoltaic power based on fuzzy sample particles is adopted. By constructing a quantile regression model of the limit learning machine, combining hierarchical clustering and sample granulation processing, the training sample set is optimized, the sample data is expanded adaptively, and the prediction accuracy is improved.

Benefits of technology

It improves the confidence and accuracy of photovoltaic power prediction, reduces the sensitivity to iterative calculations and outliers, and provides a more scientific reference for power scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115907235B_ABST
    Figure CN115907235B_ABST
Patent Text Reader

Abstract

The present invention discloses a photovoltaic power ultra-short-term non-parametric probability prediction method based on fuzzy sample particles. First, the present invention analyzes the sample characteristics of the photovoltaic power time series, combines the meteorological forecast irradiance data in the numerical weather forecast, and constructs a sample particle processing method based on merging-decomposition. In addition, a hierarchical clustering method based on sample particles is studied. According to different clusters of different sample particles, combined with the characteristics of the preset sample input to be predicted, the sample particles are adaptively expanded by a reasonable multiple and then restored to the original samples to realize the dynamic adjustment of the sample weights. Finally, based on the quantile regression model of the extreme learning machine, the photovoltaic probability prediction model is trained respectively according to the situation that the sample to be measured belongs to different clusters. The method of the present invention has good reliability and overall performance, and greatly improves the practicability and accuracy of the ultra-short-term non-parametric probability prediction technology for photovoltaic power generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technology of ultra - short - term non - parametric probability prediction of photovoltaic power, and particularly to an ultra - short - term non - parametric probability prediction method of photovoltaic power based on fuzzy sample particles. Background Art

[0002] In recent years, with the continuous increase in the installed capacity of renewable energy, photovoltaic has become one of the most important renewable energies. However, the photovoltaic power generation has uncertainty and randomness, which limits its application and development. Traditional photovoltaic power prediction focuses on deterministic prediction, that is, point prediction, lacking the description of uncertainty and unable to provide more objective and comprehensive information for power grid control and dispatching departments. Different from the point prediction method that directly predicts a definite value, the probability interval prediction method calculates the prediction range under a preset confidence interval. Compared with the traditional point prediction, the result of interval prediction has a higher credibility and can provide more scientific data reference for reasonable power dispatching, etc.

[0003] At present, probability interval prediction mainly adopts parameter - based models. The construction of the prediction interval is based on two parts: point prediction and uncertainty analysis. After the point prediction is completed, it is assumed that the photovoltaic power prediction error satisfies a certain distribution, such as β - distribution, standard normal distribution, etc. Then, according to the pre - assumed distribution, the prediction error is calculated, and then added to the point prediction value to form the calculation of the upper and lower limits of the interval. However, the actual photovoltaic power has large fluctuations and strong complexity, and it is difficult to determine the actual error distribution.

[0004] The probability interval prediction method based on the quantile regression model has been increasingly valued by technical personnel. The traditional linear quantile regression method is often used for regression analysis and prediction in statistical regression analysis. In order to improve the traditional quantile regression model, currently, methods include using the extreme learning machine model to improve the quantile regression method and improving the linear model to a non - linear model.

[0005] The clustering method is one of the most effective data - mining techniques, which can effectively improve the accuracy of model training. Clustering theories such as hierarchical clustering method have been widely used in power prediction technology, and currently, it is mainly for deterministic prediction models. The existing probability prediction models based on the clustering method are mainly for parameter - based model methods, that is, first, deterministic prediction is carried out based on the clustering method, and then the prediction error is analyzed to calculate the prediction interval. Since the deterministic prediction method is limited by the accuracy of the error assumption, and the outliers of calculating the sample distance through multiple iterations are likely to reduce the sample clustering accuracy, it is of great significance to study the probability prediction method based on the clustering method and non - parametric models. Summary of the Invention

[0006] Objective of the Invention: Aiming at the problems existing in the prior art, the present invention provides a non-parametric probability prediction method for ultra-short-term photovoltaic power based on fuzzy sample particles, so as to improve the credibility of photovoltaic power prediction.

[0007] Technical Solution: The non-parametric probability prediction method for ultra-short-term photovoltaic power based on fuzzy sample particles according to the present invention is characterized by including the following steps:

[0008] (1) Construct and initialize a quantile regression model based on an extreme learning machine, and establish a fuzzy sample distance based on the historical photovoltaic power time series and the corresponding change amount of irradiance prediction value, so as to form a clustering sample set and a training sample set;

[0009] (2) Use the matrix formed by all samples of the clustering sample set as the initial sample particles of the sample particle set, and based on the sample particle difference criterion, iterate and loop the process of decomposition and combination of sample particles until all sample particles can no longer be decomposed and synthesized;

[0010] (3) Cluster all sample particles of the sample particle set according to hierarchical clustering to obtain N clusters;

[0011] (4) Calculate the sample expansion multiples when the sample to be predicted belongs to different clusters according to the similarity between different clusters, reconstruct the training sample set according to the sample expansion multiples, and then train the quantile regression model based on the extreme learning machine according to the reconstructed training sample set to obtain extreme learning machines for different clusters;

[0012] (5) According to the cluster to which the sample to be predicted belongs, input the sample to be predicted into the extreme learning machine of the belonging cluster to predict the probability prediction interval of photovoltaic power.

[0013] Further, step (1) specifically includes the following steps:

[0014] (1.1) Construct and initialize the hidden layer coefficients and thresholds of the extreme learning machine;

[0015] (1.2) Set the upper and lower quantile percentages and α of the confidence interval of the quantile regression model;

[0016] (1.3) Normalize the historical photovoltaic power time series and the meteorological forecast irradiance data;

[0017] (1.4) Calculate the regional meteorological forecast irradiance prediction value based on the normalized meteorological forecast irradiance data:

[0018]

[0019] where R iDenote the predicted value of the regional meteorological forecast irradiance at the \(i\)-th moment, \(M\) denote the number of power stations, \(C\) u and \(I\) u respectively denote the capacity of the \(u\)-th power station and the predicted value of the meteorological forecast irradiance at the current moment;

[0020] (1.5) Calculate the change in the predicted irradiance value \(\Delta R\) i :

[0021] \(\Delta R\) i \(= R\) i \(- R\) i-1

[0022] In the formula, \(R\) i-1 denotes the predicted value of the regional meteorological forecast irradiance at the \((i - 1)\)-th moment calculated according to the formula listed in step (1.4), and represents the potential change trend of future photovoltaic power in terms of the fuzzy change in the predicted irradiance value;

[0023] (1.6) Construct a training sample set based on the normalized historical photovoltaic power time series as follows:

[0024]

[0025] In the formula, \(D\) denotes the sample set, \(x\) i denotes the historical photovoltaic power time series before the \(i\)-th moment, as the input, \(y\) i denotes the power at the \(i\)-th moment, as the output, and the two form the \(i\)-th training sample, \(S\) denotes the number of samples;

[0026] (1.7) Construct a clustering sample set as:

[0027]

[0028] In the formula, \(St\) denotes the clustering sample set.

[0029] Furthermore, step (2) specifically includes the following steps:

[0030] (2.1) Take the matrix formed by all the samples in the clustering sample set as the initial sample particle and add it to the sample particle set;

[0031] (2.2) Sequentially perform the following judgment and decomposition process on each sample particle in the sample particle set:

[0032] Decompose it into several pairs of sample sub-particle candidates in different ways to form a sample sub-particle candidate pair set, and calculate the difference index value between each pair of candidates. If the maximum value of the difference index values of all candidate pairs is greater than or equal to the threshold, then decompose the current sample particle into the two sub-particles in the sample sub-particle candidate pair corresponding to the maximum difference index value, otherwise do not decompose;

[0033] (2.3) For every two arbitrary sample particles in the sample particle set, calculate the difference index value between these two sample particles. If the minimum value of all difference index values is less than the threshold, then merge the two sample particles corresponding to the minimum value into one sample particle; otherwise, do not merge.

[0034] (2.4) Return and execute steps (2.2) to (2.3) until all sample particles in the sample particle set cannot be decomposed or merged.

[0035] Further, step (2.2) specifically includes the following steps:

[0036] (2.2.1) Select a sample particle from the sample particle set, denoted as sk.

[0037] (2.2.2) Select any variable r from all variables of the sample particle sk as the variable to be analyzed.

[0038] (2.2.3) Sort all samples within the sample particle sk in ascending order according to the value of the variable to be analyzed.

[0039] (2.2.4) Take the value of the variable to be analyzed of the v-th sample after sorting as the decomposition threshold, where v = 1,..., n sk - 1, and form a sample sub-particle sk from the samples within the sample particle sk whose values of the variable to be analyzed are greater than the decomposition threshold 1,v,r , and form another sample sub-particle sk from the remaining samples 2,v,r , and the two sample sub-particles form a pair of sample sub-particle candidates to be added to the sample sub-particle candidate pair set SK, where n sk represents the number of samples within the sample particle sk;

[0040] (2.2.5) Return and execute steps (2.2.2) to (2.2.4) until all variables of the sample particle sk are traversed, that is, r takes each value in the variable set R in turn, and obtain the sample sub-particle candidate pair set SK = {(sk 1,v,r , sk 2,v,r ) | r ∈ R, v = 1,..., n sk - 1};

[0041] (2.2.6) Calculate the difference index value between each candidate pair in the sample sub-particle candidate pair set and obtain the maximum value F max , if the maximum value F max is greater than or equal to the threshold, then decompose the sample particle sk into the sample sub-particle pair corresponding to the maximum value F max , and replace the sample particle sk in the sample particle set; otherwise, do not decompose.

[0042] (2.2.7) Return to execute steps (2.2.1) to (2.2.6) until all the sample particles in the sample particle set are traversed, and complete the decomposition of all the sample particles in the sample particle set in one cycle calculation.

[0043] Further, step (2.3) specifically includes the following steps:

[0044] (2.3.1) Combine any two sample particles in the sample particle set to form a sample particle pair;

[0045] (2.3.2) Calculate the difference index value between the two sample particles of all the sample particle pairs;

[0046] (2.3.3) Obtain the minimum value F of the difference index value min , if the minimum value F min is less than the threshold, then combine the two sample particles of the sample particle pair corresponding to the minimum value F min in the sample particle set into one sample particle, otherwise do not combine.

[0047] Further, the calculation method of the difference index value is:

[0048]

[0049]

[0050]

[0051]

[0052]

[0053]

[0054] In the formula, Q(e, h) represents the difference index value, e and h represent any two sample particles, F() represents the F-test value, () T represents the transpose of the matrix, A represents the sum of the squares of the sample particle differences, B represents the sum of the cross-product matrices, e p represents the vector corresponding to the p-th sample in e, h q represents the vector corresponding to the q-th sample in h; and respectively represent the average value vectors of e and h, n e and n h are respectively the total number of samples in e and h; d represents the number of sample variables;

[0055] The threshold of the difference index value is α F which is a preset ratio.

[0056] Further, step (4) specifically includes the following steps:

[0057] (4.1) Set χ = 1;

[0058] (4.2) Calculate the similarity between the centers of other clusters and the center of the χ-th cluster when assuming that the sample to be predicted belongs to the χ-th cluster:

[0059]

[0060] In the formula, Sim χ,k represents the similarity between the center of the cluster where the sample to be predicted is located and the center of the k-th cluster, respectively represent the center of the cluster where the sample to be predicted is located the center of the k-th cluster of the j-th variable value, d represents the number of variables of the clustering samples; the center of each cluster is the mean value of the samples in each cluster, and the formula for its j-th variable value is: g l,j represents the j-th variable of the l-th sample in the cluster, T g,* then represents the number of samples in the *-th cluster; measures the difference of the power time series of the clustering samples; measures the difference of the irradiance change trend, are the d-th variables of each clustering sample, representing ΔR in their respective samples, and using the fuzzy coefficient k R to evaluate the distance between the two differences, and k R is optimized according to the prediction performance;

[0061] (4.3) Calculate the sample expansion multiple according to the similarity as follows:

[0062]

[0063] In the formula, E χ,k represents the expansion multiple of the k-th cluster when the sample to be predicted belongs to the χ-th cluster, ε represents the segmentation coefficient, represents rounding down;

[0064] (4.4) According to the sample expansion multiple, reconstruct the training sample set assuming that the sample to be predicted belongs to the χ-th cluster as:

[0065]

[0066] In the formula, S k represents the training samples of the k-th cluster;

[0067] (4.5) Build a quantile regression model based on the extreme learning machine according to the reconstructed training sample set, and train the output coefficients of the extreme learning machine assuming that the sample to be predicted belongs to the χ-th cluster;

[0068] (4.6) Determine whether χ is N. If not, set χ = χ + 1 and return to execute step (4.2). If so, the training is completed and the training results of the extreme learning machine models for all clusters are obtained.

[0069] Further, step (4.5) specifically includes the following steps:

[0070] (4.5.1) Construct a quantile regression model based on the extreme learning machine as:

[0071]

[0072] s.t.

[0073]

[0074] 0 ≤ g(x i , w α ) ≤ 1

[0075]

[0076]

[0077] In the formula, T is the sample size of the reconstructed training sample set, α is the rated confidence level, and are variables to be trained, and α respectively represent the upper and lower quantile percentages of the confidence interval, w α , and w α are the output coefficients in the corresponding percentage cases, g(x i , *) is the output value of the extreme learning machine corresponding to the input x i and the output coefficient *, x i , y i are the input and output of the i-th training sample in the training sample set respectively;

[0078] (4.5.2) Train the quantile regression model based on the extreme learning machine according to the reconstructed training sample set to obtain the output coefficients of the extreme learning machine assuming that the sample to be predicted belongs to the χ-th cluster.

[0079] Beneficial effects: Compared with the prior art, the significant advantages of the present invention are as follows: In view of the adverse effects that multiple iterative calculations and outliers of sample distances are likely to have on the accuracy of sample clustering, the present invention first granulates the samples, combines samples with high similarity according to sample differences to form sample granules, reduces the number of hierarchical clustering iterations, and reduces the sensitivity of the clustering calculation process to the distances of abnormal samples. In addition, traditional clustering methods mainly optimize the training samples of deterministic prediction models and lack the direct optimization of the training samples of non-parametric probability prediction. The present invention proposes to adopt an adaptive sample expansion strategy to improve the accuracy of the training sample usage of non-parametric probability prediction methods. The granular clustering method proposed by the present invention can improve the deterministic prediction performance, and in combination with adaptive sample expansion, can provide more accurate training data for the probability prediction model. By comparing with the deterministic prediction method based on traditional clustering strategies and traditional probability prediction methods, it is respectively verified that the deterministic prediction and non-parametric probability prediction models based on the method proposed by the present invention have better prediction accuracy, overall performance and reliability, greatly improving the credibility of photovoltaic power prediction. Brief Description of the Drawings

[0080] Figure 1 It is a schematic flowchart of the method of the present invention;

[0081] Figure 2 It is the ultra-short-term 1-hour prediction interval for sunny days of Dataset 1 of the present invention;

[0082] Figure 3 It is the ultra-short-term 1-hour prediction interval for rainy days of Dataset 1 of the present invention;

[0083] Figure 4 It is the ultra-short-term 1-hour prediction interval for cloudy days of Dataset 1 of the present invention;

[0084] Figure 5 It is the ultra-short-term 1-hour prediction interval for sunny days of Dataset 2 of the present invention;

[0085] Figure 6 It is the ultra-short-term 1-hour prediction interval for rainy days of Dataset 2 of the present invention;

[0086] Figure 7 It is the ultra-short-term 1-hour prediction interval for cloudy days of Dataset 2 of the present invention. Detailed Embodiment

[0087] This embodiment provides a photovoltaic power ultra-short-term non-parametric probability prediction method based on fuzzy sample granules, which is a quantile regression model based on granular clustering and adaptive sample expansion for ultra-short-term prediction of photovoltaic power non-parametric probability intervals. It can be applied to predict loads, wind power / photovoltaic power output, etc. in other ranges and fields at multiple time scales. As Figure 1 shown, the specific steps are as follows:

[0088] (1) Construct and initialize the quantile regression model based on the extreme learning machine, and establish a fuzzy sample distance based on the historical photovoltaic power time series and the change amount of the corresponding irradiance prediction value to form a clustering sample set and a training sample set.

[0089] Specifically, it includes the following steps:

[0090] (1.1) Construct and initialize the hidden layer coefficients and thresholds of the extreme learning machine;

[0091] (1.2) Set the upper and lower quantile percentages and α of the confidence interval of the quantile regression model;

[0092] (1.3) Normalize the historical photovoltaic power time series and the meteorological forecast irradiance data;

[0093] (1.4) Calculate the regional meteorological forecast irradiance prediction value based on the normalized meteorological forecast irradiance data:

[0094]

[0095] In the formula, R i represents the regional meteorological forecast irradiance prediction value at the i-th moment, M represents the number of stations, C u and I u respectively represent the capacity of the u-th station and the meteorological forecast irradiance prediction value at the current moment;

[0096] (1.5) Calculate the change amount ΔR i of the irradiance prediction value:

[0097] ΔR i = R i - R i-1

[0098] In the formula, R i-1 represents the regional meteorological forecast irradiance prediction value at the (i - 1)-th moment calculated according to the formula listed in step (1.4), and the change amount of the irradiance prediction value is used to fuzzily represent the potential change trend of future photovoltaic power;

[0099] (1.6) Construct the training sample set based on the normalized historical photovoltaic power time series as follows:

[0100]

[0101] In the formula, D represents the sample set, x i represents the historical photovoltaic power time series before the i-th moment as the input, y i represents the power at the i-th moment as the output, and the two form the i-th training sample, and S represents the number of samples;

[0102] (1.7) Construct the clustering sample set as follows:

[0103]

[0104] In the formula, St represents the clustering sample set.

[0105] (2) Use the matrix formed by all samples in the clustering sample set as the initial sample particle of the sample particle set, and iterate the process of decomposition and combination of sample particles based on the sample particle difference criterion until all sample particles can no longer be decomposed and synthesized.

[0106] This step specifically includes the following steps:

[0107] (2.1) Use the matrix formed by all samples in the clustering sample set as the initial sample particle and add it to the sample particle set.

[0108] (2.2) Perform the following judgment and decomposition process on each sample particle in the sample particle set in turn: Decompose it into several pairs of sample sub-particle candidates in different ways to form a sample sub-particle candidate pair set, and calculate the difference index value between each candidate pair. If the maximum value of the difference index values of all candidate pairs is greater than or equal to the threshold, decompose the current sample particle into the two sub-particles in the sample sub-particle candidate pair corresponding to the maximum value of the difference index value; otherwise, do not decompose.

[0109] Specifically, it includes the following steps:

[0110] (2.2.1) Select a sample particle from the sample particle set, denoted as sk;

[0111] (2.2.2) Select any variable r from all variables of the sample particle sk as the variable to be analyzed;

[0112] (2.2.3) Sort all samples in the sample particle sk in ascending order according to the value of the variable to be analyzed;

[0113] (2.2.4) Use the value of the variable to be analyzed of the v-th sample after sorting as the decomposition threshold, where v = 1,..., n sk -1, and form a sample sub-particle sk with the samples in the sample particle sk whose values of the variable to be analyzed are greater than the decomposition threshold 1,v,r , and the remaining samples form another sample sub-particle sk 2,v,r , and the two sample sub-particles form a pair of sample sub-particle candidates and are added to the sample sub-particle candidate pair set SK, where n sk represents the number of samples in the sample particle sk;

[0114] (2.2.5) Return to execute steps (2.2.2) to (2.2.4) until all variables of the sample particle sk are traversed, that is, r takes each value in the variable set R in turn, and the sample sub-particle candidate pair set SK = {(sk 1,v,r , sk 2,v,r ) | r ∈ R, v = 1, …, n sk - 1};

[0115] (2.2.6) Calculate the difference index value between each candidate pair in the sample sub-particle candidate pair set, and obtain the maximum value F max . If the maximum value F max is greater than or equal to the threshold, decompose the sample particle sk into the sample sub-particle pair corresponding to the maximum value F max , and replace the sample particle sk in the sample particle set, otherwise do not decompose;

[0116] (2.2.7) Return to execute steps (2.2.1) to (2.2.6) until all sample particles in the sample particle set are traversed, and complete the decomposition of all sample particles in the sample particle set in one cycle calculation.

[0117] Among them, the calculation method of the difference index value is:

[0118]

[0119]

[0120]

[0121]

[0122]

[0123]

[0124] In the formula, Q(e, h) represents the difference index value, e and h represent any two sample particles, F() represents the F-test value, () T represents the transpose of the matrix, A represents the sum of the squares of the sample particle differences, B represents the sum of the cross-product matrices, e p represents the vector corresponding to the p-th sample in e, h q represents the vector corresponding to the q-th sample in h; and respectively represent the average value vectors of e and h, n e and n h are respectively the total number of samples in e and h; d represents the number of variables of the sample;

[0125] The threshold of the difference index value is α F is a preset ratio.

[0126] (2.3) For every two arbitrary sample particles in the sample particle set, calculate the difference index value between these two sample particles. If the minimum value of all difference index values is less than the threshold, then merge the two sample particles corresponding to the minimum value into one sample particle; otherwise, do not merge.

[0127] Specifically, it includes the following steps:

[0128] (2.3.1) Form sample particle pairs from every two arbitrary sample particles in the sample particle set;

[0129] (2.3.2) Calculate the difference index value between the two sample particles of all sample particle pairs;

[0130] (2.3.3) Obtain the minimum value F of the difference index value min , if the minimum value F min is less than the threshold, then merge the two sample particles of the sample particle pair corresponding to the minimum value F in the sample particle set into one sample particle; otherwise, do not merge. min

[0131] (2.4) Return to execute steps (2.2) to (2.3) until all sample particles in the sample particle set cannot be decomposed or merged. Note: Let one return execution be regarded as one loop calculation. In each loop calculation, each sample particle in the decomposition process is judged only once according to the maximum difference value; the merging process is judged only once according to the minimum difference value of the sample particle set.

[0132] (3) Cluster all sample particles in the sample particle set according to hierarchical clustering to obtain N clusters.

[0133] (4) Calculate the sample expansion multiples when the sample to be predicted belongs to different clusters according to the similarity between different clusters, reconstruct the training sample set according to the sample expansion multiples, and then train the quantile regression model based on the extreme learning machine with the reconstructed training sample set to obtain extreme learning machines for different clusters.

[0134] Specifically, it includes the following steps:

[0135] (4.1) Set χ = 1;

[0136] (4.2) Calculate the similarity between the centers of other clusters and the center of the χ-th cluster assuming that the sample to be predicted belongs to the χ-th cluster:

[0137]

[0138] In the formula, Sim χ,kIndicates the similarity between the cluster center where the sample to be predicted is located and the k-th cluster center. Respectively represent the center of the cluster where the sample to be predicted is located The center of the k-th cluster The j-th variable value of, d represents the number of variables of the clustering samples; the center of each cluster is the mean of the samples in each cluster, and the formula for its j-th variable value is: g l,j Represents the j-th variable of the l-th sample in the cluster, T g,* Then represents the number of samples within the *-th cluster; Measures the difference in the power time series of the clustering samples; Measures the difference in the irradiance change trend, Are respectively the d-th variables of each clustering sample, representing ΔR in their respective samples, with the fuzzy coefficient k R Evaluates the distance between the two differences, k R Is optimized according to the prediction performance;

[0139] (4.3) The sample expansion multiple is calculated according to the similarity as follows:

[0140]

[0141] In the formula, E χ,k Represents the expansion multiple of the k-th cluster when the sample to be predicted belongs to the χ-th cluster, ε represents the segmentation coefficient, Represents rounding down;

[0142] (4.4) According to the sample expansion multiple, the training sample set assuming that the sample to be predicted belongs to the χ-th cluster is reconstructed as:

[0143]

[0144] In the formula, S k Represents the training samples of the k-th cluster;

[0145] (4.5) Build a quantile regression model based on the extreme learning machine according to the reconstructed training sample set, and train the output coefficients of the extreme learning machine assuming that the sample to be predicted belongs to the χ-th cluster; specifically, it includes the following steps:

[0146] (4.5.1) Construct a quantile regression model based on the extreme learning machine as:

[0147]

[0148] s.t.

[0149]

[0150] 0 ≤ g(x i , wα ) ≤ 1

[0151]

[0152]

[0153] Wherein, T is the sample size of the reconstructed training sample set, α is the rated confidence level, and are variables to be trained, and α represent the upper and lower quantile percentages of the confidence interval respectively, w α , and w α are the output coefficients under the corresponding percentage conditions respectively, g(x i , *) is the output value of the extreme learning machine corresponding to the input x i and the output coefficient *, x i , y i are the input and output of the i-th training sample in the training sample set respectively;

[0154] (4.5.2) Train the quantile regression model based on the extreme learning machine according to the reconstructed training sample set to obtain the output coefficients of the extreme learning machine assuming that the sample to be predicted belongs to the χ-th cluster.

[0155] (4.6) Determine whether χ is N. If not, set χ = χ + 1 and return to execute step (4.2). If so, the training is completed and the training results of the extreme learning machine models for all clusters are obtained.

[0156] (5) According to the cluster to which the sample to be predicted belongs, input the sample to be predicted into the extreme learning machine of the belonging cluster to predict the probability prediction interval of the photovoltaic power.

[0157] To enable those skilled in the art to better understand the technical solutions described in the present invention and also to verify the effectiveness of the method of the present invention, the following takes the photovoltaic power generation in an actual area as an example for a detailed introduction. First, taking the traditional prediction models based on K-means and hierarchical clustering, and the deterministic prediction model based on information granules (IGNN) as comparisons, it is verified that the proposed sample granule clustering (GC) has good sample processing effects and helps to establish a high-precision deterministic prediction model; secondly, taking the traditional probability prediction methods as comparison references, including the probability prediction method based on clustering theory (CM), the neural network based on information granules (IGNN), the prediction interval based on optimal granules (OGPIs), and the statistical upscaling method (UM), the performance of the proposed probability prediction method (PM) is verified.

[0158] Two groups of data sets are used to verify the performance of the prediction model:

[0159] Dataset 1: The photovoltaic power time series of a certain area in northern China is used to verify the ultra-short-term probability interval prediction performance of the model. The data resolution is 15 minutes. The data from March to June 2019 is used as a case study. According to irradiance, humidity, and cloud cover, typical days are classified into sunny days, rainy days, and cloudy days. For each case study test, 30 days of data are taken. The data of the first 21 days are used for training, and the data of the last 9 days are used for testing.

[0160] Dataset 2: The data from October to February of the Global Energy Forecasting Competition in 2014, with a data resolution of 1 hour. Typical days are screened according to different weather classifications. Each set of data is 90 days, of which 60 days of data are used for training and 30 days of data are used for testing.

[0161] The evaluation indicators for deterministic prediction performance are Root Mean Squared Error (RMSE) and Mean Absolute Error (MAE), which are calculated as follows:

[0162]

[0163]

[0164] In the formula, N p is the total number of samples in the test set; y i and f(x i ) are the measured power value and the predicted value, respectively.

[0165] To evaluate the interval prediction performance, two indicators are generally observed: Average Coverage Error (ACE) and Score. The calculation method of ACE is as follows:

[0166] |ACE| = |PICP - PINC|

[0167] where PICP is the actual interval coverage rate, PINC is the nominal confidence level, and the closer the ACE value is to 0, the better.

[0168] While the coverage error is close to 0, the interval width should be as narrow as possible. The calculation formula is as follows:

[0169]

[0170] where is the width of the i-th predicted value under the interval confidence interval α. The Average Width (AW) represents the sharpness of the prediction interval. and are the upper and lower boundaries of the interval corresponding to the i-th sample, respectively.

[0171] Interval scoring takes into account two factors: interval coverage deviation and interval width. The calculation formula is as follows:

[0172]

[0173]

[0174] is the comprehensive score of the interval, which is a negative value. The closer the value is to 0, the better the comprehensive performance of the prediction interval. For the indicators evaluating the prediction interval, the main thing is to look at the comprehensive score of the interval. The average coverage deviation is generally used as a reference for the interval coverage performance, that is, the coverage degree needs to be close to the rated confidence level to prove that the prediction interval has good reliability.

[0175] Table 1 Comparison of Deterministic Prediction Errors for Dataset 1

[0176]

[0177] Table 2 Comparison of Deterministic Prediction Errors for Dataset 2

[0178]

[0179] Table 3 Comparison of Prediction Interval Performances for Different Confidence Levels of Dataset 1

[0180]

[0181]

[0182] Table 4 Comparison of Prediction Interval Performances for Different Confidence Levels of Dataset 2

[0183]

[0184] Tables 1 and 2 show the effectiveness of the sample particle clustering proposed in the present invention for the deterministic prediction method. Tables 3 and 4 show the effectiveness of the probability prediction based on sample particles and the adaptive sample expansion strategy proposed in the present invention. The optimal indicators of the methods proposed in the tables are bolded and blackened. Based on two datasets, the performance of the deterministic prediction model is observed by combining the root mean square error and the average absolute value. The performance of the prediction interval is observed by combining the interval comprehensive score and the reliability. The comprehensive performance is the decisive indicator for evaluating the prediction interval. At the same time, the accuracy of the interval coverage is combined to comprehensively compare the results of the probability prediction. By comparing the performance of the prediction models by different methods, it is verified that the method of the present invention has good prediction performance.

[0185] Figure 1 The probability prediction method proposed in the present invention is shown. Each module part has been clearly described in the previous text. The whole process is clear and simple, and has strong universality. The effectiveness of the proposed method has been verified by actual example tests.

[0186] Figures 2 to 7 , which reflects the prediction effect with a 90% confidence interval one hour ahead from January to June. The data is used after being normalized by the capacity of the wind farm group. It can be seen that the time series of wind power is complex, and it is difficult to accurately estimate the error distribution by modeling with fixed preset parameters to accurately give the prediction interval. As can be seen from the figure, the interval constructed by the method of the present invention has a good prediction effect. In summary, the present invention can achieve a very short-term prediction interval of wind power and can be used in practical engineering applications.

Claims

1. A very short-term non-parametric probability prediction method for photovoltaic power based on fuzzy sample particles, characterized in that, The steps include: (1) Construct and initialize a quantile regression model based on an extreme learning machine, and establish a fuzzy sample distance based on the historical photovoltaic power time series and the corresponding irradiance prediction value change to form a clustering sample set and a training sample set; (2) The matrix formed by all samples in the clustered sample set is used as the initial sample particle of the sample particle set, and the process of decomposing and merging the sample particles is iterated based on the sample particle difference criterion until all sample particles can no longer be decomposed and synthesized; (3) All sample particles in the sample particle set are clustered according to hierarchical clustering to obtain N clusters; (4) Calculating the sample expansion multiples when the predicted samples belong to different clusters based on the similarities between different clusters, reconstructing the training sample set based on the sample expansion multiples, and then training the quantile regression model based on the extreme learning machine based on the reconstructed training sample set to obtain extreme learning machines of different clusters; (5) According to the cluster to which the sample to be predicted belongs, the sample to be predicted is input into the extreme learning machine of the cluster to which it belongs, and the probability prediction interval of the photovoltaic power is predicted; Wherein, step (2) specifically comprises the following steps: (2.1) The matrix formed by all samples in the clustered sample set is used as the initial sample particle and added to the sample particle set; (2.2) The following judgment and decomposition process is performed on each sample particle in the sample particle set in turn: Decompose it into several pairs of sample sub-particle candidates in different ways to form a set of sample sub-particle candidate pairs, and calculate the difference index value between each candidate pair. If the maximum difference index value of all candidate pairs is greater than or equal to the threshold, decompose the current sample particle into two sub-particles in the sample sub-particle candidate pair corresponding to the maximum difference index value, otherwise do not decompose; (2.3) For every two random sample particles in the sample particle set, calculate the difference index value of the two sample particles. If the minimum value of all difference index values is less than the threshold, merge the two sample particles corresponding to the minimum value into one sample particle, otherwise do not merge; (2.4) Return to execute steps (2.2) to (2.3) until all sample particles in the sample particle set cannot be decomposed or merged.

2. The ultra-short-term non-parametric probability prediction method for photovoltaic power based on fuzzy sample particles according to claim 1, wherein, Step (1) specifically includes the following steps: (1.1) Construct and initialize the hidden layer coefficients and thresholds of the extreme learning machine; (1.2) Set the upper and lower quantile percentages of the confidence interval of the quantile regression model and α; (1.3) Normalize the historical photovoltaic power time series and the weather forecast irradiance data; (1.4) Calculate the regional weather forecast irradiance prediction value based on the normalized weather forecast irradiance data: where R i represents the predicted value of regional meteorological forecast irradiance at the \(i\)-th moment, \(M\) represents the number of power stations, and \(C u and \(I u represent the capacity of the \(u\)-th power station and the predicted value of meteorological forecast irradiance at the current moment, respectively; (1.5) Calculate the change in the predicted irradiance value, ΔR i : ΔR i = R i - R i-1 where R i-1 represents the predicted value of the regional meteorological forecast irradiance at the (i - 1)-th moment calculated according to the formula listed in step (1.4), and the potential change trend of the future photovoltaic power is represented by the fuzzification of the change amount of the irradiance prediction value; (1.6) The training sample set is constructed based on the normalized historical photovoltaic power time series as follows: where D represents the sample set, and x i represents the historical photovoltaic power time series before the i-th moment and serves as the input, and y i represents the power at the i-th moment and serves as the output. The two form the i-th training sample, and S represents the number of samples; (1.7) The clustering sample set is constructed as: In the formula, St represents the cluster sample set; y i Do not participate in the calculation of the distance of clustering samples. The actual calculation variable is x i and ΔR i .

3. A very short-term non-parametric probability prediction method for photovoltaic power based on fuzzy sample particles according to claim 1, characterized in that Step (2.2) specifically includes the following steps: (2.2.1) Select a sample particle from the sample particle set and set it as sk; (2.2.2) Select any variable r from all variables of the sample particle sk as the variable to be analyzed; (2.2.3) Sort all samples in the sample particle sk from small to large according to the value of the variable to be analyzed; (2.2.4) Take the value of the variable to be analyzed of the v-th sample after sorting as the decomposition threshold, where v = 1, …, n sk -1, and form a sample sub-particle sk from the samples within the sample particle sk whose variable values to be analyzed are greater than the decomposition threshold 1,v,r , and form another sample sub-particle sk from the remaining samples 2,v,r , and the two sample sub-particles form a pair of sample sub-particle candidates to be added to the sample sub-particle candidate pair set SK, where n sk represents the number of samples within the sample particle sk; (2.2.5) Return to execute steps (2.2.2) to (2.2.4) until all variables of the sample particle sk are traversed, that is, r takes each value in the variable set R in turn, and the sample sub-particle candidate pair set SK = {(sk 1,v,r , sk 2,v,r ) | r ∈ R, v = 1, …, n sk -1}; (2.2.6) Calculate the difference index values between each candidate pair in the sample sub-particle candidate pair set, and obtain the maximum value F from them. max , if the maximum value F max is greater than or equal to the threshold, decompose the sample particle sk into the sample sub-particle pair corresponding to the maximum value F max , and replace the sample particle sk in the sample particle set, otherwise do not decompose. (2.2.7) Return to execute steps (2.2.1) to (2.2.6) until all sample particles in the sample particle set are traversed, and complete the decomposition of all sample particles in the sample particle set in one cycle calculation.

4. A photovoltaic power ultra-short-term non-parametric probability prediction method based on fuzzy sample particles according to claim 1, characterized in that (2.3) The steps specifically include the following steps: (2.3.1) Form sample particle pairs from any two sample particles in the sample particle set; (2.3.2) Calculate the difference index values between the two sample particles of all sample particle pairs; (2.3.3) Obtain the minimum value F of the difference index min , if the minimum value F min is less than the threshold, then merge the two sample particles of the sample particle pair corresponding to the minimum value F min in the sample particle set into one sample particle; otherwise, do not merge.

5. A non-parametric probability prediction method for ultra-short-term photovoltaic power based on fuzzy sample particles as claimed in claim 1, characterized in that The calculation method of the difference index value is: In the formula, Q(e, h) represents the value of the difference index, e and h represent any two sample particles, F() represents the F-test value, and () T represents the transpose of the matrix, A represents the sum of the squares of the differences of the sample particles, B represents the sum of the cross-product matrices, and e p represents the vector corresponding to the p-th sample in e, and h q represents the vector corresponding to the q-th sample in h; and respectively represent the mean vectors of e and h, and n e and n h are respectively the total number of samples in e and h; d represents the number of variables of the sample; The threshold value of the difference index value is α F which is a preset ratio.

6. The ultra-short-term non-parametric probability prediction method for photovoltaic power based on fuzzy sample particles as claimed in claim 1, wherein (4) The steps specifically include the following steps: (4.1) Set χ = 1; (4.2) Calculate the similarity between the centers of other clusters and the center of the χ-th cluster when assuming that the sample to be predicted belongs to the χ-th cluster: where Sim χ,k represents the similarity between the center of the cluster where the sample to be predicted is located and the center of the k-th cluster, respectively represent the center of the cluster where the sample to be predicted is located the center of the k-th cluster the j-th variable value of, d represents the number of variables of the clustering samples; the center of each cluster is the mean of the samples in each cluster, and the formula for its j-th variable value is: g l,j represents the j-th variable of the l-th sample in the cluster, then represents the number of samples in the *-th cluster; measures the difference in the power time series of the clustering samples; measures the difference in the irradiance change trend, are respectively the d-th variables of each clustering sample, representing ΔR in their respective samples, with the fuzzy coefficient k R evaluates the distance between the two differences, k R is optimized according to the prediction performance; (4.3) Calculate the sample expansion multiple according to the similarity as follows: where E χ,k represents the expansion multiple of the k-th cluster when the sample to be predicted belongs to the χ-th cluster, and ε represents the segmentation coefficient, represents rounding down; (4.4) According to the sample expansion multiple, reconstruct the training sample set assuming that the sample to be predicted belongs to the χ-th cluster as: where S k represents the training samples of the k-th cluster; (4.5) Build a quantile regression model based on the extreme learning machine according to the reconstructed training sample set, and train the output coefficients of the extreme learning machine assuming that the sample to be predicted belongs to the χ-th cluster; (4.6) Determine whether χ is N. If not, set χ = χ + 1 and return to execute step (4.2). If so, the training is completed, and the training results of the extreme learning machine models of all clusters are obtained.

7. A very short-term non-parametric probability prediction method for photovoltaic power based on fuzzy sample particles according to claim 6, characterized in that (4.5) The steps specifically include the following steps: (4.5.1) Construct a quantile regression model based on the extreme learning machine as: 0≤g(x i ,w α )≤1 Where T is the sample size of the reconstructed training sample set, α is the rated confidence level, and are variables to be trained, and α represent the upper and lower quantile percentages of the confidence interval respectively, w α 、 and w α are the output coefficients in the corresponding percentage cases respectively, g(x i ,*) is the output value of the extreme learning machine corresponding to the input x i and the output coefficient *, x i ,y i are the input and output of the i-th training sample in the training sample set respectively; (4.5.2) Train the quantile regression model based on the extreme learning machine according to the reconstructed training sample set to obtain the output coefficients of the extreme learning machine assuming that the sample to be predicted belongs to the χ-th cluster.

Citation Information

Patent Citations

  • Extreme TS fuzzy reasoning method and system based on extreme learning machine

    CN108665070A

  • Wind power ultra-short-term probability prediction method based on conditional quantile regression model

    CN113256018A