Photovoltaic output probability distribution prediction method

By constructing a quantile regression model of neural networks, combining diameter analysis and dimensionality reduction processing, the problem of limited accuracy of existing photovoltaic power prediction methods is solved, and high accuracy prediction and probability distribution analysis of photovoltaic output are achieved, meeting the needs of the power grid scheduling department.

CN119994850APending Publication Date: 2025-05-13CHINA HUANENG GRP CO LTD +2
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411831134.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing photovoltaic power prediction methods have limited accuracy, which is difficult to reflect the uncertainty of photovoltaic output and probability distribution characteristics, and cannot meet the actual needs of the power grid dispatch department.

Method used

By obtaining historical data for preprocessing, combining diameter analysis and dimensionality reduction processing, a quantile regression model of neural network is constructed to predict the value of photovoltaic output under different quantiles, and a probability distribution function of photovoltaic output is established to provide confidence intervals under confidence.

Benefits of technology

It improves the accuracy of photovoltaic power prediction, provides probability distribution information of photovoltaic output, enhances the applicability and flexibility of the prediction model, and promotes the stable operation and optimization management of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119994850A_ABST
    Figure CN119994850A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of photovoltaic power generation power prediction. The invention provides a photovoltaic output probability distribution prediction method, which comprises the following steps: acquiring historical data, and preprocessing the historical data to obtain a preprocessed data set; performing dimension reduction processing on the preprocessed data set to obtain a dimension-reduced data set; constructing an initial prediction model based on a neural network quantile regression model, and training the initial prediction model through the dimension reduction data set to obtain a prediction model; predicting the photovoltaic output value at the to-be-predicted moment through the prediction model to obtain predicted values under a plurality of quantiles; establishing a probability distribution function of photovoltaic output based on the predicted values under a plurality of quantiles; and comparing the obtained photovoltaic output value with an actual photovoltaic output value, evaluating the prediction performance of the prediction model, and optimizing the prediction model according to an evaluation result. The problems that an existing photovoltaic power prediction method is limited in accuracy, and photovoltaic output uncertainty and probability distribution characteristics are difficult to reflect are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of photovoltaic power generation prediction, and in particular to a photovoltaic output probability distribution prediction method. Background Art

[0002] With the transformation of the global energy structure and the rapid development of renewable energy, photovoltaic power generation, as a clean and renewable energy form, has an increasing installed capacity and power generation in the power system. However, the output of photovoltaic power generation is affected by a variety of meteorological factors, such as solar radiation intensity, temperature, wind speed, cloud cover, etc., and has significant volatility and uncertainty. This uncertainty poses a huge challenge to the power generation planning and real-time dispatching of the power system. In order to meet this challenge, the power grid dispatching department requires power stations (including wind power and photovoltaic power) to report power forecast results, including short-term and ultra-short-term power forecasts, in order to better manage and optimize the operation of the power system. The accuracy of power forecasting is crucial to the stable operation and economic benefits of the power system. However, most of the existing power forecasting methods are based on simple regression analysis of historical data and meteorological factors, with limited prediction accuracy, which is difficult to meet the actual needs of the power grid dispatching department.

[0003] In order to improve the accuracy of photovoltaic power prediction, some researchers have begun to try to use advanced mathematical and machine learning methods for prediction. However, when these methods are applied to photovoltaic power prediction, they often face problems such as high data dimensions and complex nonlinear relationships, resulting in unsatisfactory prediction results. In addition, most existing prediction methods can only give deterministic prediction values ​​and cannot reflect the uncertainty and probability distribution characteristics of photovoltaic output, which is insufficient for risk assessment and decision making of power systems. Summary of the invention

[0004] The purpose of the present invention is to provide a photovoltaic output probability distribution prediction method, aiming to solve the problem that the existing photovoltaic power prediction method has limited accuracy and is difficult to reflect the uncertainty and probability distribution characteristics of photovoltaic output.

[0005] The present invention is achieved through the following technical solutions:

[0006] A photovoltaic output probability distribution prediction method comprises the following steps:

[0007] Acquire historical data, preprocess the historical data, and obtain a preprocessed data set;

[0008] Perform path analysis on the correlation between meteorological factors and photovoltaic output values ​​in the preprocessed data set, perform dimensionality reduction processing on the preprocessed data set, and obtain a reduced dimensionality data set;

[0009] Based on the neural network quantile regression model, an initial prediction model is constructed, and the initial prediction model is trained through a dimension reduction data set to obtain a prediction model;

[0010] The photovoltaic output value at the time to be predicted is predicted by the prediction model to obtain the prediction values ​​under several quantiles;

[0011] Based on the predicted values ​​at several quantiles, the probability distribution function of photovoltaic output is established, and the confidence interval at the corresponding confidence level is given;

[0012] The photovoltaic output value obtained from the prediction results is compared with the actual photovoltaic output value to evaluate the prediction performance of the prediction model, and the parameters and structure of the prediction model are optimized based on the evaluation results.

[0013] Optionally, the specific process of acquiring historical data, preprocessing the historical data, and obtaining a preprocessed data set is:

[0014] A historical data set containing meteorological parameters and corresponding photovoltaic output values ​​at historical moments is collected, and the historical data set is normalized to obtain a preprocessed data set.

[0015] Optionally, the specific process of performing path analysis on the correlation between the meteorological factors and the photovoltaic output values ​​in the preprocessed data set and performing dimensionality reduction processing on the preprocessed data set to obtain the dimensionality reduction data set is:

[0016] Calculate the direct path coefficient and indirect path coefficient between each meteorological factor and the photovoltaic output in the preprocessed data set; evaluate the direct impact of each meteorological factor on the photovoltaic output through the direct path coefficient and the indirect impact of other meteorological factors on the photovoltaic output;

[0017] Based on the evaluation results, the key meteorological factors that have a significant impact on the photovoltaic output are screened out, and the minor meteorological factors are eliminated to obtain a reduced dimension data set.

[0018] Optionally, the dimension reduction data set is divided into a training set and a test set; the daily category is set according to the daily total irradiance in the historical data; the training set and the test set are divided according to the daily category to which each day belongs, and the training set and the test set of the corresponding daily category are obtained.

[0019] Optionally, the specific process of constructing an initial prediction model based on a neural network quantile regression model and training the initial prediction model through a dimension reduction data set to obtain the prediction model is as follows:

[0020] According to the structure of the neural network quantile regression model, set the number of nodes in the input layer, hidden layer and output layer, as well as the corresponding activation function;

[0021] According to the day category to which the time to be predicted belongs, the k-nearest neighbor algorithm is used to select the historical data corresponding to the meteorological factors at the time to be predicted from the training set of the corresponding day category as the training sample set;

[0022] Initialize the connection weight vector and penalty factor of the neural network quantile regression model, use the meteorological parameters of the screened training sample set and the corresponding photovoltaic output value as input and output respectively, and obtain the initial prediction model;

[0023] The initial prediction model is trained through the training sample set to obtain the prediction model.

[0024] Optionally, according to the day category to which the time to be predicted belongs, the k-nearest neighbor algorithm is used to filter out historical data corresponding to the meteorological factors of the time to be predicted from the test set of the corresponding day category as a test sample set; the prediction model is tested using the test sample set.

[0025] Optionally, the specific process of predicting the photovoltaic output value at the time to be predicted by using the prediction model to obtain the predicted values ​​under several quantiles is:

[0026] The meteorological parameters at the time to be predicted are input into the trained prediction model. The characteristics of the neural network quantile regression model are used to calculate the corresponding prediction values ​​for several preset quantiles. Each prediction value represents the photovoltaic output prediction at the corresponding quantile point.

[0027] Optionally, the specific process of establishing the probability distribution function of photovoltaic output based on the predicted values ​​under several quantiles and providing the confidence interval under the corresponding confidence level is:

[0028] Through the kernel density estimation method, the probability density function of photovoltaic output is estimated with the predicted values ​​under several quantiles as sample points; based on the estimated probability density function, the probability distribution function of photovoltaic output is constructed; based on the probability distribution function, the quantiles under the corresponding confidence level are determined, and the corresponding confidence intervals are calculated; the confidence intervals are used to reflect the uncertainty range of the photovoltaic output prediction value under the corresponding confidence level.

[0029] Optionally, the specific process of comparing the photovoltaic output value obtained by the prediction result with the actual photovoltaic output value to evaluate the prediction performance of the prediction model is:

[0030] Collect actual PV output data that matches the predicted values ​​output by the prediction model in time and space;

[0031] Calculate the error between the predicted value and the actual photovoltaic output value to obtain the corresponding error evaluation index;

[0032] Quantify the prediction accuracy and precision of the prediction model based on the calculated error evaluation index;

[0033] Compare the error evaluation indicators under different prediction models or different parameter settings to evaluate the performance of the prediction model.

[0034] Optionally, the specific process of optimizing the parameters and structure of the prediction model according to the evaluation results is:

[0035] Identify flaws and potential improvements in the prediction model based on error evaluation metrics and prediction model performance;

[0036] Adjust and optimize the parameters of the neural network quantile regression model to determine the optimal parameter combination;

[0037] The structure of the prediction model is optimized by adjusting the node configuration of the input layer, hidden layer and output layer of the prediction model and selecting the corresponding activation function based on the optimal parameter combination;

[0038] During the optimization process, the prediction performance of the model is continuously monitored, and through iterative optimization, the prediction model with the best performance is obtained for subsequent photovoltaic output probability distribution prediction tasks.

[0039] The technical solution of the present invention has at least the following advantages and beneficial effects:

[0040] Improve the accuracy of photovoltaic power prediction: By introducing path analysis to conduct an in-depth analysis of the correlation between meteorological factors and photovoltaic output values, and combining dimensionality reduction processing technology, the data dimension is effectively reduced, reducing the interference of redundant information on the prediction results, thereby improving the accuracy of the prediction model; at the same time, the prediction model is constructed using the neural network quantile regression model, which can capture the complex nonlinear relationship between photovoltaic output and meteorological factors, further improving the prediction accuracy.

[0041] Providing probability distribution information of PV output: It can predict the values ​​of PV output at different quantiles based on the quantile regression model, and then construct the probability distribution function of PV output, which not only reflects the uncertainty of PV output, but also provides more comprehensive information support for risk assessment and decision-making of power systems.

[0042] Enhance the applicability and flexibility of the prediction model: The prediction model is constructed through a neural network, which has strong adaptability and generalization capabilities and can cope with changes in different meteorological conditions and photovoltaic power station characteristics; at the same time, the prediction model can be evaluated and optimized based on the comparison results between the actual photovoltaic output value and the predicted value, continuously improving the prediction performance and ensuring the reliability and stability of the prediction results.

[0043] Promote stable operation and optimized management of the power system: By providing accurate PV output prediction results and probability distribution information, it helps the grid dispatching department to better arrange power generation plans and real-time dispatch, and reduce the risks caused by the uncertainty of PV output; at the same time, it also helps to improve the overall operating efficiency and economic benefits of the power system, promote the transformation of the energy structure and the sustainable development of renewable energy. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 Schematic diagram of a photovoltaic output probability distribution prediction method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0045] The following is a specific implementation method in conjunction with the drawings.

[0046] Example 1

[0047] Reference Figure 1 , a photovoltaic output probability distribution prediction method, comprising the steps of:

[0048] Step 1: Obtain historical data, preprocess the historical data, and obtain a preprocessed data set.

[0049] In this embodiment, the specific process is:

[0050] Collect historical data sets containing meteorological parameters and corresponding photovoltaic output values ​​at historical moments, and normalize the historical data sets to obtain preprocessed data sets. Data sets containing meteorological parameters and corresponding photovoltaic output values ​​at historical moments can be collected from historical databases. Meteorological parameters include temperature, humidity, wind speed, wind direction, irradiance, etc., and photovoltaic output refers to the actual power generation of the photovoltaic power station at the corresponding moment. Ensure that the collected data covers a sufficiently long period of time to reflect the changes in photovoltaic output under different seasons and weather conditions. Check whether there are missing values ​​or outliers in the data set. For missing values, interpolation methods (such as linear interpolation, nearest neighbor interpolation, etc.) can be used to fill them, or records with missing values ​​can be deleted according to the data distribution. For outliers, they need to be identified and processed according to the actual situation, such as deleting or replacing them with reasonable values. Ensure that the timestamps between meteorological parameters and photovoltaic output values ​​are consistent to avoid data dislocation or time mismatch problems. Since the units and value ranges of different meteorological parameters may be different, in order to eliminate the impact of such differences on model training, the data needs to be normalized. Min-Max normalization and Z-score normalization can be used. Min-Max normalization scales the data to between 0 and 1, and Z-score normalization performs normalization based on the mean and standard deviation of the data. For each meteorological parameter and PV output value, its maximum and minimum values ​​(or mean and standard deviation) are calculated and then converted according to the selected normalization method.

[0051] Step 2: Perform path analysis on the correlation between meteorological factors and photovoltaic output values ​​in the preprocessed data set, and perform dimensionality reduction processing on the preprocessed data set to obtain a reduced dimensionality data set.

[0052] In this embodiment, the specific process is:

[0053] The direct path coefficient and indirect path coefficient between each meteorological factor and the photovoltaic output value in the preprocessed data set are calculated; the direct path coefficient can be calculated by linear regression or multiple regression method, taking each meteorological factor as the independent variable and the corresponding photovoltaic output value as the dependent variable; the indirect path coefficient can be obtained by analyzing the correlation between the independent variables and calculating the regression coefficient of each meteorological factor on other meteorological factors; the direct impact of each meteorological factor on the photovoltaic output value and the indirect impact of other meteorological factors on the photovoltaic output value are evaluated by the direct path coefficient and the indirect impact of other meteorological factors on the photovoltaic output value are comprehensively considered, that is, according to the direct path coefficient, the direct impact of each meteorological factor on the photovoltaic output value is evaluated, and according to the indirect path coefficient, the indirect impact of other meteorological factors on the photovoltaic output value is comprehensively considered, so as to evaluate the comprehensive impact of each meteorological factor.

[0054] Based on the evaluation results, the key meteorological factors that have a significant impact on the photovoltaic output are screened out, and the secondary meteorological factors are eliminated to obtain a reduced dimension data set. According to the comprehensive evaluation results of the direct path coefficient and the indirect path coefficient, a threshold can be set to screen the key meteorological factors that have a significant impact on the photovoltaic output; the direct and indirect effects of each meteorological factor are compared with the set threshold, and the meteorological factors that exceed the threshold are retained, and the secondary meteorological factors are eliminated; according to the screening results, a reduced dimension data set is constructed. Cross-validation and other methods can be used to verify whether the reduced dimension data set still retains enough information to accurately predict the photovoltaic output; if the verification results show that the prediction performance is poor, the screening threshold can be adjusted and the key meteorological factors can be screened again.

[0055] Step 3: Based on the neural network quantile regression model, an initial prediction model is constructed, and the initial prediction model is trained through the dimensionality reduction data set to obtain a prediction model.

[0056] In this embodiment, the dimension reduction data set is divided into a training set and a test set. The training set accounts for most of the data, such as 70%-80%, and is used to train the model; the test set accounts for the remaining data, such as 20%-30%, and is used to verify the performance of the model; according to the daily total irradiance in the historical data, the daily category is set, and the daily category includes sunny days and rainy days, which helps the model to better learn the photovoltaic output characteristics under different weather conditions; according to the daily category to which each day belongs, the training set and the test set are divided respectively, and the training set and the test set of the corresponding daily category are obtained.

[0057] According to the structure of the neural network quantile regression model, the number of nodes in the input layer, hidden layer, and output layer, as well as the corresponding activation function, are set; the number of nodes in the input layer matches the number of meteorological factors after dimensionality reduction. Select appropriate activation functions, such as ReLU, Sigmo id, etc., to improve the nonlinear expression ability of the model. For quantile regression, the output layer is set to multiple nodes, each node corresponding to a predicted value under a quantile.

[0058] According to the day category to which the time to be predicted belongs, the distance between samples is calculated by the k-nearest neighbor algorithm to determine the similarity, and the historical data corresponding to the meteorological factors at the time to be predicted is selected from the training set of the corresponding day category as the training sample set;

[0059] Initialize the connection weight vector and penalty factor of the neural network quantile regression model. The connection weight vector determines the connection strength between neurons, and the penalty factor is used to control the complexity of the model to prevent overfitting. Take the meteorological parameters of the selected training sample set and the corresponding photovoltaic output value as input and output respectively to obtain the initial prediction model; that is, take the meteorological parameters of the selected training sample set as input and the corresponding photovoltaic output value as output;

[0060] The initial prediction model is trained with the training sample set to obtain the prediction model. During the training process, the prediction error is minimized by adjusting the connection weight vector and the penalty factor. The goal of quantile regression is to learn a function that can output the prediction value at different quantiles. Therefore, during the training process, each quantile needs to be trained separately.

[0061] In this embodiment, according to the day category to which the time to be predicted belongs, the k-nearest neighbor algorithm is used to select historical data corresponding to the meteorological factors at the time to be predicted from the test set of the corresponding day category as the test sample set; the prediction model is tested by the test sample set. The performance of the model is evaluated by comparing the error between the predicted value and the actual photovoltaic output value.

[0062] Step 4: Use the prediction model to predict the photovoltaic output value at the prediction time and obtain the prediction values ​​under several quantiles.

[0063] In this embodiment, the specific process is:

[0064] The meteorological parameters at the time to be predicted are input into the trained prediction model. Before prediction, several quantiles need to be preset, which will be used to generate prediction values ​​at different confidence levels. For example, quantiles such as 0.1, 0.25, 0.5 (median), 0.75 and 0.9 can be selected. The corresponding prediction values ​​are calculated for several preset quantiles by using the characteristics of the neural network quantile regression model. Each prediction value represents the prediction of photovoltaic output at the corresponding quantile point, reflecting the possible values ​​of photovoltaic output at different confidence levels.

[0065] Step 5: Based on the predicted values ​​at several quantiles, establish the probability distribution function of photovoltaic output and give the confidence interval at the corresponding confidence level.

[0066] In this embodiment, the specific process is:

[0067] Through the kernel density estimation method, the probability density function of photovoltaic output is estimated with the predicted values ​​under several quantiles as sample points; based on the estimated probability density function, the probability distribution function of photovoltaic output is constructed. This function describes the probability distribution of photovoltaic output at different values ​​and can be used for subsequent probability analysis and risk assessment; based on the probability distribution function, the quantiles under the corresponding confidence level are determined, and the corresponding confidence intervals are calculated; the confidence interval is used to reflect the uncertainty range of the photovoltaic output prediction value under the corresponding confidence level. According to the required confidence level (such as 90%, 95%, etc.), the corresponding quantiles can be found in the probability distribution function, and then the difference between the two quantiles is calculated to obtain the confidence interval.

[0068] Step 6: Compare the photovoltaic output value obtained from the prediction results with the actual photovoltaic output value, evaluate the prediction performance of the prediction model, and optimize the parameters and structure of the prediction model based on the evaluation results.

[0069] In this embodiment, the specific process is:

[0070] Collect actual PV output data that matches the predicted values ​​output by the prediction model in time and space; these data should come from the same PV power station and be recorded in the same time period.

[0071] The error between the predicted value and the actual photovoltaic output value is calculated to obtain the corresponding error evaluation index; the error evaluation index includes mean square error (MSE), root mean square error (RMSE), mean absolute error (MAE), etc.

[0072] According to the calculated error evaluation index, the prediction accuracy and precision of the prediction model are quantified; the error evaluation indexes under different prediction models or different parameter settings are compared to evaluate the performance of the prediction model.

[0073] Based on the error evaluation indicators and the performance of the prediction model, identify the defects and potential improvement space in the prediction model, including analyzing the prediction performance of the model in different time periods and weather conditions, as well as the model's ability to handle outliers.

[0074] The parameters of the neural network quantile regression model are adjusted and optimized to determine the optimal parameter combination; the optimal parameter combination can be determined through methods such as cross-validation to improve the prediction performance of the model.

[0075] By adjusting the node configuration of the input layer, hidden layer and output layer of the prediction model and selecting the corresponding activation function based on the optimal parameter combination, the structure of the prediction model is optimized to further improve the nonlinear expression ability and prediction accuracy of the model.

[0076] During the optimization process, the prediction performance of the model is continuously monitored. Through iterative optimization, that is, by continuously iteratively adjusting the parameters and structure, after each iteration, the performance of the model is re-evaluated and adjusted according to the evaluation results. The prediction model is gradually optimized until the performance is optimal. The prediction model with the best performance is obtained for the subsequent photovoltaic output probability distribution prediction task. The optimized prediction model is applied to the subsequent photovoltaic output probability distribution prediction task. Ensure that the model can accurately predict the photovoltaic output value at different confidence levels and provide strong support for the operation management, energy scheduling and risk assessment of photovoltaic power stations.

[0077] Example 2

[0078] Based on Example 1, in this embodiment, the k-nearest neighbor algorithm is used to screen out data similar to the meteorological factors at the time to be predicted as a training set, and the meteorological parameters are used as input to use the neural network quantile regression model for prediction to obtain the quantiles at different quantile points. Finally, the probability distribution of photovoltaic output is constructed using kernel density approximation. When the meteorological factors at different times are similar, the corresponding photovoltaic output values ​​are close. There is a certain correlation between photovoltaic outputs at similar times on different days. Similar moments refer to moments when the time values ​​and meteorological parameters of the two moments are similar. And in the probability distribution estimation, the selection effect of the sample data has a great influence on the accuracy of the estimated probability distribution. The higher the similarity between the sample data and the variable to be estimated, the closer the probability distribution obtained by statistical theory is to the actual distribution.

[0079] In this embodiment, the k-nearest neighbor algorithm is used to analyze the similarity of sample data. The k-nearest neighbor algorithm is a simple and practical data mining algorithm that measures the similarity of samples by distance. The shorter the distance, the higher the similarity of the samples. Using Euclidean distance as the metric, for n m-dimensional sample sets X = {x1, x2, ..., x n}, the distance between two samples is defined as follows:

[0080]

[0081] Among them, D[x i ,x j ] represents the Euclidean distance between samples and; x i and x j Respectively represent two m-dimensional sample points; represents the dimension of the sample, that is, the number of meteorological factors; and Respectively represent samples x i and x j The value in the kth dimension.

[0082] The path analysis method is used to study the correlation between various meteorological factors and photovoltaic output values. At the same time, it can effectively solve the problem of indirect influence of independent variables on dependent variables and screen out the main variables as the input of the model. The direct path coefficient represents the direct influence of the independent variable itself on the dependent variable, and the indirect path coefficient represents the influence of the independent variable on the dependent variable indirectly through other variables. After comprehensively considering the interaction between variables, the simple correlation coefficient of the independent variable to the dependent variable can be obtained. For multiple regression analysis, there are independent variables x1, x2,…, x l and dependent variable y, they all have n sets of observations. From the path analysis theory, we can deduce that the independent variable x m The direct path coefficient expression for the dependent variable y is shown in equation (2):

[0083]

[0084] in, Represents the independent variable x m Direct path coefficient for dependent variable y; b m Represents the independent variable x m The partial regression coefficient in the multiple regression model is the value of x when other independent variables remain unchanged. m For every unit increase, the average change in y is b m Units; x mi Represents x m The i-th observation value of ; Represents x m The sample mean of y i represents the i-th observation value of y; represents the sample mean of y;

[0085] Partial regression coefficient b m The expression of is shown in formula (3):

[0086]

[0087] Independent variable x m Through the variable xm+1 The indirect path coefficient expression for the dependent variable y is shown in formula (4):

[0088]

[0089] in, Represents the independent variable x m Through the variable x m+1 The indirect path coefficient for the dependent variable y; Represents x m With x m+1 The correlation coefficient of Represents the independent variable x m+1 The direct path coefficient for the dependent variable y.

[0090] Correlation coefficient The expression of is shown in formula (5):

[0091]

[0092] According to the definition of regression coefficient and path coefficient, and the principle of least squares method, we can get x m The simple correlation coefficient with y is expressed as follows:

[0093]

[0094] in, Represents the independent variable x m Simple correlation coefficient with the dependent variable y; Represents the independent variable x m Through other independent variables x j The sum of the indirect effects of (where j≠m) on the dependent variable y. Each Represents x m By x j The indirect path coefficient for y.

[0095] The neural network quantile regression model is adopted, and the expression is shown in the following formula (7):

[0096] q Y (τ∣X)=f[X,V(τ),W(τ)](7)

[0097] Among them, Q Y(τ|X) represents the quantile of the photovoltaic output Y at the quantile τ under the condition of a given meteorological factor (or independent variable) X; f[X,V(τ),W(τ)] represents a neural network function, which takes the meteorological factor X as input and outputs the predicted value of the photovoltaic output Y at the quantile τ; X represents a vector containing several meteorological factors, such as temperature, humidity, wind speed, radiation, etc., which are used as input variables to predict the photovoltaic output; V(τ)={v ij (τ)} i=1,2,…,s;j=1,2,…,t , V(τ) represents the connection weight vector from the input layer to the hidden layer; W(τ) = {w ij (τ)} i=1,2,…,s;j=1,2,…,t , W(τ) represents the connection weight vector between the hidden layer and the output layer; τ represents the quantile, that is, a specific position in the probability distribution. For example, τ = 0.5 represents the median, and τ = 0.9 represents the 90% quantile. By changing the value of τ, the predicted value of photovoltaic output at different probability levels can be obtained.

[0098] Expanding formula (7) yields the following formula (8):

[0099]

[0100] Among them, w jk (τ) represents the connection weight from the jth node in the hidden layer to the kth node in the output layer; g1 and g2 represent the activation functions of the hidden layer and the output layer, respectively; represents the weighted sum of all nodes in the input layer to the jth node in the hidden layer; the estimation of the weight parameters V(τ) and W(τ) of the neural network quantile regression model can be transformed into the following formula (9), which is solved by minimizing the objective function. Formula (9) is as follows:

[0101]

[0102] in, Indicates that when the actual value Y i Greater than or equal to the predicted value f(X i ,V,W), the loss function is τ multiplied by the absolute value of the prediction error; Indicates that when the actual value Y i Less than the predicted value f(X i ,V,W), the loss function is (1-τ) multiplied by the absolute value of the prediction error; λ1 and λ2 represent penalty factors, which are used to balance the complexity of the model and the prediction error. The optimal penalty factor and the number of hidden layer nodes can be determined by the cross-validation method; Represents the connection weight v from the input layer to the hidden layer ij Perform L2 regularization to prevent overfitting of the model; represents the connection weight w from the hidden layer to the output layerj L1 regularization is also used to prevent overfitting. By minimizing the above objective function, the optimal connection weights V(τ) and W(τ) can be solved to minimize the prediction error of the model at different quantiles while avoiding overfitting.

Claims

1. A photovoltaic output probability distribution prediction method, characterized in that: Includes steps: Acquire historical data, preprocess the historical data, and obtain a preprocessed data set; Perform path analysis on the correlation between meteorological factors and photovoltaic output values ​​in the preprocessed data set, perform dimensionality reduction processing on the preprocessed data set, and obtain a reduced dimensionality data set; Based on the neural network quantile regression model, an initial prediction model is constructed, and the initial prediction model is trained through a dimension reduction data set to obtain a prediction model; The photovoltaic output value at the time to be predicted is predicted by the prediction model to obtain the prediction values ​​under several quantiles; Based on the predicted values ​​at several quantiles, the probability distribution function of photovoltaic output is established, and the confidence interval at the corresponding confidence level is given; The photovoltaic output value obtained from the prediction results is compared with the actual photovoltaic output value to evaluate the prediction performance of the prediction model, and the parameters and structure of the prediction model are optimized based on the evaluation results.

2. The photovoltaic output probability distribution prediction method according to claim 1, characterized in that: The specific process of obtaining historical data, preprocessing the historical data, and obtaining the preprocessed data set is as follows: A historical data set containing meteorological parameters and corresponding photovoltaic output values ​​at historical moments is collected, and the historical data set is normalized to obtain a preprocessed data set.

3. The photovoltaic output probability distribution prediction method according to claim 1, characterized in that: The specific process of performing path analysis on the correlation between the meteorological factors and the photovoltaic output value in the preprocessed data set and performing dimensionality reduction processing on the preprocessed data set to obtain the dimensionality reduction data set is as follows: Calculate the direct path coefficient and indirect path coefficient between each meteorological factor and the photovoltaic output in the preprocessed data set; evaluate the direct impact of each meteorological factor on the photovoltaic output through the direct path coefficient and the indirect impact of other meteorological factors on the photovoltaic output; Based on the evaluation results, the key meteorological factors that have a significant impact on the photovoltaic output are screened out, and the minor meteorological factors are eliminated to obtain a reduced dimension data set.

4. The photovoltaic output probability distribution prediction method according to claim 1, characterized in that: The dimension reduction data set is divided into a training set and a test set; the daily category is set according to the daily total irradiance in the historical data; the training set and the test set are divided according to the daily category to which each day belongs, and the training set and the test set of the corresponding daily category are obtained.

5. The photovoltaic output probability distribution prediction method according to claim 4, characterized in that: The specific process of constructing an initial prediction model based on the neural network quantile regression model and training the initial prediction model through a dimension reduction data set to obtain the prediction model is as follows: According to the structure of the neural network quantile regression model, set the number of nodes in the input layer, hidden layer and output layer, as well as the corresponding activation function; According to the day category to which the time to be predicted belongs, the k-nearest neighbor algorithm is used to select the historical data corresponding to the meteorological factors at the time to be predicted from the training set of the corresponding day category as the training sample set; Initialize the connection weight vector and penalty factor of the neural network quantile regression model, use the meteorological parameters of the screened training sample set and the corresponding photovoltaic output value as input and output respectively, and obtain the initial prediction model; The initial prediction model is trained through the training sample set to obtain the prediction model.

6. The photovoltaic output probability distribution prediction method according to claim 5, characterized in that: According to the day category to which the time to be predicted belongs, the historical data corresponding to the meteorological factors at the time to be predicted is selected from the test set of the corresponding day category by using the k-nearest neighbor algorithm as the test sample set; The prediction model is tested using a test sample set.

7. The photovoltaic output probability distribution prediction method according to claim 1, characterized in that: The specific process of predicting the photovoltaic output value at the time to be predicted by the prediction model to obtain the predicted values ​​under several quantiles is as follows: The meteorological parameters at the time to be predicted are input into the trained prediction model. The characteristics of the neural network quantile regression model are used to calculate the corresponding prediction values ​​for several preset quantiles. Each prediction value represents the photovoltaic output prediction at the corresponding quantile point.

8. The photovoltaic output probability distribution prediction method according to claim 1, characterized in that: The specific process of establishing the probability distribution function of photovoltaic output based on the predicted values ​​under several quantiles and giving the confidence interval under the corresponding confidence level is as follows: Through the kernel density estimation method, the probability density function of photovoltaic output is estimated with the predicted values ​​under several quantiles as sample points; based on the estimated probability density function, the probability distribution function of photovoltaic output is constructed; based on the probability distribution function, the quantiles under the corresponding confidence level are determined, and the corresponding confidence intervals are calculated; the confidence intervals are used to reflect the uncertainty range of the photovoltaic output prediction value under the corresponding confidence level.

9. The photovoltaic output probability distribution prediction method according to claim 1, characterized in that: The specific process of comparing the photovoltaic output value obtained by the prediction result with the actual photovoltaic output value and evaluating the prediction performance of the prediction model is as follows: Collect actual PV output data that matches the predicted values ​​output by the prediction model in time and space; Calculate the error between the predicted value and the actual photovoltaic output value to obtain the corresponding error evaluation index; Quantify the prediction accuracy and precision of the prediction model based on the calculated error evaluation index; Compare the error evaluation indicators under different prediction models or different parameter settings to evaluate the performance of the prediction model.

10. The photovoltaic output probability distribution prediction method according to claim 9, characterized in that: The specific process of optimizing the parameters and structure of the prediction model according to the evaluation results is as follows: Identify flaws and potential improvements in the prediction model based on error evaluation metrics and prediction model performance; Adjust and optimize the parameters of the neural network quantile regression model to determine the optimal parameter combination; The structure of the prediction model is optimized by adjusting the node configuration of the input layer, hidden layer and output layer of the prediction model and selecting the corresponding activation function based on the optimal parameter combination; During the optimization process, the prediction performance of the model is continuously monitored, and through iterative optimization, the prediction model with the best performance is obtained for subsequent photovoltaic output probability distribution prediction tasks.

Citation Information

Cited By

  • Dynamic evaluation method for uncertainty of new energy output power

    CN120930889A

  • Photovoltaic power uncertainty prediction method and system

    CN121688870A

  • Photovoltaic power uncertainty prediction method and system

    CN121688870B