A method and system for predicting new energy power generation errors considering meteorological conditions

CN121390398BActive Publication Date: 2026-09-01STATE GRID JIANGSU ELECTRIC POWER CO LTD YANGZHONG POWER SUPPLY BRANCH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511415381.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2026-09-01
Estimated Expiration
2045-09-30

AI Technical Summary

Technical Problem

另一些方法虽然尝试对误差进行概率建模,但大多依赖于假设误差服从单一的高斯分布或其他参数化分布,并通过拟合这些分布的参数来描述不确定性

Benefits of technology

[0053]1、本发明通过对原始数据进行系统的预处理,包括缺失数据填充和异常数据修正,显著提升了输入数据的质量和完整性,为后续概率模型的建立提供了坚实可靠的数据基础,从而从源头上保证了整个建模方法的准确性和鲁棒性,避免了因数据质量问题导致的模型失真;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121390398B_ABST
    Figure CN121390398B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for analyzing prediction errors in new energy power generation considering meteorological conditions, belonging to the field of computer data processing and prediction technology. It includes acquiring historical data on new energy power generation, prediction error data, and meteorological data, preprocessing them to generate an initial dataset, calculating the statistics of the prediction error, and, in conjunction with the meteorological data in the initial dataset, using kernel density estimation to estimate the joint probability density of the meteorological data and prediction error data to generate a joint probability density distribution. Then, using Bayes' theorem, it calculates the conditional probability distribution of the prediction error under preset meteorological conditions to generate a conditional probability model. Based on climate and location, it performs multidimensional conditional probability modeling on the conditional probability model to generate a multidimensional conditional probability model. This invention combines kernel density estimation with Bayesian theory and incorporates multidimensional spatiotemporal factors for refined modeling, enabling accurate quantification of the uncertainty in new energy power generation prediction and achieving forward-looking risk warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer data processing and prediction technology, and in particular to a method and system for predicting errors in new energy power generation that takes meteorological conditions into account. Background Technology

[0002] New energy power generation, especially photovoltaic and wind power, is an important component of the future energy structure. However, its output exhibits significant fluctuations and intermittency, primarily influenced by meteorological conditions. Accurate forecasting of new energy power generation is crucial for ensuring the safe and stable operation of the power system. Forecasting error refers to the deviation between predicted and actual power generation. Modeling and analyzing the characteristics of forecasting error, particularly its correlation with meteorological factors, is a key technical step in quantifying forecasting uncertainty and improving the grid's capacity to accommodate new energy sources. This typically requires advanced computer data processing and modeling methods.

[0003] In existing technologies, the analysis of prediction errors in new energy power generation typically focuses on improving deterministic prediction methods or using simple statistical methods to describe the error distribution. Some methods directly predict power using algorithms such as neural networks and support vector machines, while the analysis of prediction errors is limited to calculating overall indicators such as root mean square error and mean absolute error. Other methods attempt to probabilistically model the errors, but most rely on assuming that the errors follow a single Gaussian distribution or other parameterized distributions, and describe the uncertainty by fitting the parameters of these distributions.

[0004] However, simple statistical indicators cannot reveal the dynamic characteristics of prediction errors as they change with external conditions such as weather. Secondly, the imposed parametric distribution assumptions often do not match the complex error distributions in reality, especially under extreme weather conditions, where prediction errors often exhibit non-Gaussian characteristics such as skewness, fat tails, or multimodality, making it difficult for traditional models to accurately capture risks. Furthermore, most existing models neglect the systematic impact of seasonal climate change and micro-meteorological differences across geographical locations on the distribution of prediction errors, resulting in insufficient generalization ability and scenario adaptability. Summary of the Invention

[0005] Purpose of the invention: This invention provides a method for analyzing the prediction error of new energy power generation considering meteorological conditions. It adopts a technical approach that combines kernel density estimation with Bayesian theory and incorporates multidimensional spatiotemporal factors for refined modeling. This method can accurately quantify the uncertainty of new energy power generation prediction and achieve forward-looking risk warning. This invention also provides a system for analyzing the prediction error of new energy power generation considering meteorological conditions.

[0006] Technical Solution: First, this invention provides a method for predicting errors in new energy power generation considering meteorological conditions. This method includes:

[0007] An initial dataset is constructed based on historical data of new energy power generation, prediction error data and meteorological data, and error statistics are generated using the initial dataset to construct a statistical data set.

[0008] The initial dataset is analyzed and processed using a preset Gaussian kernel function to generate an initial density estimate. The joint probability density distribution is obtained based on the set of statistical parameters and the initial density estimate.

[0009] The prior probability distribution of the prediction error and the likelihood probability distribution of the preset meteorological conditions are extracted from the joint probability density distribution to generate probability components, and a multidimensional conditional probability model is obtained based on the generated probability components.

[0010] The prediction confidence interval is calculated by combining the multidimensional conditional probability model with meteorological data for the prediction period, and the prediction confidence interval is used to predict the output of new energy power generation and provide risk warning.

[0011] Furthermore, the step of generating error statistics using the initial dataset to construct a statistical data set includes:

[0012] Calculate the mean of the prediction errors in the initial dataset and generate the mean statistic;

[0013] Calculate the standard deviation of the prediction error and generate the standard deviation statistic;

[0014] Determine the upper and lower limits of the prediction error and generate extreme value statistics;

[0015] The mean statistic, the standard deviation statistic, and the extreme value statistic are combined to generate an error statistic, thereby obtaining a set of statistical data.

[0016] Furthermore, the step of analyzing and processing the initial dataset using a preset Gaussian kernel function to generate an initial density estimate, and obtaining the joint probability density distribution based on the set of statistical parameters and the initial density estimate, includes:

[0017] The joint observation samples in the initial dataset are processed using a preset Gaussian kernel function. Each sample point in the joint observation sample is a multi-dimensional vector that contains meteorological data at a specific time and the corresponding prediction error.

[0018] A kernel function is placed at the location of each sample point, and all these kernel functions are superimposed to form an initial density estimate;

[0019] The broadband matrix in the initial density estimate is optimized using error statistics, thereby optimizing the initial density estimate and generating a multidimensional joint probability density estimate.

[0020] Furthermore, the step of optimizing the broadband parameters in the initial density estimate using error statistics to optimize the initial density estimate and generate a multidimensional joint probability density estimate includes:

[0021] When a new initial dataset becomes available, a new set of error statistics is generated by recalculating the prediction error data in this new dataset, including new mean statistics, standard deviation statistics, and extreme value statistics.

[0022] The bandwidth parameter in the kernel function is adjusted by using the newly generated error statistic. The bandwidth parameter that needs to be adjusted is set as a function that is proportional to the standard deviation statistic.

[0023] Using this set of bandwidth parameters adaptively adjusted by error statistics, kernel density estimation is recalculated on the new initial dataset to form an optimized density estimate that reflects a more accurate joint probability density distribution.

[0024] Furthermore, the step of extracting the prior probability distribution of the prediction error and the likelihood probability distribution of the preset meteorological conditions from the joint probability density distribution to generate a probability component includes:

[0025] The probability distribution of the prediction error itself is calculated using the joint probability density distribution, which is obtained by integrating or summing the joint probability density distribution along the dimensions of all meteorological conditions; the meteorological conditions are a vector consisting of one or more meteorological variables.

[0026] The likelihood probability distribution of the preset meteorological conditions is obtained by combining the probability distribution of the prediction error itself with the joint probability density distribution, thereby obtaining the probability component.

[0027] Furthermore, the method of obtaining a multidimensional conditional probability model based on the generation probability component includes:

[0028] The probability component is used to calculate the posterior probability distribution of the prediction error under any given preset meteorological conditions using Bayes' theorem.

[0029] Based on the calculated posterior probability distribution, a conditional probability model is generated. The conditional probability model is used to describe the probability density of the prediction error value for any set of input meteorological conditions.

[0030] Seasonal and geographical factors are acquired to generate multidimensional factor data; the multidimensional factor data is integrated with the conditional probability model to perform multidimensional conditional probability modeling, generating a multidimensional conditional probability model that constructs a unique conditional probability model for each specific combination of factors.

[0031] Furthermore, the step of using the multidimensional conditional probability model combined with meteorological data for the prediction period to calculate the prediction confidence interval, and then using the prediction confidence interval to predict the output of new energy power generation and provide risk warnings, includes:

[0032] Obtain the historical initial dataset and use the historical initial dataset to backtest the multidimensional conditional probability model to generate verification results;

[0033] Based on the verification results, the multidimensional conditional probability model is used to generate a probability distribution of prediction error for the meteorological data of the period to be predicted.

[0034] The prediction confidence interval is calculated based on the probability distribution of the prediction error, and the prediction confidence interval is used to predict the output of new energy power generation and provide risk warning.

[0035] Furthermore, calculating the prediction confidence interval based on the probability distribution of the prediction error includes:

[0036] A prediction confidence interval with a fixed confidence level is determined, and the prediction error confidence interval is determined based on the quantile function. This interval provides the quantization boundary of the prediction uncertainty.

[0037] The predicted power value at the point is combined with the confidence interval of the prediction error to obtain the prediction interval of the actual power generation in the future. The predicted power value at the point is obtained through the prediction model.

[0038] Risk classification and early warning are determined based on the width of the confidence interval of the prediction error and a preset width threshold.

[0039] Furthermore, the step of determining a prediction confidence interval with a fixed confidence level, and determining the prediction error confidence interval based on the quantile function, provides a quantized boundary for the prediction uncertainty, including:

[0040] Calculate a prediction confidence interval with a confidence level of (1-α) and find two error values ​​e. lower and e upper This satisfies the following formula:

[0041]

[0042] Where α is the significance level, P(E|W,S,L) represents the output of the multidimensional conditional probability model, L represents geographical location factors, E represents prediction error, W represents a vector consisting of one or more meteorological variables, i.e., preset meteorological conditions, and S represents seasonal factors; the obtained [e lower ,e upper Set as the confidence interval for the prediction error.

[0043] Furthermore, the step of determining the risk classification warning based on the width of the confidence interval of the prediction error and a preset width threshold includes:

[0044] When the confidence interval of the prediction error is [e lower ,e upper The corresponding interval width (e) upper -e lower When the width is narrower, i.e. less than the preset minimum width threshold, it indicates that the uncertainty of the prediction is small, and the actual output in the future will likely be close to the predicted power value at the point, which is a low-risk warning.

[0045] When the interval width is moderate, that is, within the preset maximum and minimum width thresholds, it indicates that there is a certain degree of uncertainty, which is a medium-risk warning.

[0046] When the interval width is very large, that is, greater than the maximum value of the preset width threshold, or the probability distribution shows obvious skewness or fat tails, it indicates that there is a risk of extreme overestimation or underestimation of future output, which is a high-risk warning.

[0047] On the other hand, the present invention also provides a new energy power generation prediction error analysis system that takes meteorological conditions into account, the system comprising:

[0048] The data processing module is used to construct an initial dataset based on historical data of new energy power generation, prediction error data and meteorological data, and to generate error statistics using the initial dataset, thereby constructing a statistical data set.

[0049] The kernel density estimation module is used to analyze and process the initial dataset using a preset Gaussian kernel function, generate an initial density estimate, and obtain a joint probability density distribution based on the set of statistical parameters and the initial density estimate.

[0050] The Bayesian computation module is used to extract the prior probability distribution of the prediction error and the likelihood probability distribution of the preset meteorological conditions from the joint probability density distribution, thereby generating probability components and obtaining a multidimensional conditional probability model based on the generated probability components.

[0051] The early warning module is used to calculate the prediction confidence interval by combining the multidimensional conditional probability model with meteorological data for the prediction period, and to perform new energy power generation output prediction and risk warning based on the prediction confidence interval.

[0052] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0053] 1. This invention significantly improves the quality and completeness of input data by systematically preprocessing the original data, including filling in missing data and correcting abnormal data. This provides a solid and reliable data foundation for the subsequent establishment of probabilistic models, thereby ensuring the accuracy and robustness of the entire modeling method from the source and avoiding model distortion caused by data quality issues.

[0054] 2. This invention uses a non-parametric kernel density estimation method to construct a joint probability density distribution and combines it with Bayes' theorem to generate a conditional probability model. This technical approach eliminates the pre-assumptions about the form of data distribution and can flexibly capture the complex nonlinear and multimodal correlation between prediction errors and meteorological factors, enabling the model to more realistically reflect the random characteristics of the physical world and improve the accuracy and adaptability of probability prediction.

[0055] 3. This invention innovatively introduces multi-dimensional factors such as season and geographical location to extend the conditional probability model and construct a refined multi-dimensional conditional probability model. This scenario-based modeling strategy can accurately depict the unique patterns of prediction errors under different spatiotemporal backgrounds, making the model more targeted and adaptable to different scenarios, thereby significantly improving prediction performance in diverse application environments.

[0056] 4. This invention transforms complex probability models into intuitive hierarchical early warning indicators and performs risk assessment by calculating prediction confidence intervals, thus realizing the transformation from theoretical models to practical decision support tools. This approach enables users such as power grid dispatchers to clearly understand and apply the uncertainty information of predictions, thereby conducting forward-looking risk management and scientific decision-making, and enhancing the safe and stable operation capability of the power system.

[0057] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description

[0058] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0059] Figure 1 This is a flowchart illustrating a method for predicting errors in new energy power generation considering meteorological conditions, according to an embodiment of the present invention.

[0060] Figure 2This is a joint probability density distribution diagram of wind speed and prediction error according to an embodiment of the present invention;

[0061] Figure 3 This is a conditional probability distribution diagram of prediction error under different wind speed conditions according to an embodiment of the present invention;

[0062] Figure 4 This is a comparison diagram of the multidimensional conditional probability models in embodiments of the present invention;

[0063] Figure 5 This is a schematic diagram of the structure of a new energy power generation prediction error analysis system that takes meteorological conditions into account, according to an embodiment of the present invention. Detailed Implementation

[0064] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0065] Reference Figure 1 One embodiment of the present invention proposes a method for analyzing the prediction error of new energy power generation considering meteorological conditions. It adopts a combination of kernel density estimation and Bayesian theory, and incorporates multidimensional spatiotemporal factors for refined modeling, which can accurately quantify the uncertainty of new energy power generation prediction and achieve forward-looking risk warning.

[0066] The method described in this embodiment specifically includes:

[0067] Acquire historical data, prediction error data, and meteorological data of new energy power generation, and preprocess them to generate an initial dataset;

[0068] Based on the initial dataset, the statistics of the prediction error are calculated, and the error statistics are generated.

[0069] Based on the error statistics and the meteorological data in the initial dataset, the joint probability density of the meteorological data and the prediction error data is estimated using the kernel density estimation method, and a joint probability density distribution is generated.

[0070] Based on the joint probability density distribution, the conditional probability distribution of the prediction error under the preset meteorological conditions is calculated using Bayes' theorem, and a conditional probability model is generated.

[0071] Based on climate and location, multidimensional conditional probability modeling is performed on the conditional probability model to generate a multidimensional conditional probability model.

[0072] Specifically, this invention significantly enhances the quantitative understanding and proactive management capabilities regarding the uncertainties of renewable energy power generation. Through the constructed joint probability model, this method no longer provides a single, deterministic prediction value, but instead generates a prediction error distribution containing rich probabilistic information. This allows managers to intuitively understand the probability of all possible future power generation outcomes, dynamically assess the reliability of prediction results based on specific weather forecasts, and generate forward-looking risk warnings. This shift from "point prediction" to "probabilistic prediction" provides a more scientific and refined decision-making basis for grid dispatching, electricity market trading, and power plant operation and maintenance. It helps optimize resource allocation, reduce system operation risks, and effectively address the inherent volatility and intermittency challenges of renewable energy power generation, thus promoting the safe, stable, and economical operation of the power system.

[0073] Optionally, generating the initial dataset includes:

[0074] Acquire historical data, forecast error data, and meteorological data for new energy power generation, and identify missing and abnormal data;

[0075] The missing data is filled using the nearest neighbor interpolation method to generate the filled data;

[0076] The abnormal data is corrected using median filtering and Z-score detection to generate corrected data;

[0077] Integrate the populated data and the corrected data to generate the initial dataset.

[0078] Specifically, the method for generating the initial dataset begins with acquiring basic data, encompassing three dimensions: historical data on new energy power generation, such as the actual power generation of photovoltaic power plants or wind farms in specific past time series; prediction error data, i.e., the deviation between the predicted power and the actual power generation in the same historical period; and meteorological data closely related to power generation, such as light intensity, wind speed, wind direction, temperature, and humidity. After acquiring this multi-source data, a preliminary data quality check is conducted, focusing on identifying two common data quality issues: missing data and outlier data. Missing data refers to null or unrecorded values ​​in the time series or specific data records. Outlier data refers to values ​​that significantly deviate from the normal data fluctuation range due to measurement equipment failure, data transmission errors, or other interference factors.

[0079] For the identified missing data, this method employs nearest neighbor imputation. The process first identifies the K nearest neighbors (K NN) of a sample containing a missing value within the dataset, comparing it to the K NN across all other non-missing features. Similarity is typically measured using standards such as Euclidean distance. After finding the K NN, the average or median of the values ​​for the corresponding missing features from these neighboring samples is used as the imputation value to fill the missing positions in the original data, thus generating fully imputed data.

[0080] To correct identified outliers, this method employs a combination of median filtering and Z-score detection. First, for time-series data, median filtering is used. This technique uses a sliding window to traverse the entire data sequence, replacing the original value at the window's center point with the median value of each data point within the window. This effectively filters out short-term, intense impulse noise or isolated extreme values ​​while preserving the data's edge information. Second, Z-score detection is used to identify global statistical anomalies. This method calculates the Z-score value for each data point using the following formula:

[0081]

[0082] Where Z is the Z-score of the data point x to be detected, x is the data point to be detected, μ is the mean of the dataset to which the data point belongs, and σ is the standard deviation of the dataset. By presetting a threshold, such as 3, when the absolute value of the Z-score of a data point exceeds the threshold, the point is considered to be outlier. For data points identified as outliers, the mean or median of neighboring values, or local interpolation methods can be used for replacement and correction. After completing the above filtering and detection correction, the corrected data is obtained.

[0083] Finally, the padded data generated after nearest neighbor imputation is integrated with the corrected data generated after median filtering and Z-score detection to form a high-quality dataset with complete data and no significant outliers. This dataset serves as the initial dataset required for subsequent joint probability modeling.

[0084] Optionally, the generation error statistics include:

[0085] Calculate the mean of the prediction errors in the initial dataset and generate the mean statistic;

[0086] Calculate the standard deviation of the prediction error and generate the standard deviation statistic;

[0087] Determine the upper and lower limits of the prediction error and generate extreme value statistics;

[0088] The mean statistic, the standard deviation statistic, and the extreme value statistic are combined to generate the error statistic.

[0089] Specifically, first, the mean of the prediction errors is calculated to generate a mean statistic. The mean reflects the central tendency or average level of the prediction errors and can reveal whether the prediction model has a systematic bias, i.e., a long-term tendency to overestimate or underestimate. The calculation method is to sum the prediction error values ​​at all times in the initial dataset and then divide by the total number of error data samples. The formula is as follows:

[0090]

[0091] Where, μ e This is the generated mean statistic, representing the mathematical expectation of the prediction error; e i represents the value of the i-th prediction error sample in the initial dataset, which is calculated from the actual power and the corresponding predicted power in the historical data of new energy power generation; N is the total number of prediction error samples in the initial dataset. e e i They all have the same dimensions as power generation.

[0092] Secondly, the standard deviation of the prediction error is calculated to generate the standard deviation statistic. The standard deviation measures the dispersion of the prediction error data relative to its mean, and is a key indicator of prediction uncertainty or volatility. A larger standard deviation means a wider range of fluctuation in the prediction error and poorer stability of the prediction results. The calculation formula is as follows:

[0093]

[0094] Where, σ e This is the generated standard deviation statistic, where sqrt is the square root calculation function. This formula calculates the sample standard deviation, which is then corrected by N-1 to obtain an unbiased estimate of the population standard deviation. Next, upper and lower limits for the prediction error are determined to generate extreme value statistics.

[0095] This step involves iterating through all prediction error data in the initial dataset to identify and record the maximum and minimum values. These two values ​​define the boundaries of prediction error fluctuations in historical observation data, providing a direct basis for assessing extreme prediction risks. The upper and lower limits together constitute the extreme value statistics.

[0096] Finally, the mean, standard deviation, and extreme value statistics calculated in the preceding steps are combined to form a structured dataset or parameter set. This set fully describes the basic statistical characteristics of the prediction error and ultimately generates error statistics, providing core parameter support for subsequent joint probability density distribution estimation.

[0097] Optionally, the generation of the joint probability density distribution includes:

[0098] The meteorological data and prediction error in the initial dataset are processed using a preset Gaussian kernel function to generate an initial density estimate;

[0099] A joint probability density distribution is generated based on the error statistic and the initial density estimate.

[0100] Specifically, this step aims to construct a joint probability density distribution between meteorological data and prediction error data based on the initial dataset and error statistics generated in the previous steps. The core of this process is the use of kernel density estimation, a nonparametric statistical technique that can estimate the probability density function of the data from the data itself, without pre-assuming the data's distribution form, such as a Gaussian distribution. Figure 2 As shown, this invention uses the kernel density estimation method, which can flexibly capture the complex nonlinear and multimodal correlation between meteorological data such as wind speed and prediction error. The shades of color in the figure represent the probability density of a specific wind speed and error combination, intuitively showing the inherent statistical relationship between the two, and avoiding the rigid assumptions about the distribution pattern made by traditional parameterization methods.

[0101] The process first uses a pre-defined Gaussian kernel function to process the joint observation samples in the initial dataset. Each sample point is a multi-dimensional vector containing meteorological data at a specific time and the corresponding prediction error.

[0102] The basic idea of ​​kernel density estimation is to place a kernel function (in this method, a Gaussian kernel function) at the location of each sample point, and then superimpose all these kernel functions to form an estimate of the overall probability density. This initial superposition generates a multidimensional joint probability density estimate, which is the initial density estimate.

[0103] The calculation of this multidimensional joint probability density estimate can be expressed by the following formula:

[0104]

[0105] Among them, f hat (X) is the joint probability density estimate at point X in the multidimensional space. X is a vector containing a specific combination of meteorological data and prediction error values, such as [light intensity, temperature, prediction error]. N is the total number of samples in the initial dataset. i K is the vector of the i-th observation sample in the initial dataset. H It is a multidimensional kernel function, defined by a Gaussian kernel function K and a bandwidth matrix H. The form of the Gaussian kernel function is similar to the density function of a multidimensional normal distribution.

[0106] The bandwidth matrix H is a key parameter that controls the smoothness of the estimated probability density curve. An initial, pre-defined H can be chosen based on standard empirical rules, such as Silverman's rule, to generate the initial density estimate. Subsequently, this method uses the error statistic calculated in the previous steps to optimize the initial density estimate to generate the final joint probability density distribution.

[0107] Specifically, the standard deviation and extreme value statistics in the error statistics provide direct data-driven basis for adjusting the bandwidth matrix H. The diagonal elements of the bandwidth matrix H are typically proportional to the standard deviation of the corresponding dimension. Therefore, using the standard deviation statistics of the prediction error dimension to set or optimize the bandwidth parameters in the bandwidth matrix corresponding to the prediction error dimension can make the smoothness of the density estimate more consistent with the actual discrete characteristics of the prediction error data. This adjustment based on error statistics corrects and optimizes the initial density estimate, thereby generating a joint probability density distribution that more accurately reflects the inherent statistical correlation between meteorological variables and prediction errors.

[0108] Optionally, generating the joint probability density distribution based on the error statistic and the initial density estimate includes:

[0109] When a new initial dataset is obtained, the bandwidth parameter of the Gaussian kernel function is adjusted by the error statistics to optimize the initial density estimate and generate an optimized density estimate.

[0110] Based on the optimized density estimate, a joint probability density distribution is generated.

[0111] Specifically, this step elaborates on how to dynamically utilize error statistics to optimize the generation process of the joint probability density distribution, the core of which lies in an adaptive update mechanism.

[0112] This process is triggered when a new initial dataset is acquired, for example, after the system has accumulated a period of new operational data, requiring model updates to reflect the latest system characteristics. When a new initial dataset becomes available, the prediction error data in this new dataset is recalculated to generate a new set of error statistics, including new mean, standard deviation, and extreme value statistics. These new statistics reflect the distribution characteristics of recent prediction errors. The bandwidth parameter of the Gaussian kernel function is then adjusted using the newly generated error statistics. The bandwidth parameter is a crucial smoothing parameter in kernel density estimation, directly determining the shape of the final probability density curve. A fixed bandwidth may not be able to adapt to changes in the data distribution over time.

[0113] This method uses newly calculated error statistics, especially the standard deviation statistic, to dynamically adjust the bandwidth.

[0114] A common adjustment criterion is to set the bandwidth as a function proportional to the standard deviation of the data. For example, for the dimension of prediction error, the bandwidth parameter h... e Based on the newly calculated standard deviation σ e Adjustments need to be made, namely:

[0115]

[0116] in, These are the updated bandwidth parameters. This is the latest standard deviation statistic calculated from the new initial dataset, where c is a proportionality constant determined empirically or through cross-validation. In this way, when the volatility of the prediction error increases... As the value increases, the bandwidth also increases accordingly, making the density estimation smoother to capture a wider distribution; conversely, when the error volatility decreases... As the value decreases, the bandwidth also decreases, allowing density estimation to more precisely characterize the concentrated data distribution.

[0117] Using this set of bandwidth parameters adaptively adjusted by error statistics, the kernel density estimation is recalculated on the new initial dataset. Since the bandwidth parameters have been optimized based on the latest characteristics of the data, the density estimate obtained in this calculation is no longer the initial density estimate, but an optimized density estimate.

[0118] Finally, this optimized density estimate is used as the final output. This output is a more accurate joint probability density distribution that reflects the latest data characteristics, and it will be used for subsequent conditional probability calculations.

[0119] Optionally, the generating conditional probability model includes:

[0120] Extract the prior probability distribution of the prediction error and the likelihood probability distribution of the preset meteorological conditions from the joint probability density distribution to generate a probability component;

[0121] The posterior probability is calculated by applying Bayes' theorem to the aforementioned probability components, thereby generating a posterior probability distribution.

[0122] Based on the posterior probability distribution, a conditional probability model is generated.

[0123] Specifically, the goal of this step is to construct a conditional probability model of prediction error under specific meteorological conditions based on the joint probability density distribution generated in the preceding steps, by applying Bayes' theorem. Based on the joint probability density distribution, and using Bayes' principle, the conditional probability distribution of prediction error under preset meteorological conditions can be obtained, such as... Figure 3As shown, the probability distribution curves of the prediction error are displayed when the wind speed is 5 m / s, 10 m / s, and 15 m / s. It can be seen that the mean, variance, and shape of the error distribution all change with the change of wind speed. This is the conditional probability model generated by this invention, which can accurately characterize the uncertainty under specific meteorological scenarios.

[0124] This process realizes the transformation from joint probability to conditional probability, which is the core link in performing prediction uncertainty analysis in specific scenarios.

[0125] First, the required probability components for calculation need to be extracted from the established joint probability density distribution. This joint probability density distribution is denoted as f(E,W), where E represents the prediction error and W represents a vector consisting of one or more meteorological variables, i.e., the preset meteorological conditions.

[0126] The first step is to extract the prior probability distribution of the prediction error, denoted as P(E). This is the probability distribution of the prediction error itself without considering any specific meteorological information. It can be obtained by integrating or summing the joint probability density distribution f(E,W) along the dimensions of all meteorological variables W, i.e.:

[0127] P(E) = ∫f(E,W)dW

[0128] This process is mathematically called marginalization, and the resulting P(E) describes the overall distribution characteristics of the prediction error.

[0129] Secondly, it is necessary to extract the likelihood probability distribution of a specific meteorological condition W given a prediction error E, denoted as P(W|E). According to the definition of conditional probability, it can be directly calculated from the joint probability density distribution and the prior probability distribution, as shown in the formula:

[0130]

[0131] These two parts, namely the prior probability distribution P(E) of the prediction error and the likelihood probability distribution P(W|E) of the preset meteorological conditions, together constitute the probability components required for applying Bayes' formula.

[0132] Next, Bayes' theorem is applied to calculate the posterior probability, which is the conditional probability distribution of the prediction error E given the observed values ​​of the preset meteorological conditions W, denoted as P(E|W). This is the core of the conditional probability model. Bayes' theorem is expressed as follows:

[0133]

[0134] Where P(E|W) is the posterior probability distribution; P(W|E) is the likelihood probability; P(E) is the prior probability; and P(W) is the marginal probability distribution of meteorological condition W, also known as the evidence factor, which is obtained by integrating the joint probability density distribution f(E,W) along the dimension of the prediction error E.

[0135] P(W) = ∫f(E,W)dE

[0136] By substituting each probability component into Bayes' formula, the posterior probability distribution of the prediction error E under any given preset meteorological conditions W can be calculated.

[0137] Finally, based on the calculated posterior probability distribution, a conditional probability model is generated. This model is not a single numerical value, but rather a function or a probability distribution curve that describes the probability density of possible values ​​for the prediction error E for any set of input weather conditions W. This model can be stored as a queryable data structure or a computational function, and when new weather forecast data is input, it can instantly output the corresponding conditional probability distribution of the prediction error.

[0138] Optionally, the generation of the multidimensional conditional probability model includes:

[0139] Acquire seasonal and geographical factors to generate multidimensional factor data;

[0140] The multidimensional factor data is integrated with the conditional probability model to perform multidimensional conditional probability modeling and generate a multidimensional conditional probability model.

[0141] Specifically, this step aims to expand the dimensionality of the conditional probability model generated in the previous steps to improve its applicability and accuracy. This process upgrades the original model based on instantaneous weather conditions into a multidimensional conditional probability model by introducing seasonal and geographical location factors. For example... Figure 4 As shown, under the same wind speed (12 m / s), the prediction error distribution in the "winter-coastal" scenario exhibits a wider shape and a lower mean, while the "summer-inland" scenario shows different distribution characteristics. This indicates that the multidimensional model can accurately capture the unique patterns of prediction errors under different spatiotemporal backgrounds, significantly improving the model's scenario adaptability.

[0142] Generating a multidimensional conditional probability model first requires acquiring and processing additional multidimensional factor data. The first category is seasonal factors. This can be obtained by dividing historical data by season, such as spring, summer, autumn, and winter. For each season, its corresponding historical renewable energy generation data, prediction error data, and meteorological data can be extracted separately to form a specific seasonal subset. The second category is geographical location factors. For multiple renewable energy power plants deployed in different geographical locations, it is necessary to collect specific location information for each plant, such as latitude and longitude, altitude, and topographic features. Simultaneously, each plant has its own independent initial dataset. These seasonal divisions and geographical location identifiers together constitute the multidimensional factor data.

[0143] Next, these multidimensional factor data will be integrated with the established conditional probability model to implement multidimensional conditional probability modeling. The core idea of ​​the integration is to build a dedicated conditional probability model for each specific combination of factors, such as "summer" and "Station A". Operationally, this means that the data needs to be grouped.

[0144] Taking seasonal factors as an example, the entire initial dataset will be divided into four subsets, corresponding to spring, summer, autumn, and winter, respectively. For each seasonal subset, the complete process of generating conditional probability models described above is repeated, starting from data preprocessing, calculating the error statistics for that season, establishing the joint probability density distribution of weather and error for that season, and finally generating a conditional probability model for prediction errors specific to that season. In this way, the original single conditional probability model P(E|W) is transformed into four models, namely P(E|W,S=Spring), P(E|W,S=Summer), P(E|W,S=Autumn), and P(E|W,S=Winter), where S represents the seasonal factor.

[0145] In this embodiment, W does not refer to a single meteorological indicator, but rather a data vector containing one or more sets of meteorological variables closely related to the power output of new energy sources. When this embodiment mentions "the joint probability density of meteorological data and prediction error data," it actually refers to the joint probability density between the one-dimensional variable of prediction error and the multi-dimensional vector W composed of multiple meteorological indicators.

[0146] In this embodiment, the output is the joint probability density function P(E,W), where E represents the prediction error and W represents the meteorological condition vector. This function describes the probability density of a specific prediction error value e and a specific meteorological condition w occurring simultaneously. W is one of the two basic variables in the joint distribution; without W, the correlation between the two cannot be established.

[0147] like Figure 2This is a two-dimensional example. The x-axis represents "wind speed" (a component of W), and the y-axis represents "prediction error" E. This graph shows the joint probability density of P(wind speed, prediction error). Without the wind speed data (W), this graph cannot be generated. The conditional probability density function is P(E|W). W is the given condition here. The purpose of the entire Bayesian calculation is to derive the conditional probability P(E|W) from the joint probability P(E,W). This model answers the question: if the future weather conditions are known to be W, what is the probability distribution of the prediction error E? W is the model input, and the probability distribution of E is the model output.

[0148] like Figure 3 The image shows three different curves, corresponding to the conditional probability distribution of the prediction error E under three different preset meteorological conditions W: wind speed of 5 m / s, wind speed of 10 m / s, and wind speed of 15 m / s. It can be seen that as the wind speed W changes, the probability distribution shape (mean, variance, etc.) of E also changes significantly. If W had no effect, these three curves should overlap.

[0149] Similarly, for geographical location factors, a separate conditional probability model is constructed for each subset of data from different geographical locations, such as station A and station B. If both seasonal and geographical location dimensions are considered, a separate model needs to be built for each "season-location" combination, such as "summer-station A" and "winter-station B". The final generated model will then be in the form of P(E|W,S,L), where L represents the geographical location factor. This set of models constitutes the final multidimensional conditional probability model. In practical applications, the specific conditional probability model corresponding to the season of the period to be predicted and the geographical location of the target station is selected and used for calculation.

[0150] Optionally, the method further includes:

[0151] Obtain the historical initial dataset and use the historical initial dataset to backtest the multidimensional conditional probability model to generate verification results;

[0152] Based on the verification results, a tiered early warning index is generated for predicting and issuing early warnings for new energy power generation.

[0153] Specifically, after constructing the multidimensional conditional probability model, this method adds a crucial application extension step: applying the model to actual new energy power generation forecasting and early warning. This process includes two main parts: model validation and early warning indicator generation.

[0154] First, backtesting is performed to validate the model. This step aims to evaluate the historical performance and predictive ability of the multidimensional conditional probability model. Operationally, this requires obtaining an independent, unused historical initial dataset, typically from an earlier period or a reserved test dataset. Then, iterate through each time point in this historical initial dataset, extracting the actual meteorological data, seasonal information, and geographical location information for that specific time point.

[0155] Using this information as input, the corresponding multidimensional conditional probability model is invoked, and the model outputs the conditional probability distribution of the prediction error under the given historical weather conditions.

[0156] Then, the predicted probability distribution is compared with the actual prediction error value at that time point. Through backtesting with a large number of sample points, a series of probability prediction evaluation indicators can be calculated, such as coverage, interval width, and continuous graded probability score (CRPS). These indicators together constitute the validation results, quantitatively proving the accuracy and reliability of the model.

[0157] Secondly, based on the validation results, tiered early warning indicators are generated to achieve predictive early warning for new energy power generation. This step is the core of the model's transition from theory to application. When it is necessary to predict a certain period in the future, the meteorological forecast data for that period is first obtained, and the season and geographical location are determined. This information is input into a validated multidimensional conditional probability model, which generates a conditional probability distribution of the prediction error for that period. Based on this probability distribution, the prediction confidence interval at different confidence levels can be calculated. For example, the range of prediction errors with a 95% probability can be calculated. This interval provides a quantitative description of the uncertainty of future actual power generation. Combined with the point prediction value, i.e., the deterministic predicted power given by the prediction system, this error interval can be transformed into a prediction interval for power generation output.

[0158] Finally, based on this probabilistic information, graded early warning indicators are set. For example, different early warning levels can be defined. Level 1 (low risk): the probability distribution of the prediction error is concentrated and narrow, and the prediction confidence interval is small, indicating that the prediction results are very reliable. Level 2 (medium risk): the prediction confidence interval widens, indicating increased uncertainty, and the actual output may deviate significantly. Level 3 (high risk): the prediction confidence interval is very wide, or the probability distribution shows a bimodal or fat-tailed pattern, indicating the possibility of extreme weather events, and the prediction error may be very large. Dispatchers and power plant operation and maintenance personnel can formulate response strategies in advance based on the issued early warning levels, such as adjusting reserve capacity, arranging maintenance plans, or participating in the ancillary services market, thereby proactively managing the uncertainty risks of new energy power generation.

[0159] Optionally, generating tiered early warning indicators for new energy power generation forecasting and early warning based on the verification results includes:

[0160] Based on the verification results, the multidimensional conditional probability model is used to generate a probability distribution of prediction error for the meteorological data of the period to be predicted.

[0161] Calculate the prediction confidence interval based on the probability distribution of the prediction error;

[0162] The predicted confidence interval is used to predict the output of new energy power generation and provide risk warnings.

[0163] Specifically, the probability information output by the model is transformed into prediction confidence intervals with practical guiding significance, and risk warnings are issued accordingly.

[0164] First, for the future period to be predicted, a validated multidimensional conditional probability model is used to generate the probability distribution of the prediction error. This process begins with acquiring the forecast meteorological data, seasonal affiliation, and geographical location information for that period.

[0165] Using this information as input, a matching multidimensional conditional probability model is invoked. The model will output a continuous or discrete probability density function (PDF) or cumulative distribution function (CDF), denoted as P(E|W,S,L). This function fully describes the relative probability of all possible values ​​of the prediction error E under this specific condition.

[0166] Secondly, the prediction confidence interval is calculated based on the probability distribution of this prediction error. The confidence interval is a range determined according to a preset confidence level, such as 90% or 95%, indicating a degree of certainty that the actual prediction error will fall within this range. The calculation method involves solving for the inverse function of the cumulative distribution function (CDF), also known as the quantile function. For example, to calculate a prediction confidence interval with a confidence level of (1-α), two error values ​​e need to be found. lower and e upper , so that:

[0167]

[0168] Where α is the significance level; for example, for a 95% confidence level, α = 0.05. The solution obtained is [e...]. lower ,e upper This represents the confidence interval for the prediction error. This interval provides a quantification boundary for the prediction uncertainty.

[0169] Finally, by combining the predicted power value with the confidence interval of the prediction error, the prediction interval of the actual future power generation can be directly obtained. The predicted power value is given by the conventional prediction model and is denoted as P. forecast The predicted range for actual output is:

[0170] [Pforecast +e lower ,P forecast +e upper ],

[0171] This range visually illustrates the potential fluctuation range of future power generation. Risk warnings are based on the characteristics of this range.

[0172] When the calculated confidence interval [e lower ,e upper The width of ]e upper -e lower When the width is narrower, i.e. less than the preset minimum width threshold, it indicates that the uncertainty of the prediction is small, the actual output in the future will likely be close to the predicted value, the system risk is low, and it is a low-risk warning.

[0173] When the interval width is moderate, i.e., within the preset maximum and minimum width thresholds, it indicates a certain degree of uncertainty, requiring attention from dispatchers and consideration of preparing appropriate reserve capacity, thus serving as a medium-risk warning.

[0174] When the interval width is very large, exceeding the preset maximum width threshold, or when the probability distribution exhibits significant skewness or fat-tailed characteristics, it indicates that future power output may deviate significantly, even posing a risk of extreme overestimation or underestimation. Skewness or fat-tailed characteristics can be identified by calculating indicators such as skewness and kurtosis. In such cases, a high-risk warning should be issued, prompting dispatchers to adopt more conservative strategies, such as increasing spinning reserves and adjusting generation plans, to cope with potential large power fluctuations. By setting different interval width thresholds, automated tiered warnings can be achieved.

[0175] Based on the same inventive concept, such as Figure 5 As shown, the present invention also provides a new energy power generation prediction error analysis system that takes meteorological conditions into account, the system comprising:

[0176] The data processing module is used to acquire historical data of new energy power generation, prediction error data and meteorological data, and to preprocess them to generate an initial dataset;

[0177] The error statistics module is used to calculate the statistics of the prediction error based on the initial dataset and generate error statistics.

[0178] The kernel density estimation module is used to estimate the joint probability density of meteorological data and prediction error data based on the error statistics and meteorological data in the initial dataset, and to generate a joint probability density distribution.

[0179] The Bayesian calculation module is used to calculate the conditional probability distribution of the prediction error under preset meteorological conditions based on the joint probability density distribution and using the Bayesian formula to generate a conditional probability model.

[0180] The model building module is used to perform multidimensional conditional probability modeling on the conditional probability model based on climate and location, and generate a multidimensional conditional probability model.

[0181] To verify the feasibility of this invention in practice, it was applied to a large-scale wind farm cluster located in different geographical environments. This cluster includes wind farms in both coastal and inland terrains, exhibiting significant seasonal and regional differences in wind resources, resulting in strong fluctuations in power generation and making forecasting difficult. Traditional point forecasting methods cannot quantify the uncertainty of forecasts, posing a significant challenge to grid dispatching. The group aims to use the method of this invention to probabilistically model wind power forecasting errors, thereby achieving more refined risk warnings and dispatching decision support.

[0182] In this embodiment, historical data on renewable energy generation, prediction error data, and corresponding high-resolution meteorological data of the wind farm cluster were collected. The meteorological data included wind speed, wind direction, temperature, and humidity. First, the data processing module preprocessed the acquired raw data, using nearest neighbor interpolation to fill in missing data caused by sensor communication interruptions, and combined median filtering and Z-score detection to correct abnormal data caused by equipment failures. Finally, a high-quality initial dataset was generated.

[0183] Based on the initial dataset, the error statistics module calculated the mean, standard deviation, and extreme values ​​of the prediction error, generating error statistics. Subsequently, the kernel density estimation module used a Gaussian kernel function and, in conjunction with the error statistics, particularly the standard deviation, optimized the bandwidth parameter to perform joint probability density estimation on the meteorological data and prediction error data. The Bayesian calculation module then used this joint probability density to calculate the conditional probability distribution of the prediction error under specific meteorological conditions. In particular, to reflect the influence of climate and location, the model building module divided the data by season (winter and summer) and geographical location (coastal station A and inland station B), establishing a dedicated multidimensional conditional probability model for each "season-location" combination. Finally, the model was backtested using historical data for validation, and tiered early warning indicators were generated.

[0184] Analysis of data from the winter of 2024 revealed that coastal power station A experienced frequent low temperatures and strong winds due to the influence of strong cold air, resulting in a significantly increased standard deviation of its prediction error and a wider probability distribution. In summer, influenced by sea and land breezes, the error distribution, while fluctuating, remained relatively concentrated. For inland power station B, the winter error volatility was lower than that of coastal power station A, but the probability of short-term extreme errors was higher in summer due to convective weather. This invention's multidimensional model successfully captured these spatiotemporal differences. For example, on August 15, 2024, a weather forecast indicated a strong typhoon would affect coastal power station A. This invention used the "Summer-A Power Station" model for analysis, predicting an unusually wide 95% confidence interval for the prediction error during that period, far exceeding typical weather conditions. Based on this, the system issued the highest-level high-risk warning, prompting the dispatch center to adopt a conservative power generation plan and reserve sufficient backup capacity. Backtesting results showed that this warning effectively prevented a power grid frequency fluctuation event caused by a significant prediction deviation.

[0185] Data shows that the system of this invention significantly improves the ability to manage power generation uncertainties. Backtesting results show that the 95% prediction confidence interval generated by the model achieves an actual coverage of 94.8% across the entire validation set, demonstrating high reliability. The early warning system based on this probabilistic prediction, compared to traditional threshold-based alarm methods, can improve the early warning accuracy of potential extreme power fluctuation events to over 95%, and extends the average early warning lead time by 3 hours. After the dispatch center optimizes its dispatch strategy based on these tiered early warning information, the cost of reserve capacity call-up due to prediction errors is reduced by approximately 18%.

[0186] Table 1. Comparison of Prediction Error Statistics for Different Seasons and Locations

[0187] winter Coastal Station A -5.2 25.8 -98.5~+89.1 winter Inland Station B -3.1 18.3 -75.4~+70.2 summer Coastal Station A 2.5 15.6 -60.3~+65.7 summer Inland Station B 4.1 19.5 -80.1~+82.4

[0188] Table 2 Comparison of Probability Predictions and Actual Results under Key Meteorological Events

[0189] 2025-01-20 Cold wave passes 280 190~350 215 High risk 2024-08-15 Typhoon impact 150 20~250 45 High risk 2024-10-05 The weather is stable 220 205~235 228 Low risk

[0190] Table 3 Performance Verification Data of Early Warning System

[0191] High risk 25 24 96.0 4.5 Medium risk 118 109 92.4 6.2 Low risk 589 578 98.1 N / A

[0192] As can be seen from the data in Tables 1 to 3 above, the method proposed in this invention has achieved significant technical results. Table 1 clearly shows the significant impact of different seasons and geographical locations on the statistical characteristics of prediction errors. For example, the standard deviation of the error at the coastal A power station in winter is 25.8 MW, which is much higher than in other scenarios. This fully demonstrates the necessity and effectiveness of the multidimensional conditional probability modeling of this invention. The data in Table 2 verifies the excellent prediction capability of this invention in specific key events. Under extreme weather conditions such as typhoons, although the predicted output of 150 MW deviates greatly from the actual output of 45 MW, the probability prediction range of 20–250 MW of this invention successfully covers the actual results and issues high-risk warnings in a timely manner, demonstrating its great value in risk management. Table 3 quantifies the overall performance of the early warning system. The early warning accuracy rate for high-risk events is as high as 96.0%, and it can provide an average lead time of 4.5 hours. This indicates that this method can successfully transform the abstract probability model into a reliable and operable decision support tool, providing strong technical support for ensuring the safe and stable operation of the power grid under large-scale grid connection of new energy sources.

[0193] It should be noted that the electrical connections between the various units described above do not necessarily represent direct or indirect connections. Any indirect connection method can be applied to the embodiments of the present invention as long as it achieves the purpose of the present invention. The above descriptions are merely exemplary embodiments of the present invention and should not be construed as limiting the scope of the present invention.

[0194] All equivalent changes and modifications made in accordance with the teachings of this invention are still within the scope of this invention. Those skilled in the art will readily conceive of other embodiments of this invention upon considering the specification and the disclosure of practical truth. This application is intended to cover any variations, uses, or adaptations of this invention that follow the general principles of this invention and include common knowledge or conventional techniques in the art not described herein.

Claims

1. A method for predicting errors in new energy power generation considering meteorological conditions, characterized in that, The method includes: An initial dataset is constructed based on historical data of new energy power generation, prediction error data and meteorological data, and error statistics are generated using the initial dataset to construct a statistical data set. The initial dataset is analyzed and processed using a preset Gaussian kernel function to generate an initial density estimate. Based on the statistical data set and the initial density estimate, the joint probability density distribution is obtained. The prior probability distribution of the prediction error and the likelihood probability distribution of the preset meteorological conditions are extracted from the joint probability density distribution to generate probability components, and a multidimensional conditional probability model is obtained based on the generated probability components. The prediction confidence interval is calculated by combining the multidimensional conditional probability model with meteorological data for the prediction period, and the prediction confidence interval is used to predict the output of new energy power generation and provide risk warning. The method for obtaining a multidimensional conditional probability model based on a generative probability component includes: The probability component is used to calculate the posterior probability distribution of the prediction error under any given preset meteorological conditions using Bayes' theorem. Based on the calculated posterior probability distribution, a conditional probability model is generated. The conditional probability model is used to describe the probability density of the prediction error value for any set of input meteorological conditions. Seasonal and geographical factors are acquired to generate multidimensional factor data; the multidimensional factor data is integrated with the conditional probability model to perform multidimensional conditional probability modeling, generating a multidimensional conditional probability model that constructs a unique conditional probability model for each specific combination of factors.

2. The method for predicting new energy power generation considering meteorological conditions according to claim 1, characterized in that, The step of generating error statistics using the initial dataset to construct a statistical data set includes: Calculate the mean of the prediction errors in the initial dataset and generate the mean statistic; Calculate the standard deviation of the prediction error and generate the standard deviation statistic; Determine the upper and lower limits of the prediction error and generate extreme value statistics; The mean statistic, the standard deviation statistic, and the extreme value statistic are combined to generate an error statistic, thereby obtaining a set of statistical data.

3. The method for predicting new energy power generation considering meteorological conditions according to claim 2, characterized in that, The process of analyzing and processing the initial dataset using a preset Gaussian kernel function to generate an initial density estimate, and obtaining the joint probability density distribution based on the statistical data set and the initial density estimate, includes: The joint observation samples in the initial dataset are processed using a preset Gaussian kernel function. Each sample point in the joint observation sample is a multi-dimensional vector that contains meteorological data at a specific time and the corresponding prediction error. A kernel function is placed at the location of each sample point, and all these kernel functions are superimposed to form an initial density estimate; The bandwidth matrix in the initial density estimate is optimized using error statistics, thereby optimizing the initial density estimate and generating a multidimensional joint probability density estimate.

4. The method for predicting new energy power generation considering meteorological conditions according to claim 3, characterized in that, The step of optimizing the bandwidth parameter in the initial density estimate using error statistics to generate a multidimensional joint probability density estimate includes: When a new initial dataset becomes available, a new set of error statistics is generated by recalculating the prediction error data in this new dataset, including new mean statistics, standard deviation statistics, and extreme value statistics. The bandwidth parameter in the kernel function is adjusted by using the newly generated error statistic. The bandwidth parameter that needs to be adjusted is set as a function that is proportional to the standard deviation statistic. Using this set of bandwidth parameters adaptively adjusted by error statistics, kernel density estimation is recalculated on the new initial dataset to form an optimized density estimate that reflects a more accurate joint probability density distribution.

5. The method for predicting new energy power generation considering meteorological conditions according to claim 1, characterized in that, The step of extracting the prior probability distribution of the prediction error and the likelihood probability distribution of the preset meteorological conditions from the joint probability density distribution to generate a probability component includes: The probability distribution of the prediction error itself is calculated using the joint probability density distribution, which is obtained by integrating or summing the joint probability density distribution along the dimensions of all meteorological conditions; the meteorological conditions are a vector consisting of one or more meteorological variables. The likelihood probability distribution of the preset meteorological conditions is obtained by combining the probability distribution of the prediction error itself with the joint probability density distribution, thereby obtaining the probability component.

6. The method for predicting new energy power generation considering meteorological conditions according to claim 1, characterized in that, The process of calculating the prediction confidence interval using the multidimensional conditional probability model combined with meteorological data for the prediction period, and then using the prediction confidence interval to predict the output of new energy power generation and provide risk warnings, includes: Obtain the historical initial dataset and use the historical initial dataset to backtest the multidimensional conditional probability model to generate verification results; Based on the verification results, the multidimensional conditional probability model is used to generate a probability distribution of prediction error for the meteorological data of the period to be predicted. The prediction confidence interval is calculated based on the probability distribution of the prediction error, and the prediction confidence interval is used to predict the output of new energy power generation and provide risk warning.

7. The method for predicting new energy power generation considering meteorological conditions according to claim 6, characterized in that, The calculation of the prediction confidence interval based on the probability distribution of the prediction error includes: A prediction confidence interval with a fixed confidence level is determined, and the prediction error confidence interval is determined based on the quantile function. This interval provides the quantization boundary of the prediction uncertainty. The predicted power value at the point is combined with the confidence interval of the prediction error to obtain the prediction interval of the actual power generation in the future. The predicted power value at the point is obtained through the prediction model. Risk classification and early warning are determined based on the width of the confidence interval of the prediction error and a preset width threshold.

8. The method for predicting new energy power generation considering meteorological conditions according to claim 7, characterized in that, The process involves determining a prediction confidence interval with a fixed confidence level and then determining the prediction error confidence interval based on a quantile function. This interval provides a quantization boundary for the prediction uncertainty, including: Calculate a confidence level of Find the prediction confidence interval and two error values. and This satisfies the following formula: in, It is the significance level. This represents the output of a multidimensional conditional probability model. Representing geographical location factors, Represents prediction error. A vector representing one or more meteorological variables, i.e., preset meteorological conditions. Representing seasonal factors; the solution obtained Set as the confidence interval for the prediction error.

9. The method for predicting new energy power generation considering meteorological conditions according to claim 8, characterized in that, The step of determining the risk classification warning based on the width of the confidence interval of the prediction error and a preset width threshold includes: When the confidence interval of the prediction error Corresponding interval width When the width is narrower, i.e. less than the preset minimum width threshold, it indicates that the uncertainty of the prediction is small, and the actual output in the future will likely be close to the predicted power value at the point, which is a low-risk warning. When the interval width is moderate, that is, within the preset maximum and minimum width thresholds, it indicates that there is a certain degree of uncertainty, which is a medium-risk warning. When the interval width is very large, that is, greater than the maximum value of the preset width threshold, or the probability distribution shows obvious skewness or fat tails, it indicates that there is a risk of extreme overestimation or underestimation of future output, which is a high-risk warning.

10. An error analysis system based on the new energy power generation prediction error analysis method considering meteorological conditions according to claim 1, characterized in that, The system includes: The data processing module is used to construct an initial dataset based on historical data of new energy power generation, prediction error data and meteorological data, and to generate error statistics using the initial dataset, thereby constructing a statistical data set. The kernel density estimation module is used to analyze and process the initial dataset using a preset Gaussian kernel function, generate an initial density estimate, and obtain a joint probability density distribution based on the statistical data set and the initial density estimate. The Bayesian computation module is used to extract the prior probability distribution of the prediction error and the likelihood probability distribution of the preset meteorological conditions from the joint probability density distribution, thereby generating probability components and obtaining a multidimensional conditional probability model based on the generated probability components. The early warning module is used to calculate the prediction confidence interval by combining the multidimensional conditional probability model with meteorological data for the prediction period, and to perform new energy power generation output prediction and risk warning based on the prediction confidence interval.

Citation Information

Patent Citations

  • New energy prediction error modeling method and system based on naive Bayes classification

    CN115965135A

  • Wind power prediction error estimation method based on meteorological prediction error

    CN119578611A