A method and related apparatus for photovoltaic power output probability prediction

By obtaining historical and predicted weather factor data and photovoltaic output data, calculating the correlation coefficient and collinearity strength, screening the target weather variables, and constructing a vine copula model, the problem of neglecting factor correlation in existing photovoltaic output probability prediction methods is solved, and a more accurate photovoltaic output probability prediction and optimized scheduling are achieved.

CN118982438BActive Publication Date: 2025-10-10GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411235318.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-04
Publication Date
2025-10-10
Estimated Expiration
2044-09-04

AI Technical Summary

Technical Problem

The existing photovoltaic output probability prediction method is difficult to consider the correlation of photovoltaic-related factors, which leads to the deviation of the optimization scheduling scheme from reality and fails to meet the economic and safety requirements of distribution network operation.

Method used

By obtaining historical and predicted weather factor data and photovoltaic output data, calculating the correlation coefficient and collinearity strength, screening target weather variables, building a Vine Copula model, determining the photovoltaic output confidence interval, and improving prediction accuracy.

Benefits of technology

It achieves more accurate photovoltaic output probability prediction, provides a more reasonable optimization scheduling plan, avoids the occurrence of power backflow and voltage limit exceeding, and improves the economy and safety of power grid operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118982438B_ABST
    Figure CN118982438B_ABST
Patent Text Reader

Abstract

The application provides a photovoltaic output probability prediction method and related devices. The method comprises: obtaining historical weather factor data, predicted weather factor data and historical photovoltaic output data of a region to be predicted; obtaining a photovoltaic output point prediction value according to the data and a neural network point prediction model; calculating a correlation coefficient and a collinearity intensity between the photovoltaic output and the weather factor, the photovoltaic output point prediction value by using the historical weather factor data, the historical photovoltaic output data and the photovoltaic output point prediction value; determining a target weather variable from the weather factor according to the correlation coefficient and the collinearity intensity; constructing a vine Copula model by using the target weather variable, the photovoltaic output and the photovoltaic output point prediction value; and determining a photovoltaic output confidence interval according to the vine Copula model, the predicted weather factor data, the photovoltaic output point prediction value and a preset confidence level, thereby improving the accuracy of the photovoltaic output probability prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of distributed photovoltaic power generation in distribution networks, and in particular to a photovoltaic output probability prediction method and related devices. Background Art

[0002] As the power grid undergoes an increasingly rapid energy transition, photovoltaic power generation has become a key clean energy source and is entering a phase of large-scale development. PV power generation is influenced by numerous factors, including high uncertainty, intermittency, and volatility. These characteristics pose a threat to the safe and stable operation of the power grid. Correlation modeling of PV output and its influencing factors is necessary to provide efficient and accurate probabilistic forecasts. This can improve the economics and efficiency of power grid operations and provide technical support for dispatching departments to proactively adjust power grid operations to respond to emergencies.

[0003] In recent years, with the widespread adoption of Numerical Weather Prediction (NWP) in photovoltaic power plants and the maturity of its technology, effective and accurate data on solar radiation, temperature, cloud cover, rainfall, and wind speed, synchronized with photovoltaic output, can be obtained. Smart meters can then monitor the power generation level of photovoltaic power plants in real time, collect photovoltaic output data, and connect to power grid operators via remote communication technologies. This power generation data is regularly reported to the grid operations center, allowing for timely adjustments to power generation and enabling the grid to achieve rapid and orderly load balancing. Therefore, existing technologies utilize this data to construct relevant probabilistic models to achieve probabilistic predictions of photovoltaic output.

[0004] However, existing PV output probability prediction methods struggle to provide accurate output boundaries by taking into account the correlation of PV-related factors. This results in deviating from the optimal scheduling scheme and failing to meet the economic and safety requirements of distribution network operation. Therefore, it is necessary to propose a new PV output probability prediction method. Summary of the Invention

[0005] The present invention provides a photovoltaic output probability prediction method and related devices for improving the accuracy of photovoltaic output probability prediction.

[0006] In one aspect, the present invention provides a photovoltaic output probability prediction method, comprising:

[0007] Obtain historical weather factor data, forecast weather factor data, and historical photovoltaic output data for the area to be predicted;

[0008] Determining a photovoltaic output point prediction value based on the historical weather factor data, the predicted weather factor data, and the historical photovoltaic output data;

[0009] Calculate the correlation coefficient and collinearity strength between photovoltaic output, weather factors, and the photovoltaic output point prediction value using the historical weather factor data, the historical photovoltaic output data, and the photovoltaic output point prediction value;

[0010] determining a target weather variable from the weather factors according to the correlation coefficient and the collinearity strength;

[0011] Constructing a Vine Copula model using the target weather variable, the photovoltaic output, and the photovoltaic output point prediction value;

[0012] A photovoltaic output confidence interval is determined based on the vine copula model, the predicted weather factor data, the photovoltaic output point prediction value, and a preset confidence level.

[0013] Optionally, determining the photovoltaic output point prediction value according to the historical weather factor data, the predicted weather factor data, and the historical photovoltaic output data includes:

[0014] Synchronously aligning the historical weather factor data, the predicted weather factor data, and the historical photovoltaic output data in a time dimension;

[0015] According to a preset division ratio, the historical weather factor data, the predicted weather factor data, and the historical photovoltaic output data after synchronization and alignment are divided into a training set, a validation set, and a set to be predicted;

[0016] The training set and the validation set are used to train a pre-built photovoltaic output point prediction model, and the set to be predicted is input into the trained photovoltaic output point prediction model to output the photovoltaic output point prediction value.

[0017] Optionally, the using the historical weather factor data, the historical photovoltaic output data, and the photovoltaic output point prediction value to calculate the correlation coefficient and collinearity strength between photovoltaic output, weather factors, and the photovoltaic output point prediction value includes:

[0018] Normalizing the historical weather factor data, the historical photovoltaic output data, and the photovoltaic output point prediction value respectively;

[0019] Calculating the Kendall correlation coefficient and variance inflation factor between the photovoltaic output, the weather factors, and the photovoltaic output point prediction values ​​based on the normalized historical weather factor data, the historical photovoltaic output data, and the photovoltaic output point prediction values;

[0020] The Kendall correlation coefficient was used as the correlation coefficient, and the variance inflation factor was used as the collinearity strength.

[0021] Optionally, constructing a vine tree using the target weather variable, the photovoltaic output, and the photovoltaic output point prediction value specifically includes:

[0022] S51. Using a nonparametric kernel density estimation method, perform probability integral transformation on the target weather variable, the photovoltaic output, and the photovoltaic output point prediction value, respectively, to obtain the marginal probability distribution of the target weather variable, the marginal probability distribution of the photovoltaic output, and the marginal probability distribution of the photovoltaic output point prediction value;

[0023] S52: using a reinforcement learning algorithm to determine an optimal connection relationship among the target weather variable, the photovoltaic output, and the photovoltaic output point prediction value;

[0024] S53, using the marginal probability distribution of the target weather variable, the marginal probability distribution of the photovoltaic output, and the marginal probability distribution of the photovoltaic output point prediction value as vine tree nodes, and connecting the vine tree nodes using the optimal connection relationship to obtain a target vine tree;

[0025] S54, using the maximum likelihood estimation method and the Akaike Information Criterion method to determine the optimal Copula function and the parameter value of the optimal Copula function for each edge in the target vine tree, and using the optimal Copula function, the parameter value of the optimal Copula function, and each vine tree node to respectively calculate the conditional probability of each edge in the target vine tree;

[0026] S55. Taking the conditional probability of each edge as the vine tree node of the next vine tree;

[0027] S56. Using a reinforcement learning algorithm, determine the optimal connection relationship between the vine tree nodes of the next vine tree;

[0028] S57. Connecting the vine tree nodes of the next vine tree according to the optimal connection relationship between the vine tree nodes of the next vine tree to obtain the next vine tree;

[0029] S58: Update the next vine tree to the target vine tree in step S54, and jump to S59;

[0030] S59. Loop through steps S54 to S58 until the updated target vine tree in step S58 has only one edge. Then, stop the loop and output the target vine tree and each updated target vine tree as the vine copula model.

[0031] Optionally, the determining of the optimal Copula function of each edge in the target vine tree and the parameter value of the optimal Copula function by using the maximum likelihood estimation method and the Akaike information criterion method includes:

[0032] Obtaining multiple Copula functions for each edge in the target vine tree respectively;

[0033] Calculating the parameter value of each copula function using maximum likelihood estimation method;

[0034] Calculate the AIC value of each Copula function according to each Copula function and the parameter value of each Copula function;

[0035] The Copula function with the smallest AIC value is taken as the optimal Copula function, and the parameter value of the Copula function with the smallest AIC value is taken as the function value of the optimal Copula function.

[0036] Optionally, the determining of the photovoltaic output confidence interval according to the vine copula model, the predicted weather factor data, the photovoltaic output point prediction value, and a preset confidence level specifically includes:

[0037] The vine copula model is reversed to obtain the quantile regression expression of the photovoltaic output;

[0038] The predicted weather factor data, the predicted value of the photovoltaic output point, and a preset confidence level are input into the quantile regression expression, and a photovoltaic output confidence interval is output.

[0039] The present invention also provides a photovoltaic output probability prediction device, comprising:

[0040] An acquisition module is used to obtain historical weather factor data, predicted weather factor data, and historical photovoltaic output data of the area to be predicted;

[0041] A first determination module is configured to determine a photovoltaic output point prediction value based on the historical weather factor data, the predicted weather factor data, and the historical photovoltaic output data;

[0042] A calculation module, configured to calculate the correlation coefficient and collinearity strength between the historical weather factor data, the historical photovoltaic output data, and the photovoltaic output point prediction value;

[0043] a second determining module, determining a target weather variable from the weather factors according to the correlation coefficient and the collinearity strength;

[0044] A construction module, using the target weather variable, the photovoltaic output, and the photovoltaic output point prediction value to construct a vine copula model;

[0045] The prediction module is used to determine the photovoltaic output confidence interval based on the vine copula model, the predicted weather factor data, the photovoltaic output point prediction value, and a preset confidence level.

[0046] Optionally, the first determining module includes:

[0047] an alignment module, configured to synchronize the historical weather factor data, the predicted weather factor data, and the historical photovoltaic output data in a time dimension;

[0048] a partitioning module, configured to partition the synchronized and aligned historical weather factor data, the predicted weather factor data, and the historical photovoltaic output data into a training set, a validation set, and a set to be predicted according to a preset partitioning ratio;

[0049] The training module is used to use the training set and the verification set to train the pre-built photovoltaic output point prediction model, and input the set to be predicted into the trained photovoltaic output point prediction model to output the photovoltaic output point prediction value.

[0050] The present invention further provides an electronic device, comprising a processor and a memory:

[0051] The memory is used to store program code and transmit the program code to the processor;

[0052] The processor is configured to execute the method described above according to the instructions in the program code.

[0053] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store program code, and the program code is used to execute the method described above.

[0054] It can be seen from the above technical solutions that the present invention has the following advantages:

[0055] The present invention provides a photovoltaic output probability prediction method, comprising: obtaining historical weather factor data, predicted weather factor data, and historical photovoltaic output data of a to-be-predicted area; determining a photovoltaic output point prediction value based on the historical weather factor data, the predicted weather factor data, and the historical photovoltaic output data; calculating a correlation coefficient and collinearity strength between photovoltaic output, weather factors, and the photovoltaic output point prediction value using the historical weather factor data, the historical photovoltaic output data, and the photovoltaic output point prediction value; determining a target weather variable from the weather factors based on the correlation coefficient and the collinearity strength; constructing a vine copula model using the target weather variable, the photovoltaic output, and the photovoltaic output point prediction value; and determining a photovoltaic output confidence interval based on the vine copula model, the predicted weather factor data, the photovoltaic output point prediction value, and a preset confidence level.

[0056] In the present invention, by acquiring historical weather factor data, predicted weather factor data, and historical photovoltaic output data of the area to be predicted, and determining the predicted value of the photovoltaic output point based on the historical weather factor data, the predicted weather factor data, and the historical photovoltaic output data, the prediction of the photovoltaic output point is achieved; and by using the historical weather factor data, the historical photovoltaic output data, and the predicted value of the photovoltaic output point, the correlation coefficient and the collinearity strength between the photovoltaic output, weather factors, and the predicted value of the photovoltaic output point are calculated, thereby determining the degree of correlation between the photovoltaic output, weather factors, and the predicted value of the photovoltaic output point; and according to the correlation coefficient and the collinearity strength, the target weather variable is determined from the weather factors, thereby achieving the screening of weather factors. Thus, weather factors with a high degree of correlation with photovoltaic output are retained; and the target weather variable, the photovoltaic output, and the photovoltaic output point prediction value are used to construct a vine copula model, thereby realizing the construction of a photovoltaic output correlation model with the strongest multi-factor dependency structure, which can simultaneously characterize the correlation between photovoltaic output and multiple influencing factors, and provide strong technical support for more accurate photovoltaic output probability prediction; and according to the vine copula model, the predicted weather factor data, the photovoltaic output point prediction value, and the preset confidence level, the photovoltaic output confidence interval is determined, thereby obtaining a photovoltaic probability prediction result including the photovoltaic output confidence interval and the photovoltaic output point prediction value, realizing the prediction of photovoltaic output probability and improving the accuracy of the photovoltaic output probability prediction result. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description only constitute some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0058] Figure 1 A step flow chart of a photovoltaic output probability prediction method provided by an embodiment of the present application;

[0059] Figure 2 Another step flow chart of a photovoltaic output probability prediction method provided by an embodiment of the present application;

[0060] Figure 3 A Kendall correlation coefficient thermodynamic diagram provided by an embodiment of the present application;

[0061] Figure 4 A VIF multicollinearity feature analysis data diagram provided by an embodiment of the present application;

[0062] Figure 5 An effect diagram of a probability integral transform provided by an embodiment of the present application;

[0063] Figure 6 An optimization process schematic diagram of finding the optimal connection relationship between variables by using Q-Learning provided by an embodiment of the present application;

[0064] Figure 7 A structural schematic diagram of a C vine Copula provided by an embodiment of the present application;

[0065] Figure 8 A conditional probability distribution calculation path diagram of a C vine Copula provided by an embodiment of the present application;

[0066] Figure 9 A structural schematic diagram of a vine Copula model provided by an embodiment of the present application;

[0067] Figure 10 A probability prediction result schematic diagram of the next six days under a given confidence level of 95% provided by an embodiment of the present application;

[0068] Figure 11 A flow schematic diagram of a photovoltaic output probability prediction method provided by an application of the present application;

[0069] Figure 12 A structural schematic diagram of a photovoltaic output probability prediction device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0070] Existing photovoltaic output probability prediction methods, such as Amazon (Salinas D, Flunkert V, Gasthaus J, et al. DeepAR: Probabilistic forecasting with autoregressiverecurrent networks[J]. International Journal of Forecasting, 2020, 36(3):1181-1191.), proposed using long short-term memory artificial neural networks (LSTM) to predict time series, and constructed a negative logarithmic loss function to optimize model parameters. The probability distribution of each prediction time step is given, and the value with the highest probability is selected as the deterministic prediction. The variance is used to represent the uncertainty. This has been verified in actual commercial applications at Amazon. Researchers from the University of La Rioja, Spain, used a semiparametric method to obtain the parameters of the probability density function of hourly photovoltaic output (Findings from University of La Rioja in the Area of ​​Energy Described (Short-term Probabilistic Forecasting Models Using Beta Distributions for Photovoltaic Plants)[J]. Electronics Newsweekly, 2023.), thereby obtaining probabilistic forecast results and applying them to photovoltaic power plants in Galicia. Researchers from the University of Texas used pinball loss as the loss function to train multi-model machine learning to predict the distribution parameters at each time point (M. Sun, C. Feng, J. Zhang, EK Chartan and B. -M. Hodge, "Probabilistic Short-term Wind Forecasting Based on Pinball Loss Optimization," 2018 IEEE International Conference on Probabilistic Methods Applied to Power Systems (PMAPS), Boise, ID, USA, 2018, pp. 1-6), and applied this method to wind power forecasting.

[0071] From the above, it can be seen that the current photovoltaic power output probability prediction method mostly uses a non-parametric method to obtain distribution parameters of a prediction target according to historical data, and is often combined with a machine learning method, which on the one hand ignores the nonlinear correlation between the photovoltaic power output and multiple influencing factors, and the obtained result can only reflect part of the uncertain information; on the other hand, the overfitting and quantile crossing problems of the machine learning often highly affect the accuracy of the result. With the improvement of the power quality requirement in the distribution network and the increase of the photovoltaic penetration rate, especially the sudden change of the photovoltaic power output, the influence on the user side is getting larger and larger, and in the day-ahead scheduling, an accurate photovoltaic power output result is needed to provide a reasonable optimization scheme to avoid the occurrence of power reverse sending, voltage out-of-limit and other situations.

[0072] Therefore, based on the above prior art, the defects existing in the prior art include:

[0073] 1) The photovoltaic power output prediction method currently applied in the day-ahead optimization scheduling field of the photovoltaic distribution network ignores the correlation between the photovoltaic power output and the influencing factors, and the prediction curve cannot represent the uncertainty thereof.

[0074] 2) The probability prediction method based on a large model deep learning has long calculation time and many parameters, and is prone to overfitting, quantile crossing and other problems.

[0075] 3) The correlation modeling of the photovoltaic influencing factors is relatively subjective, and the existing Vine Copula modeling process is prone to local optimization and cannot maximize the correlation of the model.

[0076] Therefore, the embodiment of the present application provides a photovoltaic power output probability prediction method and related device to improve the accuracy of photovoltaic power output probability prediction.

[0077] In order to make the invention purpose, features and advantages of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the embodiments described below are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0078] Please refer to Figure 1 The present application provides a photovoltaic power output probability prediction method, which comprises:

[0079] 101. Obtain historical weather factor data, predicted weather factor data and historical photovoltaic power output data of a region to be predicted.

[0080] It should be noted that in this embodiment, historical weather factor data and predicted weather factor data are obtained from the NWP system of the measured area. The historical weather factor data refers to weather factor data for historical dates based on the predicted date. The predicted weather factor data refers to weather factor data for the predicted date. The historical photovoltaic output data is photovoltaic output data read from smart meters in photovoltaic power plants. It is understood that both the historical weather factor data and the predicted weather factor data include influencing factors such as wind speed, temperature, humidity, total solar irradiance, diffuse irradiance, wind direction, and rainfall.

[0081] 102. Determine the predicted value of the photovoltaic output point based on historical weather factor data, predicted weather factor data, and historical photovoltaic output data.

[0082] This embodiment constructs a neural network model, uses historical weather factor data and historical photovoltaic output data to train the neural network point prediction model, and uses the predicted weather factor data and the trained neural network point prediction model to perform point prediction of photovoltaic output, outputting estimated photovoltaic output values ​​at each moment of the day to be predicted, that is, obtaining photovoltaic output point prediction values.

[0083] It can be understood that the photovoltaic output estimation value output by the neural network point prediction model in this embodiment is a point value. Therefore, the photovoltaic output estimation value output by the neural network point prediction model is a photovoltaic output point prediction value.

[0084] 103. Using historical weather factor data, historical photovoltaic output data, and photovoltaic output point forecast values, calculate the correlation coefficient and collinearity strength between photovoltaic output, weather factors, and photovoltaic output point forecast values.

[0085] It should be noted that the correlation coefficient may be a Kendall correlation coefficient, and the collinearity strength may be a variance inflation factor value.

[0086] This embodiment calculates the correlation coefficient and collinearity strength between photovoltaic output, photovoltaic output point prediction value and each factor in weather factors by using historical weather factor data, historical photovoltaic output data and photovoltaic output point prediction value.

[0087] 104. Determine the target weather variables from the weather factors based on the correlation coefficient and collinearity strength.

[0088] It should be noted that, according to the correlation coefficient and the phase strength calculated in step 103 , the weather factor with the greatest correlation with the photovoltaic output is determined, that is, the target weather variable is obtained.

[0089] 105. The target weather variables, photovoltaic output, and photovoltaic output point forecast values ​​are used to construct the Vine Copula model.

[0090] It should be noted that the vine copula model constructed in this embodiment is C-vine copula (C-vine), where the dimension of the vine copula model is equal to the sum of the number of factors in the target weather variable plus two.

[0091] This embodiment uses photovoltaic output and photovoltaic output point forecast values ​​to screen out target weather variables with strong correlation, constructs a vine copula model, and realizes the construction of a photovoltaic output correlation model with the strongest dependency structure.

[0092] 106. Determine the photovoltaic output confidence interval based on the Vine Copula model, predicted weather factor data, photovoltaic output point prediction value, and preset confidence level.

[0093] In this embodiment, by using the predicted weather factor data and photovoltaic output point data, and combining it with the vine Copula model, the probability information of photovoltaic output can be obtained. According to the preset confidence level and the probability information of photovoltaic output, the photovoltaic output confidence interval can be obtained, and the photovoltaic output confidence interval and the photovoltaic point prediction value are used as the photovoltaic output probability prediction result.

[0094] It should be noted that the purpose of the PV probability forecast in this embodiment is to obtain the probability information of PV output at each predicted time on the predicted day. Since the probability information is represented by a probability distribution, the PV output probability information at the predicted time can be expressed by plotting the upper and lower bounds of the confidence interval. Therefore, this embodiment combines the PV output confidence interval with the PV point prediction value to obtain the PV output probability forecast result.

[0095] In this embodiment, by obtaining historical weather factor data, predicted weather factor data, and historical photovoltaic output data of the area to be predicted, and determining the predicted value of the photovoltaic output point based on the historical weather factor data, predicted weather factor data, and historical photovoltaic output data, the prediction of the photovoltaic output point is realized; and using the historical weather factor data, the historical photovoltaic output data, and the predicted value of the photovoltaic output point, the correlation coefficient and the collinearity strength between the photovoltaic output, weather factors, and the predicted value of the photovoltaic output point are calculated, thereby realizing the determination of the degree of correlation between the photovoltaic output and the weather factors; and according to the correlation coefficient and the collinearity strength, the target weather variable is determined from the weather factors, thereby realizing the prediction of the weather factors. The results show that the proposed method can effectively prevent the occurrence of PV output errors and the prediction accuracy of photovoltaic power generation. The method can effectively prevent the occurrence of PV output errors and the prediction accuracy of photovoltaic power generation. The method can effectively prevent the occurrence of PV output errors and the prediction accuracy of photovoltaic power generation. The method can effectively prevent the occurrence of PV output errors and the prediction accuracy of photovoltaic power generation. The method can effectively prevent the occurrence of PV output errors and the prediction accuracy of photovoltaic power generation.

[0096] See also Figure 2 , an embodiment of the present invention provides a photovoltaic output probability prediction method, comprising:

[0097] 201. Obtain historical weather factor data, predicted weather factor data, and historical photovoltaic output data for the area to be predicted.

[0098] It should be noted that step 201 may refer to step 101 and will not be described in detail here.

[0099] 202. Synchronize and align historical weather factor data, predicted weather factor data, and historical photovoltaic output data in the time dimension.

[0100] It should be noted that the historical weather factor data, predicted weather factor data, and historical photovoltaic output data obtained in this embodiment are synchronized and aligned in the time dimension so that each set of data has a corresponding correlation relationship in time.

[0101] It's important to note that PV output data from smart meters and weather data from the NWP system typically have different time intervals. Synchronous alignment involves determining an appropriate time interval based on the different time intervals of each data point, sampling all data points to the same time interval, ensuring consistency between the previous and next sampling moments. For example, data with a short time interval (such as 1 minute) can be downsampled to a longer time interval (5 minutes) to maintain temporal consistency with the data with the longer time interval. This step ultimately aligns the different types of data so that each moment corresponds to a set of historical weather factor data, forecasted weather factor data, and historical PV output data.

[0102] 203. According to a preset division ratio, the historical weather factor data, the predicted weather factor data, and the historical photovoltaic output data after synchronization and alignment are divided into a training set, a validation set, and a set to be predicted.

[0103] It should be noted that the division ratio can be determined according to actual needs. In particular, the historical weather factor data and the historical photovoltaic processing data are used to divide the training set and the validation set, and the predicted weather factor data is divided to obtain the prediction set.

[0104] 204. Using the training set and the validation set, the pre-built photovoltaic output point prediction model is trained, and the set to be predicted is input into the trained photovoltaic output point prediction model to output a photovoltaic output point prediction value.

[0105] It should be noted that the input of the photovoltaic point prediction model is weather factor data, and the output is the photovoltaic output point prediction value. The photovoltaic output point prediction model in this embodiment adopts an LSTM network.

[0106] This example uses a training set and a validation set to train and validate the LSTM network, optimizing its hyperparameters and completing the training of the LSTM network. Once the LSTM network is trained, the prediction set is fed into the LSTM network, which outputs the predicted PV output point for the predicted date.

[0107] Specifically, the LSTM network inputs weather factor data for the day to be predicted and historical PV output data corresponding to the day to be predicted. The historical dates of the input PV data can be determined based on the day to be predicted. For example, if the day to be predicted is one day, the input is the historical PV data from the day before the day to be predicted. If the day to be predicted is two days, the input is the historical PV output data from the two days before the day to be predicted. For example, if the day to be predicted is June 22nd and June 23rd, the input is the historical PV data from June 20th and June 21st, and so on.

[0108] The output of the LSTM network is the predicted value of the photovoltaic output point at each time of the day to be predicted.

[0109] It should be noted that the day to be predicted can be divided into multiple time steps (i.e., prediction times), for example, the time interval is set to 30 min, and the time is divided according to the set time interval, so one day can be divided into 48 time points (i.e., 48 time points). Therefore, the data input into the LSTM network in this embodiment includes multiple time steps, and the LSTM network outputs the point prediction value at multiple times. Among them, the LSTM network uses a recursive prediction method to construct a sliding window to predict the future multi-step photovoltaic output. The recursive prediction method refers to putting the prediction result of the last step into the data of the last time step needed to be input in the next step prediction, and then predicting again to obtain the prediction value of multiple steps.

[0110] 205. Using the historical weather factor data, the historical photovoltaic output data, and the photovoltaic output point prediction value, the correlation coefficient and the collinearity strength between the photovoltaic output, the weather factor, and the photovoltaic output point prediction value are calculated.

[0111] It should be noted that step 205 specifically includes the following sub-steps:

[0112] S31. The historical weather factor data, the historical photovoltaic output data, and the photovoltaic output point prediction value are normalized respectively.

[0113] S32. Based on the normalized historical weather factor data, the historical photovoltaic output data, and the photovoltaic output point prediction value, the Kendall correlation coefficient and the variance inflation factor between the photovoltaic output, the weather factor, and the photovoltaic output point prediction value are calculated.

[0114] S33. The Kendall correlation coefficient is taken as the correlation coefficient, and the variance inflation factor is taken as the collinearity strength.

[0115] It should be noted that in this embodiment, the photovoltaic output point prediction value output by the LSTM network after the hyperparameter optimization is completed is taken as the collaborative variable to determine the weather factor with high correlation degree with the photovoltaic output.

[0116] This embodiment uses the Kendall correlation coefficient to quantify the correlation between the photovoltaic output and multiple influencing factors, obtains the selected influencing factors with high correlation degree, and uses the variance inflation factor (variance inflation factor, VIF) to determine the multicollinearity between the photovoltaic output and other factors. For subsequent vine Copula correlation modeling, provide variables with high correlation and collinearity.

[0117] Specifically, the Kendall rank correlation coefficient (Kendall rank correlation coefficient), also known as the Kendall rank correlation coefficient, is a statistic used to measure the rank correlation between two variables and can be used to reflect the nonlinear relationship between the two variables. Its calculation principle is as follows: first, the data of the two variables are sorted into matrices to obtain the rank of each data point in their respective variables; then, the rank relationships of the corresponding data points in the two variables are compared to see if they are consistent, and the number of consistent and inconsistent data points is recorded. Based on this number of consistent and inconsistent data points, the Kendall correlation coefficient of the two variables is calculated.

[0118] In the specific implementation, the relevant calculation formula of Kendall's correlation coefficient is as follows:

[0119]

[0120] Where, Indicates the number of variables where any two elements have the same order; It indicates the number of inconsistencies; Indicates the number of elements.

[0121] In this embodiment, the relationship between weather factors and photovoltaic output is nonlinear, so the Kendall correlation coefficient can be used to measure the correlation between factors. This embodiment calculates the Kendall correlation coefficients between historical photovoltaic output, photovoltaic output point forecast value, and various influencing factors, and constructs a Kendall correlation coefficient heat map based on the calculated Kendall correlation coefficients. The Kendall correlation coefficient heat map is shown in the figure below. Figure 3 The color depth represents the correlation between factors. The larger the value and the brighter the color, the greater the correlation. Figure 3 In this data, PV output refers to historical PV output. Point forecast value refers to the point forecast value of PV output.

[0122] Specifically, the VIF method is a method used to evaluate the linear relationship between model input variables. Its basic idea is to calculate the variance inflation factor value of each feature and other features to feedback the collinearity strength of each feature with other features. The larger the VIF value, the stronger the correlation, which feedbacks the existence of multicollinearity between the calculated variables. Among them, the calculation principle of VIF is to fit a linear regression model with each feature as the dependent variable and other features as independent variables, and then calculate the mean square error ratio of the independent variable and the dependent variable, that is:

[0123]

[0124] Where, The goodness of fit of the relationship between the independent variable and other independent variables is measured by a simple linear regression model.

[0125] This embodiment uses the least squares method as the linear regression in the VIF method, which estimates the regression coefficient by minimizing the sum of squared errors. It is easy to understand and implement, has high computational efficiency, and can provide a clear regression coefficient.

[0126] In this implementation, PV output is used as the independent variable, and other variables except PV output are used as dependent variables. The corresponding variance inflation factors are calculated respectively, and a multicollinearity analysis histogram is constructed based on the obtained variance inflation factors. Figure 4 As shown, Figure 4 In the figure, the longer the column, the higher the degree of collinearity with PV output, and also indicates that the correlation between the variable and other variables besides PV output is also high.

[0127] 206. Determine the target weather variables from the weather factors based on the correlation coefficient and collinearity strength.

[0128] It should be noted that, based on the correlation coefficient and collinearity strength calculated in step 205, influencing factors with greater correlation and collinearity can be screened out from various influencing factors. Among them, the screening conditions can be set according to the dimensions of the vine copula model to be constructed. For example, when establishing a 5-dimensional C vine copula model, the correlation coefficient and collinearity strength are sorted from large to small, and according to the sorting order, the top four influencing factors other than photovoltaic output are selected. However, in the process of calculating the correlation between photovoltaic output and multiple influencing factors using the Kendall coefficient and observing the multicollinearity between photovoltaic output and other factors using the variance inflation factor, it is found that when the photovoltaic output point prediction value is used as an influencing factor, the correlation with photovoltaic output is extremely high. Therefore, the photovoltaic output point prediction value is used as one of the influencing factors, and the remaining three influencing factors are weather factors that are ranked higher. Based on this principle, this embodiment can determine the corresponding target weather variable from the weather factors.

[0129] Therefore, this embodiment calculates the Kendall correlation coefficient and variance inflation factor between photovoltaic output, weather factors, and photovoltaic output point prediction values ​​based on the normalized historical weather factor data, historical photovoltaic output data, and photovoltaic output point prediction values. Using the Kendall correlation coefficient as the correlation coefficient and the variance inflation factor as the collinearity strength can more clearly capture the correlation characteristics between multiple variables, thereby screening out several influencing factors with large correlation and collinearity, providing strong technical support for the construction of the vine copula model in steps 207-208.

[0130] Adopting multiple collinearity VIF will make the regression coefficient estimate unstable, and the interpretation ability of the conventional machine learning model decreases. However, the application innovatively uses the VIF method to screen variables with high degree of multicollinearity, and at the same time combines the Kendall correlation coefficient to judge the relationship between the variables and the photovoltaic output target, provides variables with high correlation degree for the establishment of the vine Copula model, so that the established vine Copula model can better depict the dependence relationship between multiple factors, and improve the accuracy of photovoltaic output probability prediction.

[0131] 207、Adopting target weather variables, photovoltaic output, and photovoltaic output point prediction value to construct a vine Copula model.

[0132] It should be noted that step 207 specifically includes the following substeps:

[0133] S51, using a non-parametric kernel density estimation method, respectively performing probability integral transformation on the target weather variables, photovoltaic output, and photovoltaic output point prediction value to obtain the marginal probability distribution of the target weather variables, the marginal probability distribution of the photovoltaic output, and the marginal probability distribution of the photovoltaic output point prediction value.

[0134] It should be noted that the non-parametric kernel density estimation method is a non-parametric statistical method for estimating the probability density function of a random variable, and its core idea is to estimate the probability density function by weighted average of the local region around the data points. Among them, the probability integral transformation is a data transformation process in probability theory, which can map the integrable region of the distribution function of the data to the unit interval (0, 1), so that the variable after transformation obeys uniform distribution.

[0135] In a specific implementation, the calculation formula of the non-parametric kernel density estimation method is as follows:

[0136]

[0137]

[0138] In the formula, indicates the marginal distribution estimation value of the random variable x; N is the total number of samples; indicates the nth sample value of the random variable x; indicates the Gaussian kernel function.

[0139] In this embodiment, the non-parametric kernel density estimation method is used to perform probability integral transformation on each influencing factor in the target weather variables, photovoltaic output, and photovoltaic output point prediction value, to obtain the marginal probability distribution of each influencing factor, the marginal probability distribution of photovoltaic output, and the marginal probability distribution of photovoltaic output point prediction value. Taking photovoltaic output as an example, as shown in Figure 5As shown, the left figure is the photovoltaic output histogram before the transformation, and the right figure is the probability distribution histogram after the transformation. From the transformation effect diagram, it can be seen that the transformation effect of this embodiment is significant.

[0140] S52. Use a reinforcement learning algorithm to determine the optimal connection relationship between the target weather variable, photovoltaic output, and photovoltaic output point prediction value.

[0141] It's important to note that the reinforcement learning algorithm, Q-Learning, is used to train intelligent agents to make decisions in a specific environment to maximize cumulative rewards. It works by continuously trying different actions, observing the rewards provided by the environment, and updating its Q-value table based on these rewards, gradually finding the optimal action selection strategy for each state. In this process, the agent both explores (trying new actions) to discover better strategies and exploits (selecting the currently considered optimal action) to obtain higher rewards. Based on this, the Q-Learning algorithm can gradually learn the optimal behavior strategy in an unknown environment to achieve the goal of maximizing cumulative rewards.

[0142] In this embodiment, a Q-value table is created and initialized to set the initial and final states. The number of rows in the Q-table corresponds to the number of states in the environment, and the number of columns corresponds to the number of actions. The Q function is used to represent the reward for performing an action in the current state, thereby obtaining the maximum Q-value for the entire process. For each step taken by the agent from the initial state to the final state, the Q-value of the different actions taken in the current state is recorded and recorded using a Q-table (i.e., a Q-value table). After reaching the final state, the agent's state is reset and exploration is resumed. Each row in the Q-table represents the current state, and each column represents an optional action. Each cell in the table stores the new reward value resulting from the current state corresponding to the future action. The Q-table is updated using the following formula:

[0143]

[0144] Where, Indicates status, Indicates action, is the learning rate, It represents the immediate reward obtained by the agent after taking an action. is the decay rate, Indicates that the agent is in state Next, we query the maximum value of the row in the Q-table. A represents the action space. and Represents the next state and the newly selected action. As can be seen from the above update formula, during the interaction between the agent and the environment, the update of the Q value in the current state utilizes the Q value of the next state.

[0145] Specifically, this embodiment first creates a Q-value table, with its rows and columns set to the target weather variable, PV output, and predicted PV output point values. The Q-value table is initialized, defining PV output as the initial state and the final state as the connection of all variables. The selected variable is used as the current state, and the remaining optional variables are used as actions. The action space is narrowed based on the generation characteristics of the largest vine tree. The Kendall correlation coefficients between the target weather variable, PV output, and predicted PV output point values ​​are then calculated. A Kendall correlation coefficient matrix is ​​established based on the calculated Kendall correlation coefficients. During the optimization process, the Kendall correlation coefficients between each variable are used as immediate rewards, and the Q-table is updated using the above formula. During the interactive process, the agent selects the variable with the highest Kendall correlation coefficient with the current state variable based on the coefficient matrix for connection. When the values ​​in the Q-table no longer change significantly, the agent is considered to have found the optimal variable connection relationship, and the optimization process ends.

[0146] Taking the construction of a 5-dimensional C-vine Copula model as an example, this step is further explained as follows:

[0147] In this embodiment, the target weather variables include: total solar radiation, wind speed, and temperature. Therefore, the variables for establishing the vine tree are photovoltaic output, photovoltaic output point prediction value, total solar radiation, wind speed, and temperature. Figure 6 As shown, the rows and columns in the Q-value table are: P (PV output), F (PV output point predicted value), R (global solar radiation), W (wind speed), and T (temperature). In the figure, CORR-MAT is the Kendall correlation coefficient matrix, which is composed of the Kendall correlation coefficients between the variables. First, taking PV output as the current state, the action space includes F (PV output point predicted value), R (global solar radiation), W (wind speed), and T (temperature). Then, F is selected to update the current state, and the Q-value table is updated using the Kendall matrix and the above update formula. This process is repeated iteratively to complete the Q-value table update, obtaining the connection relationship between the nodes with the highest correlation (i.e., the optimal connection relationship).

[0148] S53. Use the marginal probability distribution of the target weather variable, the marginal probability distribution of photovoltaic output, and the marginal probability distribution of the photovoltaic output point prediction value as vine tree nodes, respectively, and connect the vine tree nodes using the optimal connection relationship to obtain the target vine tree.

[0149] It should be noted that the vine tree nodes can be divided into root nodes and child nodes. In this embodiment, the marginal probability distribution of photovoltaic output is used as the root node of the target vine tree, and the marginal probability distribution of the target weather variable and the marginal probability distribution of the photovoltaic output point prediction value are used as child nodes. According to the optimal connection relationship obtained in the above steps, the root node and each child node are connected in sequence. The principle of vine tree construction is as follows: Figure 7As shown, the root node of vine tree 1 is variable 1, and the child nodes are variables 2 to 5. According to the order, variables 2 to 5 are connected to form vine tree 1.

[0150] This embodiment takes the construction of a 5-dimensional C-vine Copula model as an example. Figure 9 As shown, based on the optimal connection relationship obtained in step S2 above, the photovoltaic output P is used as the root node, and the photovoltaic output point predicted value F, total solar radiation R, wind speed W, and temperature T are sequentially connected to the photovoltaic output P, ​​forming the vine tree 1 in the figure. Each node is configured with the marginal probability distribution corresponding to each variable.

[0151] S54, using the maximum likelihood estimation method and the Akaike information criterion method to determine the optimal Copula function and the parameter value of the optimal Copula function for each edge in the target vine tree, and using the optimal Copula function, the parameter value of the optimal Copula function, and each vine tree node to respectively calculate the conditional probability of each edge in the target vine tree;

[0152] It should be noted that the edge of the target vine tree refers to the connecting line between two vine tree nodes, which is used to feedback the Copula function between the two vine tree nodes. The conditional probability of the edge in the target vine tree represents the conditional probability of it connecting the two vine tree nodes.

[0153] The Maximum Likelihood Estimation (MLE) method is a method that can calculate the parameters of a known distribution model that fits data and maximizes the probability of the assumed known distribution occurring. Assuming that various copula functions are set for each edge in the target vine tree, the MLE can be used to determine the parameter values ​​of each copula function for each edge in the target vine tree. The Akaike Information Criterion (AIC) can be used to measure the fitting effect of each copula function for each edge in the target vine tree, given the parameters of each copula function calculated above. Based on the fitting effect of each copula function, the optimal copula function and its parameter values ​​can be selected from the various copula functions to serve as the copula function and its parameter values ​​for that edge.

[0154] After determining the optimal Copula function and the optimal Copula function parameter values ​​for each edge, the conditional probability value of the edge is calculated using the corresponding optimal Copula function and the optimal Copula function parameter values, as well as the marginal probability distribution of the variables connected by the edge. The calculation principle is as follows:

[0155] It should be noted that the construction of a Vine Copula consists of three parts: a vine tree, a Copula function type, and function parameters. The vine tree influences the selection of function types and parameters. The embodiments of the present invention are based on the use of Vine Copula theory to establish a joint distribution of multiple variables. Vine Copula combines Copula theory with graph theory to decompose the multivariate joint probability distribution function into a series of conditional two-dimensional copulas in a cascade manner. According to the definition of the Copula function, the multivariate joint probability distribution is:

[0156]

[0157] in, To contain d-dimensional marginal distribution The joint distribution function of For the marginal cumulative distribution function of a variable; is a Copula function.

[0158] This embodiment constructs a C-vine Copula model, wherein the conditional probability distribution calculation path of C-vine Copula is as follows: Figure 8 As shown. Each edge in the vine tree is a Copula function of two node variables. When a conditional variable appears, the given conditional variable set can be Lower variable and variables The Copula function is Then their conditional distributions are for:

[0159]

[0160] In the above formula, represents the conditional distribution function, yes Eliminate variables The variable after represents the variable in the vine tree and variables The set of variables that are commonly connected. From the above formula for calculating conditional distribution, we can know that the conditional distribution of high-dimensional conditional set can be recursively calculated from low-dimensional conditional set.

[0161] After obtaining the optimal Copula function and function value for each edge, the corresponding conditional distribution (i.e., conditional probability value) can be calculated based on the aforementioned calculation formula and the corresponding edge probability distribution.

[0162] In a specific embodiment, the steps of determining the optimal Copula function and the parameter value of the optimal Copula function for each edge in the target vine tree using the maximum likelihood estimation method and the Akaike information criterion method include:

[0163] S541. Obtain multiple Copula functions for each edge in the target vine tree.

[0164] It should be noted that in this embodiment, multiple copula functions of different types are pre-set for each edge in the target vine tree, wherein the number of copula functions can be set according to the actual needs. At runtime, each copula function of each edge is obtained separately.

[0165] S542. Calculate the parameter value of each Copula function using the maximum likelihood estimation method.

[0166] It should be noted that the maximum likelihood estimation method (MLE) can transform the search for parameters of the data distribution into a non-parametric optimization problem, thereby achieving the estimation of the parameter value of each Copula function. The estimation formula is:

[0167]

[0168] Where, Indicates solving the parameter that makes the function output the maximum value value.

[0169] In this embodiment, by solving the parameter values ​​of each Copula function of each edge, a set of parameters of the Copula function of each edge on the vine tree can be obtained.

[0170] S543. Calculate the AIC value of each Copula function according to each Copula function and its parameter value.

[0171] S544. The Copula function with the smallest AIC value is used as the optimal Copula function, and the parameter value of the Copula function with the smallest AIC value is used as the function value of the optimal Copula function.

[0172] It should be noted that this embodiment uses the Akaike Information Criterion method to measure the fitting effect of each copula function (including the likelihood function of the parameter estimation result based on the MLE) on each edge. By comparing the sizes of the AIC values, the copula function with the smallest AIC value is selected as the optimal copula function to achieve the screening of copula functions.

[0173] The calculation formula of AIC value is:

[0174]

[0175] in, is the Copula function (i.e. likelihood function) under a certain parameter, and K is the number of parameters of the Copula function.

[0176] S55. Taking the conditional probability of each edge as the vine tree node of the next vine tree;

[0177] S56. Using a reinforcement learning algorithm, determine the optimal connection relationship between the vine tree nodes of the next vine tree;

[0178] S57. Connect the vine tree nodes of the next vine tree according to the optimal connection relationship between the vine tree nodes of the next vine tree to obtain the next vine tree.

[0179] It should be noted that in this embodiment, after obtaining the conditional probabilities of each edge, these are used as the vine nodes of the next vine tree. A reinforcement learning algorithm is then used to calculate the optimal connection relationships between the vine nodes of the next vine tree. Based on this optimal connection relationship, the vine nodes are connected to form a new vine tree, the structure of which is shown as vine tree 2 in the figure. The principles of steps S6 and S7 are similar to those of steps S2 and S3. For details, please refer to steps S2-S3 and will not be repeated here.

[0180] This step is further explained below using the 5-dimensional C-vine Copula model of the previous step as an example. The newly constructed vine tree is shown in vine tree 2, in which the conditional probability of photovoltaic output P and total solar radiation R is used as the root node, and the conditional probability between photovoltaic output P and photovoltaic output point predicted value F, the conditional probability between photovoltaic output P and wind speed W, and the conditional probability between photovoltaic output P and temperature T are used as child nodes.

[0181] S58. Update the next vine tree to the target vine tree in step S54, and jump to S59.

[0182] It should be noted that, in this embodiment, the new vine tree obtained in step S57 replaces the target vine tree in step S54.

[0183] S59. Loop through steps S54 to S58 until the updated target vine tree in step S58 has only one edge. Then, stop the loop and output the target vine tree and each updated target vine tree.

[0184] It should be noted that in this embodiment, steps S54 to S58 are executed on the updated target vine tree (i.e., the next vine tree in step S58) to construct a new vine tree. This process is iterated until the target vine tree has only one edge, indicating that the highest-level vine tree has been obtained. Therefore, the initial target vine tree and the new target vine tree obtained with each iteration are output. The vine copula model constructed in this embodiment includes the initial target vine tree and the new target vine tree obtained with each iteration. Therefore, this embodiment achieves the construction of the vine copula model through the above iterative steps.

[0185] This step is further explained below by taking the 5-dimensional C-vine Copula model of the previous step as an example.

[0186] like Figure 9 As shown in the figure, the constructed vine copula model includes vine trees 1, vine trees 2, vine trees 3, and vine trees 4. The edges of each vine tree are labeled with the corresponding optimal copula function and the parameter value of the optimal copula function, which can describe the dependency relationship between each pair of variables and complete the construction of the photovoltaic output correlation model with the strongest dependency structure.

[0187] 208. Determine the photovoltaic output confidence interval based on the Vine Copula model, predicted weather factor data, photovoltaic output point prediction value, and preset confidence level.

[0188] It should be noted that this embodiment gradually solves the conditional distribution of photovoltaic output for the Vine Copula model to obtain the corresponding quantile regression expression, and uses the quantile regression expression to obtain the conditional probability information of photovoltaic output on the day to be predicted and the predicted photovoltaic output point prediction value, and determines the photovoltaic output confidence interval and its corresponding photovoltaic output point prediction value according to the preset confidence level to obtain the photovoltaic output probability prediction result with multi-factor correlation.

[0189] The principle of this step is:

[0190] Given the predicted value F of the photovoltaic output point, temperature, total solar radiation, and wind speed influencing variables, find the known conditions 、 、 、 Conditional probability distribution of the actual photovoltaic output P under , and the Monte Carlo simulation method is used to generate the real value samples of photovoltaic output during the forecast period. Among them, the inverse function of its conditional distribution needs to be calculated, and its calculation formula is as follows:

[0191]

[0192] The formula reflects the effect of removing a variable in the high-dimensional conditional distribution Therefore, according to the vine copula connection relationship established in the previous steps, the variables associated with it can be gradually eliminated from the high-dimensional conditional distribution, and finally the required probability distribution function of the actual photovoltaic output can be obtained. Thus, this embodiment is a probability prediction result obtained by considering the correlation of multiple factors and considering the marginal distribution of photovoltaic output under given other influencing variables.

[0193] Step 208 specifically includes the following sub-steps:

[0194] S61. Reverse the Vine Copula model to obtain the quantile regression expression of photovoltaic output;

[0195] It should be noted that this embodiment uses the h-inverse function to solve the inverse of the photovoltaic power conditional distribution under given other conditions, thereby obtaining an analytical expression of the quantile. The h-inverse function is the inverse function of the h-function with respect to the first variable. It can invert the conditional distribution, as shown in the following formula:

[0196]

[0197]

[0198] As can be seen from the above formula, contrary to the h function, the h inverse function can eliminate variables in high-dimensional conditional distribution , we get the conditional distribution after reducing one dimension From this we can see that the h-inverse function is actually the reverse deduction of the vine copula, gradually eliminating the conditional variables starting from the highest dimension to obtain the variables that need to be simulated.

[0199] Taking the 5-dimensional C-vine Copula model in the previous step as an example, the 5-dimensional C-vine Copula model is derived as follows:

[0200]

[0201] in, Indicates the first dimensional variables, and .

[0202] By deducing the above Cvine Copula, the probability distribution function of the actual photovoltaic output (i.e., the quantile regression expression) can be obtained as follows:

[0203]

[0204] As can be seen from the above, the solution of the h inverse function is dynamic, and the value obtained from the high dimension at each step is required to solve the low-dimensional variables until the conditional variables are completely eliminated. is the conditional probability of the input, also known as the quantile , which is used to express the expected quantile regression expression output of the photovoltaic quantile at a certain moment. Given each conditional variable And enter different quantiles To the quantile regression expression, the quantile regression value of photovoltaic power can be calculated For example, if you want the quantile regression expression to output the 90% quantile at a certain moment, then set The corresponding predicted weather factor data and the predicted value of the photovoltaic output point are input into the above quantile regression expression. The output photovoltaic power quantile regression value represents the 90% quantile of the photovoltaic output at that moment.

[0205] S62: Input the predicted weather factor data, the predicted value of the photovoltaic output point, and the preset confidence level into the quantile regression expression, and output a photovoltaic output confidence interval.

[0206] It should be noted that in this embodiment, the forecast set and the predicted PV output point values ​​from step 202 can be input into the quantile regression expression. The forecast set includes the NWP forecast value for the forecast day. The preset confidence level is used to determine the desired upper and lower quantiles. Therefore, as described in step S61 above, the preset confidence level, the predicted weather factor data, the predicted PV output point values, and the quantile regression expression are used to determine the quantile critical values ​​(i.e., the upper and lower bounds of the confidence interval) for each moment on the forecast day. This allows the PV output confidence interval to be determined based on the quantile critical values.

[0207] For example, the confidence level can be set to 95%, with the corresponding upper and lower quantiles being 97.5% and 2.5% respectively. Therefore, the regression values ​​output by the quantile regression expression are the 97.5th and 2.5th percentile values ​​of all simulated values ​​of PV output at each moment of the predicted day, sorted from small to large.

[0208] Therefore, this embodiment obtains the NWP forecast value of the day to be predicted, as well as the predicted value of the PV output point, and inputs the preset confidence level into the quantile regression expression derived in step S61, and outputs the upper and lower bounds of the confidence interval at each moment in the prediction period, thereby obtaining the PV output confidence interval.

[0209] In a simulation application example, this application example obtains 6 months of historical NWP data and photovoltaic output data, and obtains 6 months of photovoltaic output point prediction training set data based on the 6 months of historical NWP data and photovoltaic output data. By creating a 5-dimensional vine Copula model of photovoltaic actual output, photovoltaic output point prediction value, solar total irradiance, temperature, and wind speed, the conditional distribution expression of photovoltaic actual output is solved. At a confidence level of 95%, the NWP data for the next 6 days and the next 6-day point prediction value obtained by LSTM prediction are input into the conditional distribution expression (quantile regression expression) to obtain the photovoltaic output confidence interval. The photovoltaic output confidence interval and photovoltaic output point prediction (i.e., photovoltaic probability prediction result) are plotted as shown below. Figure 10 shown.

[0210] In summary, the present invention can ultimately establish a photovoltaic output probability prediction result that takes into account the correlation of multiple factors. When the confidence interval output by the model can effectively cover the true value and the interval width is appropriate, then theoretically this method can be used to infer the operating trend of photovoltaic power generation in a certain distribution network area, so that the distribution network can be managed daily more scientifically, preventing the impact of photovoltaic and other new energy access on the distribution network, reducing user complaint rates, and providing effective guarantees for the economic, safe and stable operation of the distribution network containing distributed photovoltaics.

[0211] In an application example, in actual operation, see Figure 11 The photovoltaic output probability prediction method provided by the embodiment of the present invention can be divided into deterministic prediction, variable screening, vine tree establishment, vine Copula model establishment, conditional distribution acquisition and simulation.

[0212] The deterministic prediction includes obtaining NWP data for historical and forecast days, preprocessing historical PV output data recorded by smart meters, dividing the training set according to a preset ratio, and inputting it into the LSTM network for deterministic prediction to obtain the predicted value of the PV output point on the forecast day.

[0213] Variable screening includes: using the Kendall correlation coefficient to screen factors that are strongly correlated with PV output, and retaining highly collinear characteristics based on VIF.

[0214] The establishment of the vine tree includes: performing probability integral transformation on the screening variables to obtain the marginal probability distribution; using Q-Learning to determine the variable connection relationship with the maximum correlation, and using the marginal probability distribution of the variables as nodes to determine the form of the vine tree.

[0215] The establishment of the vine copula model includes: using MLE and AIC to calculate the Copula function parameters and types of the two variables connecting each edge in the vine tree; calculating the conditional probability value on each edge tree by tree and using it as the node of the next vine tree; and again determining the maximum correlation connection relationship of the new node based on Q-Learning to form a complete photovoltaic output vine copula model conditional distribution acquisition: gradually solving the conditional distribution of photovoltaic output for the vine copula model to obtain the quantile regression expression.

[0216] The conditional distribution acquisition includes: gradually solving the conditional distribution of photovoltaic output for the vine copula model to obtain the quantile regression expression.

[0217] The simulation includes: obtaining the probability information of photovoltaic output conditions on the day to be predicted based on the NWP information of the day to be predicted and the vine copula model, obtaining the confidence interval according to the preset confidence level, and obtaining the photovoltaic output probability prediction result considering the correlation of multiple factors.

[0218] See Figure 12, an embodiment of the present invention provides a photovoltaic output probability prediction device, comprising:

[0219] An acquisition module 301 is used to acquire historical weather factor data, predicted weather factor data, and historical photovoltaic output data of the area to be predicted;

[0220] A first determination module 302 is configured to determine a photovoltaic output point prediction value based on historical weather factor data, predicted weather factor data, and historical photovoltaic output data;

[0221] The calculation module 303 is used to calculate the correlation coefficient and collinearity strength between the photovoltaic output, weather factors and photovoltaic output point prediction values ​​based on historical weather factor data, historical photovoltaic output data and photovoltaic output point prediction values;

[0222] The second determination module 304 determines the target weather variable from the weather factors according to the correlation coefficient and the collinearity strength;

[0223] Building module 305, constructing a vine copula model using target weather variables, photovoltaic output, and photovoltaic output point prediction values;

[0224] The prediction module 306 is used to determine the photovoltaic output confidence interval based on the vine copula model, the predicted weather factor data, the photovoltaic output point prediction value, and the preset confidence level.

[0225] Furthermore, the first determining module 302 includes:

[0226] The alignment module is used to synchronize historical weather factor data, forecast weather factor data, and historical photovoltaic output data in the time dimension;

[0227] A partitioning module is used to divide the synchronized historical weather factor data, predicted weather factor data, and historical photovoltaic output data into a training set, a validation set, and a set to be predicted according to a preset partitioning ratio;

[0228] The training module is used to train the pre-built photovoltaic output point prediction model using the training set and the validation set, and input the set to be predicted into the trained photovoltaic output point prediction model to output the photovoltaic output point prediction value.

[0229] The present invention further provides an electronic device, comprising a processor and a memory:

[0230] The memory is used to store program codes and transmit the program codes to the processor;

[0231] The processor is configured to execute the method of any of the above embodiments according to the instructions in the program code.

[0232] The present invention also provides a computer-readable storage medium, which is used to store program code, and the program code is used to execute the method of any of the above embodiments.

[0233] Based on the above, the present invention provides more accurate and effective photovoltaic output prediction results, characterizes the uncertainty impact caused by multi-factor correlation, and provides effective output probability information and intervals to quantify photovoltaic random characteristics; and the present invention realizes multi-factor correlation modeling of photovoltaic output, realizes visualization and maximization of dependent structure; the present invention constructs a high-dimensional joint distribution of photovoltaic output through correlation modeling, which can be directly applied to sampling, simulation, and generation of photovoltaic output intervals covering multi-factor related information. In addition, the present invention more accurately estimates the output of photovoltaic power generation for distributed photovoltaic distribution networks, predicts various photovoltaic output conditions in the future in advance during day-ahead scheduling, reasonably regulates the conservatism of scheduling optimization schemes and determines planning boundaries, optimizes and controls the output of generator sets and other power resources to reduce power losses and lower operating costs; at the same time, it can reduce the impact of extreme weather on the power grid through probabilistic prediction, help distribution network operators to make scheduling and management in advance, and reduce power grid instability caused by photovoltaic power generation fluctuations.

[0234] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0235] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0236] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0237] In addition, the functional units in various embodiments of the present invention may be integrated into a single processing unit, or each functional unit may exist as a separate physical unit, or two or more functional units may be integrated into a single processing unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0238] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0239] The terms "first," "second," "third," "fourth," and the like (if any) in the specification of the present application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present application described herein can, for example, be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product, or apparatus.

[0240] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A photovoltaic output probability prediction method, characterized in that: include: Obtain historical weather factor data, forecast weather factor data, and historical photovoltaic output data for the area to be predicted; Determining a photovoltaic output point prediction value based on the historical weather factor data, the predicted weather factor data, and the historical photovoltaic output data; Calculate the correlation coefficient and collinearity strength between photovoltaic output, weather factors, and the photovoltaic output point prediction value using the historical weather factor data, the historical photovoltaic output data, and the photovoltaic output point prediction value; determining a target weather variable from the weather factors according to the correlation coefficient and the collinearity strength; Constructing a Vine Copula model using the target weather variable, the photovoltaic output, and the photovoltaic output point prediction value; Determining a photovoltaic output confidence interval based on the vine copula model, the predicted weather factor data, the photovoltaic output point prediction value, and a preset confidence level; The method of constructing a Vine Copula model using the target weather variable, the photovoltaic output, and the photovoltaic output point prediction value specifically includes: S51. Using a nonparametric kernel density estimation method, perform probability integral transformation on the target weather variable, the photovoltaic output, and the photovoltaic output point prediction value, respectively, to obtain the marginal probability distribution of the target weather variable, the marginal probability distribution of the photovoltaic output, and the marginal probability distribution of the photovoltaic output point prediction value; S52: using a reinforcement learning algorithm to determine an optimal connection relationship among the target weather variable, the photovoltaic output, and the photovoltaic output point prediction value; S53, using the marginal probability distribution of the target weather variable, the marginal probability distribution of the photovoltaic output, and the marginal probability distribution of the photovoltaic output point prediction value as vine tree nodes, and connecting the vine tree nodes using the optimal connection relationship to obtain a target vine tree; S54, using the maximum likelihood estimation method and the Akaike Information Criterion method to determine the optimal Copula function and the parameter value of the optimal Copula function for each edge in the target vine tree, and using the optimal Copula function, the parameter value of the optimal Copula function, and each vine tree node to respectively calculate the conditional probability of each edge in the target vine tree; S55. Taking the conditional probability of each edge as the vine tree node of the next vine tree; S56. Using a reinforcement learning algorithm, determine the optimal connection relationship between the vine tree nodes of the next vine tree; S57. Connecting the vine tree nodes of the next vine tree according to the optimal connection relationship between the vine tree nodes of the next vine tree to obtain the next vine tree; S58: Update the next vine tree to the target vine tree in step S54, and jump to S59; S59, looping through steps S54 to S58 until the updated target vine tree in step S58 has only one edge, stopping the loop, and outputting the target vine tree and each updated target vine tree as the vine copula model; The method of using the maximum likelihood estimation method and the Akaike information criterion method to determine the optimal Copula function of each edge in the target vine tree and the parameter value of the optimal Copula function includes: Obtaining multiple Copula functions for each edge in the target vine tree respectively; Calculating the parameter value of each copula function using maximum likelihood estimation method; Calculating the AIC value of each Copula function according to each Copula function and the parameter value of each Copula function; The Copula function with the smallest AIC value is taken as the optimal Copula function, and the parameter value of the Copula function with the smallest AIC value is taken as the function value of the optimal Copula function.

2. The method according to claim 1, characterized in that The determining of the photovoltaic output point prediction value according to the historical weather factor data, the predicted weather factor data, and the historical photovoltaic output data includes: Synchronously aligning the historical weather factor data, the predicted weather factor data, and the historical photovoltaic output data in a time dimension; According to a preset division ratio, the historical weather factor data, the predicted weather factor data, and the historical photovoltaic output data after synchronization and alignment are divided into a training set, a validation set, and a set to be predicted; The training set and the validation set are used to train a pre-built photovoltaic output point prediction model, and the set to be predicted is input into the trained photovoltaic output point prediction model to output the photovoltaic output point prediction value.

3. The method according to claim 2, characterized in that The calculating of the correlation coefficient and collinearity strength between photovoltaic output, weather factors and the photovoltaic output point prediction value by using the historical weather factor data, the historical photovoltaic output data and the photovoltaic output point prediction value comprises: Normalizing the historical weather factor data, the historical photovoltaic output data, and the photovoltaic output point prediction value respectively; Calculating the Kendall correlation coefficient and variance inflation factor between the photovoltaic output, the weather factors, and the photovoltaic output point prediction value based on the normalized historical weather factor data, the historical photovoltaic output data, and the photovoltaic output point prediction value; The Kendall correlation coefficient was used as the correlation coefficient, and the variance inflation factor was used as the collinearity strength.

4. The method according to claim 1, wherein Determining the photovoltaic output confidence interval based on the vine copula model, the predicted weather factor data, the photovoltaic output point prediction value, and the preset confidence level specifically includes: The vine copula model is reversed to obtain the quantile regression expression of the photovoltaic output; The predicted weather factor data, the predicted value of the photovoltaic output point, and a preset confidence level are input into the quantile regression expression, and a photovoltaic output confidence interval is output.

5. A photovoltaic output probability prediction device, characterized in that: include: An acquisition module is used to obtain historical weather factor data, predicted weather factor data, and historical photovoltaic output data of the area to be predicted; A first determination module is configured to determine a photovoltaic output point prediction value based on the historical weather factor data, the predicted weather factor data, and the historical photovoltaic output data; A calculation module, configured to calculate the correlation coefficient and collinearity strength between the historical weather factor data, the historical photovoltaic output data, and the photovoltaic output point prediction value; a second determining module, determining a target weather variable from the weather factors according to the correlation coefficient and the collinearity strength; A construction module, using the target weather variable, the photovoltaic output, and the photovoltaic output point prediction value to construct a vine copula model; A prediction module, configured to determine a photovoltaic output confidence interval based on the vine copula model, the predicted weather factor data, the photovoltaic output point prediction value, and a preset confidence level; The building block is specifically configured to perform the following steps: S51. Using a nonparametric kernel density estimation method, perform probability integral transformation on the target weather variable, the photovoltaic output, and the photovoltaic output point prediction value, respectively, to obtain the marginal probability distribution of the target weather variable, the marginal probability distribution of the photovoltaic output, and the marginal probability distribution of the photovoltaic output point prediction value; S52: using a reinforcement learning algorithm to determine an optimal connection relationship among the target weather variable, the photovoltaic output, and the photovoltaic output point prediction value; S53, using the marginal probability distribution of the target weather variable, the marginal probability distribution of the photovoltaic output, and the marginal probability distribution of the photovoltaic output point prediction value as vine tree nodes, and connecting the vine tree nodes using the optimal connection relationship to obtain a target vine tree; S54, using the maximum likelihood estimation method and the Akaike Information Criterion method to determine the optimal Copula function and the parameter value of the optimal Copula function for each edge in the target vine tree, and using the optimal Copula function, the parameter value of the optimal Copula function, and each vine tree node to respectively calculate the conditional probability of each edge in the target vine tree; S55. Taking the conditional probability of each edge as the vine tree node of the next vine tree; S56. Using a reinforcement learning algorithm, determine the optimal connection relationship between the vine tree nodes of the next vine tree; S57. Connecting the vine tree nodes of the next vine tree according to the optimal connection relationship between the vine tree nodes of the next vine tree to obtain the next vine tree; S58: Update the next vine tree to the target vine tree in step S54, and jump to S59; S59, looping through steps S54 to S58 until the updated target vine tree in step S58 has only one edge, stopping the loop, and outputting the target vine tree and each updated target vine tree as the vine copula model; The method of using the maximum likelihood estimation method and the Akaike information criterion method to determine the optimal Copula function of each edge in the target vine tree and the parameter value of the optimal Copula function includes: Obtaining multiple Copula functions for each edge in the target vine tree respectively; Calculating the parameter value of each copula function using maximum likelihood estimation method; Calculating the AIC value of each Copula function according to each Copula function and the parameter value of each Copula function; The Copula function with the smallest AIC value is taken as the optimal Copula function, and the parameter value of the Copula function with the smallest AIC value is taken as the function value of the optimal Copula function.

6. The device according to claim 5, characterized in that The first determining module includes: an alignment module, configured to synchronize the historical weather factor data, the predicted weather factor data, and the historical photovoltaic output data in a time dimension; a partitioning module, configured to partition the synchronized and aligned historical weather factor data, the predicted weather factor data, and the historical photovoltaic output data into a training set, a validation set, and a set to be predicted according to a preset partitioning ratio; The training module is used to use the training set and the verification set to train the pre-built photovoltaic output point prediction model, and input the set to be predicted into the trained photovoltaic output point prediction model to output the photovoltaic output point prediction value.

7. An electronic device, characterized in that: The device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the method according to any one of claims 1 to 4 according to instructions in the program code.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store program code, and the program code is used to execute the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Photovoltaic power probability prediction method and system based on copula function

    CN116470491A

  • Probability distribution prediction method and device for photovoltaic output

    CN117932208A