Method for generating typical scene of watershed water, wind and light based on improved C-vine Copula
By improving the C-vine Copula model and Latin hypercube sampling combined with K-means clustering, the problem of multi-dimensional random variable dependence and low sampling efficiency in the watershed and light system in the basin is solved, and a more accurate typical scenario set is generated to support system planning and scheduling.
Patent Information
- Application Number
- CN202510501711.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The prior art is difficult to effectively characterize the complex dependence relationship of multi-dimensional random variables in water and the low sampling efficiency of high-dimensional random variables in watershed and light systems, which leads to the inability to accurately generate typical scenario sets of watershed and light integrated systems.
The improved C-vine Copula model is used to layer-based modeling of the spatiotemporal correlation of water and scenery systems, and combined with Latin hypercube sampling and K-means clustering methods, representative typical scenarios are generated.
The complexity of the model is simplified, the sampling efficiency and coverage are improved, and the generated typical scenario set can better reflect the actual distribution characteristics, supporting the planning and scheduling of the integrated water and wind system in the basin.
Smart Images

Figure CN120493692A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of generating typical water and wind scenes, and in particular relates to a method for generating typical water and wind scenes in a watershed based on an improved C-vineCopula. Background Art
[0002] Hydropower, with its large capacity and strong regulation capabilities, can complement wind and solar power well. However, over the long term, runoff and wind and solar power generation capacity exhibit significant randomness and seasonality, and are interdependent, making accurate characterization difficult. This poses significant challenges to the planning and operation of integrated, multi-energy systems combining hydropower, wind power, and solar power within a river basin. By considering the spatiotemporal correlations between hydropower, wind power, and solar power, a representative set of high-dimensional coupled scenarios is generated, allowing the uncertainty of the hydropower, wind power, and solar power system to be quantified and characterized in a finite discrete manner, thus providing effective support for planning and the construction of stochastic scheduling models.
[0003] A method for generating typical scenarios for water, wind, and solar power in a watershed that considers multiple uncertainties is based on long-term, long-series historical runoff data from a large hydropower station and fitted output data from wind and solar power stations within its basin. This historical data is used to simulate and quantify the multiple uncertainties of water, wind, and solar power. Numerous runoff and wind and solar power output sequences are generated to maximize the statistical characteristics of the historical data, providing effective input data for power system planning, scheduling, and risk analysis. Currently, existing technologies have certain shortcomings in addressing the problem of quantifying the multiple uncertainties of water, wind, and solar power resources. When modeling multidimensional variables using ordinary Copula functions, it is often necessary to model the dependencies between all variables at once. As the number of dimensions increases, the model structure becomes more complex. Furthermore, in water, wind, and solar power systems, the complex dependencies between inflow and wind and solar power output are not only direct but also conditional. A one-time modeling approach may not be able to capture the complex dependencies between these random variables. In addition, when sampling the joint distribution of high-dimensional random variables, the Monte Carlo sampling method is used to deal with high-dimensional problems. The number of samples will increase exponentially with the increase of dimensions, and the convergence speed is slow. At the same time, the sampled samples may be unevenly distributed, resulting in poor coverage and failure to reflect the characteristics of the distribution. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method for generating typical water, wind and solar scenes in a watershed based on an improved C-vine Copula, in order to solve the problem that existing methods are difficult to apply to the generation of long-term water, wind and solar coupling scene sets for watershed water, wind and solar integration. Specifically, in the process of generating water and wind scenes, the existing Copula theory is used to fit the joint distribution of multidimensional variables through one-time modeling, which results in complex model structure and large computational complexity and low coverage when sampling the joint distribution of multidimensional random variables.
[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0006] A method for generating typical waterscape scenes in a watershed based on an improved C-vine Copula is proposed. The steps are as follows:
[0007] Step 1: Data selection: Select the required historical data based on the long-term demand for water, wind, and solar energy multi-energy complementarity. The required historical data includes long-term historical runoff data of a power station and output data of wind and solar power stations in the power station's basin.
[0008] Step 2: Data preprocessing: Preprocess the required historical data, including filling missing values and removing outliers;
[0009] Step 3: Let historical runoff and wind and solar power output be random variables X h ,X w ,X s , the data distribution is smoothed by non-parametric kernel density estimation to reflect the true distribution of the data, and X is estimated respectively. h ,X w ,X s The probability density function of , and then the marginal distribution function of each random variable is obtained;
[0010] Step 4: Based on the marginal distribution functions of water, wind, and solar power obtained in step 3, the improved C-vine copula model is used to characterize the spatiotemporal correlation between the three factors, and the joint probability distribution of water, wind, and solar power is obtained.
[0011] Step 5: Based on the joint probability distribution of water, wind and solar power in step 4, stratified sampling is performed from the multidimensional distribution using the Latin hypercube sampling method. This method divides the sample space of each dimension into uniform subintervals to obtain a more uniform random sample set.
[0012] Step 6: Perform cluster analysis on the random sample set obtained in step 5, and use the K-means clustering method to cluster all samples into several representative typical scenarios;
[0013] Step 7: Based on the typical scenarios obtained in step 6, conduct a validity evaluation on the generated typical scenarios to ensure their representativeness and practicality.
[0014] Preferably, in step 1, considering the runoff into the hydropower station and assuming that new energy sources such as wind and solar power in the reservoir basin are clustered and integrated into the power grid according to energy type, no less than thirty years of monthly average wind power and photovoltaic output series data are selected.
[0015] Preferably, in step 2, preprocessing includes removing outliers in the required historical data and filling missing values using interpolation, and using mathematical calculations to obtain the monthly average values of the inflow runoff and wind and solar power output data for the existing data sets.
[0016] Preferably, in step 3, let runoff and wind power be variables X h ,X w ,X s ;
[0017] (1)X w The process of solving the distribution function is as follows:
[0018] If x 1w ,x 2w ...x nw Let n wind power output sample points be independent and identically distributed F, and let their probability density function be f. The kernel density estimate is as follows:
[0019]
[0020] Where h>0 is a smoothing parameter called bandwidth; K(·) is the kernel function; the Gaussian kernel function is selected, and the expression is as follows:
[0021]
[0022] The data and bandwidth of each data point are used as the parameters of the kernel function through the kernel function to obtain N kernel functions, which are then linearly superimposed to form the estimation function of the kernel density. After normalization, it is the kernel density probability density function; and then the distribution function of each random variable is obtained;
[0023] (2)X h ,X s The solution process of X w same.
[0024] Preferably, in step 4, the maximum column sum method is used to hierarchically construct the dependency relationship between each random variable; the information criterion method is used to select the optimal truncation level to determine the final improved C-vine Copula model to characterize the dependency relationship between water, wind and light.
[0025] Preferably, the specific method of step 4 is:
[0026] (1) Construct a C-vine Copula model for inflow runoff and wind and solar power output, and select the Kendall coefficient to characterize the correlation between variables;
[0027] Now assume that for any pair of n-dimensional random variables (X, Y), the Kendall correlation coefficient is calculated as follows:
[0028] τ=P[(X-X')(Y-Y')>0]-P[(X-X')(Y-Y')<0] (18);
[0029] Where P(·) is the probability density function, and the random vectors (X, Y) and (X', Y') follow the same distribution;
[0030] The three random variables X of water, wind and light are calculated by formula (3) h 、X w 、X s The Kendall coefficient between each pair; then the maximum column sum method is used to select the best root node, and the weight matrix is constructed as follows:
[0031]
[0032] Among them, τ xy is the Kendall coefficient;
[0033] (2) Calculate the column sum to select the root node: The dependency relationship between the root node variable and other node variables constitutes the first layer structure of the C-vine Copula. The second layer structure is to construct the dependency relationship between the remaining variables based on the first layer root node variable;
[0034] The optimal pair Copula function for each edge of the vine structure is selected, and the mixed Copula function is estimated using the EM algorithm to estimate the weight parameters and dependence coefficients of the mixed Copula. The constructed mixed Copula form is:
[0035]
[0036] Where: u, v are the marginal distributions of random variables; λ n ∈[0,1] is the weight parameter of the model, and is the dependent parameter of the model;
[0037] (3) The information criterion method is used to select the optimal truncation level. First, different truncation levels are assumed and the overall likelihood estimate of the C-vine Copula model under the corresponding truncation level is calculated. The formula is as follows:
[0038]
[0039] Where: L i is the likelihood estimate of the i-th layer of the C-vine Copula model, m is the total number of layers of the C-vine Copula model, and L is the overall likelihood estimate;
[0040] Based on the overall likelihood estimate obtained above, AIC and BIC are calculated to determine the optimal truncation level. The formula is as follows:
[0041] AIC = 2k-xln(L) (22);
[0042] BIC = kln(n)-2ln(L) (23);
[0043] Among them, k is the number of parameters, L is the likelihood function, n is the sample size; AIC is the information criterion, BIC is the Bayesian information criterion;
[0044] According to the principle of smaller AIC and BIC, the best truncation level is selected at the second level of the C-vine Copula model; the final C-vine Copula structure is as follows: Assume that there are n-dimensional variables X = (x1, x2, ..., x n ), F(x1,x2,...,x n ) and f(x1,x2,...,x n ) are their joint distribution function and joint probability density function, F i (x i ) and f i (x i ) represent the cumulative distribution function and probability density function respectively; then the joint probability distribution of the C-vine structure of X is expressed as follows:
[0045]
[0046] Where c j,j+i|11...j-1 (·) is the given x1, x2,…, x j-1 Under the condition, variable x j and x j+i The Copula probability density function composed of the two; F(x j |x1,…,x j-1 ) are known x1, x2,…, x j-1 Under the condition that the variable x j The distribution function of .
[0047] Preferably, the specific method of step 5 is:
[0048] Latin hypercube sampling is performed on the improved C-vine Copula model. Assume that the sample size is N and the Nth sample is represented by s n =[s nh ,s nw ,s ns ], the specific steps are as follows:
[0049] (1) Assume x wis the root node random variable and x w ∈[x wd ,x wu ], x w The marginal distribution function is f w (x w ); Set the range [f w (x wd ),f w (x wu )] is divided into N equal probability intervals, and q is randomly selected from all probability subintervals. i , get the uniform variable Z1;
[0050] (2) Let U1, U2, and U3 be three sets of variables to be determined. The uniform variable obtained by the above steps is equal to U1, that is, U1 = Z1, and Z1 is regarded as the sampling point of the variables to be determined;
[0051] (3) Similarly, according to the above step (1), random numbers that obey uniform distribution are generated, which are defined as uniform variables Z2 and Z3. From the obtained improved C-vine model and the conditional distribution function of formula (11), it can be seen that the second group of variables to be determined U2 can be used Calculate, where Z2 and U1 are both known quantities, thereby converting the above nonlinear equation into a linear equation and solving it using the bisection method. The resulting solution to the equation is the sample point of the variable U2 to be determined;
[0052] (4) Similarly, since Z2 and Z3 have been defined, we can get:
[0053]
[0054] Find F(x3|x1), and Solve the equation; the result is the sample of the variable U3 to be solved;
[0055] (5) Finally, perform inverse transformation sampling on U1, U2, and U3 to obtain the water-landscape scene set S = [S h ,S w ,S s ].
[0056] Preferably, in step 6, the data set is divided into K clusters by using the K-means clustering sampling distance as the similarity evaluation index; the operation steps of performing K-means clustering based on the scene set obtained in step 5 are as follows:
[0057] (1) Initialization: Randomly select K sample points from the data set as the initial centroid;
[0058] (2) Assign sample points to the nearest centroid: For each sample point, calculate its Euclidean distance to all centroids and assign it to the cluster where the nearest centroid is located. The Euclidean distance formula is as follows:
[0059]
[0060] Among them, C m is the center of mass, X i is the sample point, d k (X i ,C m ) is the Euclidean distance from the sample point to the centroid;
[0061] (3) Update the centroid: Calculate the average value of all sample points in the cluster and use it as the new centroid;
[0062] (4) Repeat steps (2) and (3): keep assigning sample points and updating the centroid until the centroid no longer changes or the preset number of iterations is reached;
[0063] (5) Assign weights and count the number of scenes in each cluster. If a cluster contains n i scenes, then the weight p of the representative scene in the cluster i for:
[0064]
[0065] Among them, n i is the number of sample points contained in the i-th cluster, p i is the weight parameter for ;
[0066] (6) Output the centroid of each cluster and the corresponding weight parameter.
[0067] Preferably, in step 7, the generated typical scene is evaluated by considering temporal correlation, spatial correlation and randomness;
[0068] The autocorrelation coefficient is used as the time correlation indicator for evaluation; its formula is as follows:
[0069]
[0070] Where r t represents the value of month t; k represents the time interval; represents the mean; ρ k represents the autocorrelation coefficient with time delay k;
[0071] The average absolute error of the Kendall correlation coefficient is selected to evaluate the representation quality of the generated typical scene spatial correlation; the formula is as follows:
[0072]
[0073] Where: τ i is the Kendall correlation coefficient between the typical scenarios in group i; is the Kendall correlation coefficient between historical time series data; p i is the weight parameter of the i-th group of typical scenarios;
[0074] Randomness refers to the deviation between the actual scenario of runoff and wind power and photovoltaic output in the time section and the generated scenario set. The scenario is evaluated based on the distance, and the formula is as follows:
[0075]
[0076] Where: R i is the i-th typical scene sequence; R is the actual historical sequence; ||·||2 is the L2 norm of the difference between the two sequences; E represents the average Euclidean distance between the typical scene set and the actual historical sequence.
[0077] A watershed waterscape typical scene generation system based on improved C-vine Copula adopts the watershed waterscape typical scene generation method based on improved C-vine Copula.
[0078] The present invention can achieve the following beneficial effects:
[0079] 1) This paper innovatively utilizes an improved C-vine Copula model to characterize the spatiotemporal correlations between random variables. This model hierarchically models the dependency structure of all random variables, decomposing the high-dimensional joint distribution into a series of low-dimensional pair-Copula functions while truncating them at appropriate levels. This simplifies the computational and modeling process, effectively balancing model complexity and fitting accuracy.
[0080] 2) This paper uses Latin hypercube sampling to perform stratified sampling on the joint distribution of multidimensional random variables. This evenly divides the sample space of each dimension into several subintervals and randomly selects sample points within each subinterval. This ensures sampling uniformity while also improving sampling efficiency and coverage. This approach is combined with the K-means clustering algorithm, using Euclidean distance as the evaluation metric for scene similarity, to further analyze and optimize the water-wind-solar system. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] The present invention will be further described below with reference to the accompanying drawings and examples:
[0082] Figure 1 It is the overall flow chart of the present invention;
[0083] Figure 2It is a flow chart of the improved C-vine Copula model of the present invention;
[0084] Figure 3 It is a structural diagram of the improved C-vine Copula model of the present invention;
[0085] Figure 4 It is a flow chart of the Latin hypercube algorithm of the present invention;
[0086] Figure 5 It is a flow chart of the K-means clustering algorithm of the present invention. DETAILED DESCRIPTION
[0087] The preferred solution is Figures 1 to 5 As shown in the figure, a method for generating typical watershed, wind, and solar-power scenarios based on an improved C-vine copula is proposed. First, a multi-year series of historical runoff and wind and solar power output data is used as model input. Nonparametric kernel density estimation is used to fit the marginal distribution of each random variable. Second, based on the improved C-vine copula function, a sampling stratified modeling approach is used to construct the spatiotemporal correlations between the random variables, and truncation is performed at certain levels. Next, Latin hypercube sampling is combined to perform stratified sampling on the joint distribution of the multivariate coupled system to generate a set of scenarios for the water, wind, and solar multi-energy complementary system. K-means clustering is then used to cluster the large number of samples into a few representative scenarios. Finally, the resulting representative scenarios are effectively evaluated.
[0088] The following is the overall process Figure 1 To explain in detail, it mainly includes the following steps:
[0089] Step 1: Data selection. Select the required historical data based on the long-term water, wind and solar energy complementary needs of this study. The required historical data include long-term historical runoff data of a large hydropower station and output data of wind and solar power stations in the basin of the station.
[0090] Step 2: Data preprocessing: select the required historical data according to the method in step 1 and preprocess the data, including filling missing values and removing outliers.
[0091] Step 3: Let runoff and wind / solar output be random variables X h ,X w ,X s , the data distribution is smoothed by non-parametric kernel density estimation to reflect the true distribution of the data, and X is estimated respectively. h ,X w ,X s The probability density function of , and then the marginal distribution function of each random variable is obtained.
[0092] Step 4: Based on the marginal distribution functions of water, wind, and light obtained in Step 3, the C-vine Copula model is used to characterize the spatiotemporal correlations among the three factors. First, the maximum column sum method is used to hierarchically construct the dependencies between the random variables. Next, the information criterion is used to select the optimal cutoff level to determine the final improved C-vine Copula model to characterize the dependencies among the three factors.
[0093] Step 5: Based on the joint distribution of water, wind, and light in Step 4, stratified sampling is performed from the multidimensional distribution using the Latin hypercube sampling method. This method achieves more uniform sampling by dividing the sample space of each dimension into uniform subintervals. A more uniform random sample set is obtained by dividing the sample space of each dimension into uniform subintervals.
[0094] Step 6: Perform cluster analysis on the random sample set obtained in step 5, and use the K-means clustering method to cluster all samples into several representative typical scenarios.
[0095] Step 7: Based on the typical scenarios obtained in step 6, conduct a validity evaluation on the generated typical scenarios to ensure their representativeness and practicality.
[0096] Specifically, step 1 considers the runoff inflow from a large hydropower station and assumes that renewable energy sources such as wind and solar power within the reservoir basin are integrated into the grid in their own clusters. The runoff data is provided by the corresponding entity, and the wind and solar power output data is estimated based on historical meteorological data using a model developed by the National Renewable Energy Laboratory. This yields a long-term series of monthly average wind and solar power output over multiple years.
[0097] Taking into account the runoff from the hydropower station and assuming that new energy sources such as wind and solar power in the reservoir basin are clustered and integrated into the power grid according to energy type, no less than thirty years of monthly average wind power and photovoltaic output series data are selected.
[0098] Specifically, step 2 shown is as follows: data preprocessing is performed on the required historical data selected according to step 1, including removing outliers and filling missing values using interpolation. For the existing data set, mathematical calculations are used to obtain the monthly average values of the inflow runoff and wind and solar power output data.
[0099] Specifically, let runoff and wind power output be random variables X h ,X w ,X s , with wind power output X w For example, if x 1w ,x 2w x nw Let n wind power output sample points be independent and identically distributed F, and let their probability density function be f. The kernel density estimate is as follows:
[0100]
[0101] Where h>0 is a smoothing parameter called bandwidth; K(·) is the kernel function. Common kernel functions include Epanechikov function and Gaussian function. The present invention selects Gaussian kernel function, and the expression is as follows:
[0102]
[0103] The kernel function uses the data and bandwidth of each data point as the parameters of the kernel function to obtain N kernel functions, which are then linearly superimposed to form the kernel density estimation function. After normalization, it becomes the kernel density probability density function. In turn, the distribution function of each random variable is obtained.
[0104] Specifically, in step 4, we first construct a C-vine copula model for runoff and wind and solar power output, and select an appropriate correlation coefficient to characterize the correlation between the variables. Here, we use the Kendall coefficient as the measure. Assuming any pair of n-dimensional random variables (X, Y), the Kendall correlation coefficient is calculated as follows:
[0105] τ=P[(X-X')(Y-Y')>0]-P[(X-X')(Y-Y')<0] (33)
[0106] Where P(·) is the probability density function, and the random vectors (X, Y) and (X', Y') obey the same distribution.
[0107] The above formulas are used to calculate the three random variables X of water, wind and light. h 、X w 、X s The Kendall coefficient between each pair. Then the maximum column sum method is used to select the best root node and construct the weight matrix as follows:
[0108]
[0109] Among them, τ xy is the Kendall coefficient.
[0110] The root node is selected by calculating the column sum. The dependencies between the root node variables and other node variables constitute the first-level structure of the C-vineCopula. The second-level structure constructs the dependencies between the remaining variables, conditional on the first-level root node variables. To select the optimal pair copula function for each edge of the vine structure, this paper introduces a hybrid copula function and uses the EM algorithm to estimate the weight parameters and dependency coefficients of the hybrid copula. The constructed hybrid copula is in the form of:
[0111]
[0112] Where: u, v are the marginal distributions of random variables; λ n ∈[0,1] is the weight parameter of the model, and are the dependent parameters of the model.
[0113] Next, we use the information criterion method to select the optimal truncation level. First, we assume different truncation levels and calculate the overall likelihood estimate of the C-vine Copula model under the corresponding truncation level. The formula is as follows:
[0114]
[0115] Where: L i is the likelihood estimate of the i-th layer of the C-vine Copula model, m is the total number of layers of the C-vine Copula model, and L is the overall likelihood estimate.
[0116] Based on the overall likelihood estimate obtained above, AIC (Akaike Information Criterion) and BIC (Bayesian Information Criterion) are calculated to determine the optimal truncation level. The formula is as follows:
[0117] AIC=2k-xln(L) (37)
[0118] BIC=kln(n)-2ln(L) (38)
[0119] Among them, k is the number of parameters, L is the likelihood function, n is the sample size.
[0120] According to the principle of smaller AIC and BIC, the best truncation level is selected at the second level of the C-vine Copula model. The final C-vine Copula structure is as follows: Assume that there are n-dimensional variables X = (x1, x2, ..., x n ), F(x1,x2,...,x n ) and f(x1,x2,...,x n ) are their joint distribution function and joint probability density function, F i (x i ) and f i (x i ) represent the cumulative distribution function and probability density function respectively. Then the joint probability distribution of the C-vine structure of X is expressed as follows:
[0121]
[0122] Where c j,j+i|11...j-1(·) is the given x1, x2,…, x j-1 Under the condition, variable x j and x j+i The Copula probability density function composed of the two; F(x j |x1,…,x j-1 ) are given x1, x2,…, x j-1 Under the condition that the variable x j The distribution function of .
[0123] Specifically, as shown in step 5: Latin hypercube sampling is performed on the improved C-vine Copula model. Now, assuming that the sample size is N, the Nth sample can be expressed as s n =[s nh ,s nw ,s ns ], the specific steps are as follows:
[0124] (1) Assume x w is the root node random variable and x w ∈[x wd ,x wu ], x w The marginal distribution function is f w (x w ). Set the range [f w (x wd ),f w (x wu )] is divided into N equal probability intervals, and q is randomly selected from all probability subintervals. i , and obtain the uniform variable Z1.
[0125] (2) Let U1, U2, and U3 be three sets of variables to be determined. The uniform variable obtained by the above steps is equal to U1, that is, U1 = Z1. Z1 can be regarded as the sampling point of the variables to be determined.
[0126] (3) Similarly, generate random numbers that obey uniform distribution according to the above step 1, which are defined as uniform variables Z2 and Z3. According to the improved C-vine model and the conditional distribution function of formula (11), the second group of variables to be determined U2 can be obtained using Calculate, where Z2 and U1 are both known quantities, thereby converting the above one-variable nonlinear equation into a linear equation and solving it using the bisection method. The resulting solution to the equation is the sample point of the variable U2 to be determined.
[0127] (4) Similarly, since Z2 and Z3 have been defined, we can
[0128]
[0129] Find F(x3|x1), and Solve the equation. The result is the sample of the variable U3 to be solved.
[0130] (5) Finally, perform inverse transformation sampling on U1, U2, and U3 to obtain the water-landscape scene set S = [S h ,S w ,S s ].
[0131] Specifically, in step 6, the dataset is divided into K clusters using the K-means clustering sampling distance as the similarity evaluation index. The following are the steps for performing K-means clustering based on the scene set obtained in step (5).
[0132] (1) Initialization: Randomly select K sample points from the data set as the initial centroid.
[0133] (2) Assign sample points to the nearest centroid: For each sample point, calculate its Euclidean distance to all centroids and assign it to the cluster where the nearest centroid is located. The Euclidean distance formula is as follows:
[0134]
[0135] Among them, C m is the center of mass, X i is the sample point, d k (X i ,C m ) is the Euclidean distance from the sample point to the centroid.
[0136] (3) Update the centroid: Calculate the average value of all sample points in the cluster and use it as the new centroid.
[0137] (4) Repeat steps (2) and (3): Repeat the allocation of sample points and the updating of the centroid until the centroid no longer changes or the preset number of iterations is reached.
[0138] (5) Assign weights and count the number of scenes in each cluster. If a cluster contains n i scenes, then the weight p of the representative scene in the cluster i for:
[0139]
[0140] Among them, n i is the number of sample points contained in the i-th cluster, p i is the weight parameter for .
[0141] (6) Output the centroid of each cluster and the corresponding weight parameter.
[0142] Specifically, in step (7), the present invention evaluates the generated typical scenes by taking into account temporal correlation, spatial correlation and randomness.
[0143] The autocorrelation coefficient is used as the time correlation indicator. Its formula is as follows:
[0144]
[0145] Where r t represents the value of month t; k represents the time interval; r represents the mean; ρ k represents the autocorrelation coefficient with time delay k.
[0146] The average absolute error of the Kendall correlation coefficient is selected to evaluate the representation quality of the generated typical scene spatial correlation. The formula is as follows:
[0147]
[0148] Where: τ i is the Kendall correlation coefficient between the typical scenarios in group i; is the Kendall correlation coefficient between historical time series data; p i is the weight parameter of the i-th group of typical scenarios;
[0149] Randomness refers to the deviation between the actual scenario of runoff and wind power and photovoltaic output in the time section and the generated scenario set. The scenario is evaluated based on the distance, and the formula is as follows:
[0150]
[0151] Where: R i is the i-th typical scene sequence; R is the actual historical sequence; ||·||2 is the L2 norm of the difference between the two sequences; E represents the average Euclidean distance between the typical scene set and the actual historical sequence.
[0152] The above embodiments are merely preferred technical solutions of the present invention and should not be construed as limiting the present invention. The scope of protection of the present invention shall be the technical solutions set forth in the claims, including equivalent alternatives to the technical features of the technical solutions set forth in the claims. In other words, equivalent alternatives and improvements within this scope are also within the scope of protection of the present invention.
Claims
1. A method for generating typical waterscape scenes in a watershed based on an improved C-vine Copula, characterized by The following steps are involved: Step 1: Data selection: Select the required historical data based on the long-term demand for water, wind, and solar energy multi-energy complementarity. The required historical data includes long-term historical runoff data of a power station and output data of wind and solar power stations in the power station's basin. Step 2: Data preprocessing: Preprocess the required historical data, including filling missing values and removing outliers; Step 3: Let historical runoff and wind and solar power output be random variables X h ,X w ,X s , the data distribution is smoothed by non-parametric kernel density estimation to reflect the true distribution of the data, and X is estimated respectively. h ,X w ,X s The probability density function of , and then the marginal distribution function of each random variable is obtained; Step 4: Based on the marginal distribution functions of water, wind, and solar power obtained in step 3, the improved C-vine copula model is used to characterize the spatiotemporal correlation between the three factors, and the joint probability distribution of water, wind, and solar power is obtained. Step 5: Based on the joint probability distribution of water, wind and solar power in step 4, stratified sampling is performed from the multidimensional distribution using the Latin hypercube sampling method. This method divides the sample space of each dimension into uniform subintervals to obtain a more uniform random sample set. Step 6: Perform cluster analysis on the random sample set obtained in step 5, and use the K-means clustering method to cluster all samples into several representative typical scenarios; Step 7: Based on the typical scenarios obtained in step 6, conduct a validity evaluation on the generated typical scenarios to ensure their representativeness and practicality.
2. The method for generating typical waterscape scenes based on the improved C-vine copula according to claim 1, characterized in that: In step 1, considering the runoff into the hydropower station and assuming that new energy sources such as wind and solar power in the reservoir basin are clustered and integrated into the power grid according to energy type, the monthly average wind power and photovoltaic output series data of no less than 30 years are selected.
3. The method for generating typical waterscape scenes based on the improved C-vine copula according to claim 1, characterized in that: In step 2, preprocessing includes removing outliers from the required historical data and filling missing values using interpolation. For the existing data set, mathematical calculations are used to obtain the monthly average values of inflow runoff and wind and solar power output data.
4. The method for generating typical waterscape scenes based on the improved C-vine copula according to claim 1, characterized in that: In step 3, let runoff and wind power be variables X h ,X w ,X s , (1)X w The process of solving the distribution function is as follows: If x 1w ,x 2w x nw Let n wind power output sample points be independent and identically distributed F, and let their probability density function be f. The kernel density estimate is as follows: Where h>0 is a smoothing parameter called bandwidth; K(·) is the kernel function; the Gaussian kernel function is selected, and the expression is as follows: The data and bandwidth of each data point are used as the parameters of the kernel function through the kernel function to obtain N kernel functions, which are then linearly superimposed to form the estimation function of the kernel density. After normalization, it is the kernel density probability density function; and then the distribution function of each random variable is obtained; (2)X h ,X s The solution process of X w same.
5. The method for generating typical waterscape scenes based on the improved C-vine copula according to claim 1, characterized in that: In step 4, the maximum column sum method is used to hierarchically construct the dependency relationship between each random variable; the information criterion method is used to select the optimal truncation level to determine the final improved C-vine Copula model to characterize the dependency relationship between water, wind and light.
6. The method for generating typical waterscape scenes based on the improved C-vine copula according to claim 5, characterized in that: The specific method of step 4 is: (1) Construct a C-vine Copula model for inflow runoff and wind and solar power output, and select the Kendall coefficient to characterize the correlation between variables; Now assume that for any pair of n-dimensional random variables (X, Y), the Kendall correlation coefficient is calculated as follows: τ=P[(X-X')(Y-Y')>0]-P[(X-X')(Y-Y')<0] (3); Where P(·) is the probability density function, and the random vectors (X, Y) and (X', Y') follow the same distribution; The three random variables X of water, wind and light are calculated by formula (3) h 、X w 、X s The Kendall coefficient between each pair; then the maximum column sum method is used to select the best root node, and the weight matrix is constructed as follows: Among them, τ xy is the Kendall coefficient; (2) Calculate the column sum to select the root node: The dependency relationship between the root node variable and other node variables constitutes the first layer structure of C-vineCopula. The second layer structure is to construct the dependency relationship between the remaining variables based on the first layer root node variable; The optimal pair Copula function for each edge of the vine structure is selected, and the mixed Copula function is estimated using the EM algorithm to estimate the weight parameters and dependence coefficients of the mixed Copula. The constructed mixed Copula form is: Where: u, v are the marginal distributions of random variables; λ n ∈[0,1] is the weight parameter of the model, and θ n is the dependent parameter of the model; (3) The information criterion method is used to select the optimal truncation level. First, different truncation levels are assumed and the overall likelihood estimate of the C-vine Copula model under the corresponding truncation level is calculated. The formula is as follows: Where: L i is the likelihood estimate of the i-th layer of the C-vine Copula model, m is the total number of layers of the C-vine Copula model, and L is the overall likelihood estimate; Based on the overall likelihood estimate obtained above, AIC and BIC are calculated to determine the optimal truncation level. The formula is as follows: AIC = 2k-xln(L) (7); BIC = kln(n)-2ln(L) (8); Among them, k is the number of parameters, L is the likelihood function, n is the sample size; AIC is the information criterion, BIC is the Bayesian information criterion; According to the principle of smaller AIC and BIC, the best truncation level is selected at the second level of the C-vine Copula model; the final C-vine Copula structure is as follows: Assume that there are n-dimensional variables X = (x1, x2, ..., x n ), F(x1,x2,…,x n ) and f(x1,x2,...,x n ) are their joint distribution function and joint probability density function, F i (x i ) and f i (x i ) represent the cumulative distribution function and probability density function respectively; then the joint probability distribution of the C-vine structure of X is expressed as follows: Where c j,j+i|11...j-1 (·) is the given x1, x2,…, x j-1 Under the condition, variable x j and x j+i The Copula probability density function composed of the two; F(x j |x1,…,x j-1 ) are given x1, x2,…, x j-1 Under the condition that the variable x j The distribution function of .
7. The method for generating typical waterscape scenes based on the improved C-vine copula according to claim 1, characterized in that: The specific method of step 5 is: Latin hypercube sampling is performed on the improved C-vine Copula model. Assume that the sample size is N and the Nth sample is represented by s n =[s nh ,s nw ,s ns ], the specific steps are as follows: (1) Assume x w is the root node random variable and x w ∈[x wd ,x wu ], x w The marginal distribution function is f w (x w ); Set the range [f w (x wd ),f w (x wu )] is divided into N equal probability intervals, and q is randomly selected from all probability subintervals. i , get the uniform variable Z1; (2) Let U1, U2, and U3 be three sets of variables to be determined. The uniform variable obtained by the above steps is equal to U1, that is, U1 = Z1, and Z1 is regarded as the sampling point of the variables to be determined; (3) Similarly, according to the above step (1), random numbers that obey uniform distribution are generated, which are defined as uniform variables Z2 and Z3. From the obtained improved C-vine model and the conditional distribution function of formula (11), it can be seen that the second group of variables to be determined U2 can be used Calculate, where Z2 and U1 are both known quantities, thereby converting the above nonlinear equation into a linear equation and solving it using the bisection method. The resulting solution to the equation is the sample point of the variable U2 to be determined; (4) Similarly, since Z2 and Z3 have been defined, we can get: Find F(x3|x1), and Solve the equation; the result is the sample of the variable U3 to be solved; (5) Finally, perform inverse transformation sampling on U1, U2, and U3 to obtain the water-landscape scene set S = [S h ,S w ,S s ].
8. The method for generating typical waterscape scenes based on the improved C-vine copula according to claim 7, characterized in that: In step 6, the K-means clustering sampling distance is used as the similarity evaluation index to divide the data set into K clusters. The steps for performing K-means clustering based on the scene set obtained in step 5 are as follows: (1) Initialization: Randomly select K sample points from the data set as the initial centroid; (2) Assign sample points to the nearest centroid: For each sample point, calculate its Euclidean distance to all centroids and assign it to the cluster where the nearest centroid is located. The Euclidean distance formula is as follows: Among them, C m is the center of mass, X i is the sample point, d k (X i ,C m ) is the Euclidean distance from the sample point to the centroid; (3) Update the centroid: Calculate the average value of all sample points in the cluster and use it as the new centroid; (4) Repeat steps (2) and (3): keep assigning sample points and updating the centroid until the centroid no longer changes or the preset number of iterations is reached; (5) Assign weights and count the number of scenes in each cluster. If a cluster contains n i scenes, then the weight p of the representative scene in the cluster i for: Among them, n i is the number of sample points contained in the i-th cluster, p i is the weight parameter for ; (6) Output the centroid of each cluster and the corresponding weight parameter.
9. The method for generating typical waterscape scenes based on the improved C-vine copula according to claim 1, characterized in that: In step 7, the generated typical scenarios are evaluated by considering temporal correlation, spatial correlation and randomness; The autocorrelation coefficient is used as the time correlation indicator for evaluation; its formula is as follows: Where r t represents the value of month t; k represents the time interval; represents the mean; ρ k represents the autocorrelation coefficient with time delay k; The average absolute error of the Kendall correlation coefficient is selected to evaluate the representation quality of the generated typical scene spatial correlation; the formula is as follows: Where: τ i is the Kendall correlation coefficient between the typical scenarios in group i; is the Kendall correlation coefficient between historical time series data; p i is the weight parameter of the i-th group of typical scenarios; Randomness refers to the deviation between the actual scenario of runoff and wind power and photovoltaic output in the time section and the generated scenario set. The scenario is evaluated based on the distance, and the formula is as follows: Where: R i is the i-th typical scene sequence; R is the actual historical sequence; ||·||2 is the L2 norm of the difference between the two sequences; E represents the average Euclidean distance between the typical scene set and the actual historical sequence.
10. A watershed waterscape typical scene generation system based on improved C-vine Copula, characterized by: A method for generating typical waterscape scenes in a watershed based on an improved C-vine Copula according to claim 1 is adopted.
Citation Information
Patent Citations
Water-wind-light long-term scene generation method considering space-time correlation
CN117609586A
Wind-solar combined output uncertainty modeling method and system
CN118520686A
Cited By
A method and system for generating a wind-solar load space-time correlation scenario
CN122509021A
A method and system for generating a wind-solar load space-time correlation scenario
CN122509021B