A method for generating typical scenes of water and scenery in a basin based on improved C-vine Copula

By combining the improved C-vine Copula model and Latin hypercube sampling with K-means clustering, the problem of the difficulty in representing the multidimensional random variable dependencies in the watershed water-wind-solar system is solved, generating an efficient and representative set of water-wind-solar coupled scenarios to support the planning and scheduling of the power system.

CN120493692BActive Publication Date: 2026-04-17CHINA YANGTZE POWER
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA YANGTZE POWER
Filing Date
2025-04-21
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively characterize the complex dependencies of multidimensional random variables and the low sampling efficiency of the joint distribution of high-dimensional random variables in watershed water-wind-solar systems, resulting in the inability to accurately generate coupled scene sets for integrated water-wind-solar systems.

Method used

An improved C-vine Copula model is used to hierarchically model the spatiotemporal correlation of each random variable, and combined with Latin hypercube sampling and K-means clustering algorithm, typical water landscape scenes are generated.

Benefits of technology

It simplifies model complexity, improves sampling efficiency and coverage, and generates a more representative set of water-wind-solar coupled scenarios, providing effective input data for power system planning and scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493692B_ABST
    Figure CN120493692B_ABST
Patent Text Reader

Abstract

A method for generating typical watershed hydro-wind-solar hybrid scenarios based on an improved C-vine Copula model is proposed. First, based on long-term multi-energy complementarity requirements, multi-year runoff data from a power station and power output data from wind and solar power stations within the watershed are selected, outliers are removed, and missing values ​​are filled, completing data preprocessing. Next, runoff and wind / solar output are set as random variables, and nonparametric kernel density estimation is used to obtain the marginal distribution functions of each variable. Then, the improved C-vine Copula model is used to accurately characterize the spatiotemporal correlation between water, wind, and solar resources, deriving the joint probability distribution. Then, Latin hypercube sampling is used to collect uniformly random samples stratified, and K-means clustering is used to generate typical scenarios. Finally, the effectiveness of the scenarios is evaluated from the perspectives of temporal and spatial correlation and randomness. This technology can effectively address the randomness and complexity of water, wind, and solar resources, generating realistic typical scenarios, providing valuable reference for the planning and scheduling of watershed hydro-wind-solar hybrid systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of typical water landscape scene generation technology, and specifically relates to a method for generating typical water landscape scenes based on an improved C-vineCopula. Background Technology

[0002] Hydropower, with its large capacity and strong regulation capabilities, can effectively complement wind and solar resources. However, on a long-term scale, runoff and wind / solar power generation capacity exhibit significant randomness and seasonality, and are correlated with each other, making accurate characterization difficult. This poses a significant challenge to the planning and operation of integrated hydro-wind-solar multi-energy complementary systems in watersheds. By considering the spatiotemporal correlation of hydropower, wind, and solar power, a representative high-dimensional coupled scenario set is generated to quantify and characterize the uncertainties of hydro-wind-solar systems in a finite discrete manner, thus providing effective support for the construction of planning and stochastic scheduling models.

[0003] A method for generating typical watershed hydro-wind-solar power scenarios considering multiple uncertainties is based on long-term historical runoff data from a large hydropower station and fitted output data from wind and solar power stations within its watershed. This historical data is used to simulate and quantify the multiple uncertainties of hydro-wind-solar power, generating a large number of runoff and wind / solar output sequences to preserve the statistical characteristics of historical data to the greatest extent possible, thus providing effective input data for power system planning, scheduling, and risk analysis. Currently, existing technologies have certain shortcomings in addressing the problem of quantifying multiple uncertainties in hydro-wind-solar resources. In representing the spatiotemporal correlation of multiple variables, using ordinary Copula functions to model multidimensional variables usually requires modeling the dependencies between all variables at once, and the model structure becomes more complex with increasing dimensionality. Furthermore, in hydro-wind-solar systems, the complex dependencies between inflow runoff and wind / solar output are not only direct but may also be conditional; a one-time model may not be able to capture these complex dependencies between random variables. Furthermore, when sampling the joint distribution of high-dimensional random variables, the Monte Carlo sampling method is used to handle high-dimensional problems. The number of samples increases exponentially with the increase of dimension, the convergence speed is slow, and the sampled samples may be unevenly distributed, resulting in poor coverage and failure to reflect the characteristics of the distribution. Summary of the Invention

[0004] The technical problem to be solved by this invention is to provide a method for generating typical watershed water-wind-light scenes based on an improved C-vine Copula. This method addresses the difficulty of existing methods in generating long-term water-wind-light coupled scene sets for watershed water-wind-light integration. Specifically, it addresses the technical problems in the water-wind-light scene generation process where existing Copula theory is used to fit the joint distribution of multidimensional variables through one-time modeling, resulting in complex model structure and high computational cost and low coverage when sampling the joint distribution of multidimensional random variables.

[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:

[0006] A method for generating typical watershed waterscape scenes based on an improved C-vine Copula, comprising the following steps:

[0007] Step 1: Data selection: Select the required historical data based on the long-term scale demand for multi-energy complementarity of water, wind and solar power. The required historical data includes the long-term historical runoff data of a power station and the power output data of wind and solar power stations in the watershed of the power station.

[0008] Step 2: Data preprocessing: Preprocess the required historical data, including filling missing values ​​and removing outliers;

[0009] Step 3: Let historical runoff and solar power output be random variables X. h ,X w ,X s The data distribution is smoothed using nonparametric kernel density estimation to reflect the true distribution of the data, and X is estimated accordingly. h ,X w ,X s The probability density function is obtained, and then the marginal distribution function of each random variable is obtained;

[0010] Step 4: Based on the marginal distribution function of water, scenery and light obtained in Step 3, the spatiotemporal correlation among water, scenery and light is characterized by the improved C-vine Copula model, and the joint probability distribution of water, scenery and light is obtained.

[0011] Step 5: Based on the joint probability distribution of water, wind and light in Step 4, the Latin hypercube sampling method is used to perform stratified sampling from the multidimensional distribution. By dividing the sample space of each dimension into uniform sub-intervals, a more uniform random sample set is obtained.

[0012] Step 6: Perform cluster analysis on the random sample set obtained in Step 5, and use the K-means clustering method to cluster all samples into several representative typical scenarios;

[0013] Step 7: Based on the typical scenarios obtained in Step 6, evaluate the effectiveness of the generated typical scenarios to ensure their representativeness and practicality.

[0014] Preferably, in step 1, considering the inflow of hydropower stations into the reservoir and assuming that new energy sources such as wind and solar power within the reservoir basin are clustered and connected to the power grid according to energy type, monthly average wind power and photovoltaic power output sequence data of no less than thirty years are selected.

[0015] Preferably, in step 2, the preprocessing includes removing outliers from the required historical data and using interpolation to fill missing values. For the existing dataset, mathematical calculations are used to obtain the monthly average values ​​of inflow runoff and wind and solar power output data, respectively.

[0016] Preferably, in step 3, runoff and wind and solar power output are respectively variables X. h ,X w ,X s ;

[0017] (1)X w The process of solving the distribution function is as follows:

[0018] If x 1w ,x 2w ...x nw For n independent and identically distributed wind power output sample points F, let their probability density function be f, and the kernel density estimate is as follows:

[0019]

[0020] In the formula, h > 0 is a smoothing parameter, called bandwidth; K(·) is the kernel function; the Gaussian kernel function is chosen, and its expression is as follows:

[0021]

[0022] By using the data and bandwidth of each data point as parameters of the kernel function, N kernel functions are obtained. These are then linearly superimposed to form the kernel density estimation function. After normalization, the kernel density probability density function is obtained; thus, the distribution function of each random variable is obtained.

[0023] (2)X h ,X s The solution process and X w same.

[0024] Preferably, in step 4, the maximum column sum method is used to construct the dependency relationship between each random variable in a hierarchical manner; the information criterion method is used to select the optimal truncation level to determine the final improved C-vine Copula model to characterize the dependency relationship between water, scenery and light.

[0025] Preferably, the specific method of step 4 is as follows:

[0026] (1) Construct a C-vine Copula model of inflow runoff and wind and solar power output, and select the Kendall coefficient to characterize the correlation between variables;

[0027] Now, assuming any pair of n-dimensional random variables (X, Y), the Kendall correlation coefficient is calculated using the following formula:

[0028] τ=P[(X-X')(Y-Y')>0]-P[(X-X')(Y-Y')<0] (18);

[0029] Where P(·) is the probability density function, and the random vectors (X,Y) and (X',Y') follow the same distribution;

[0030] The three random variables X (water, wind, and light) are calculated using formula (3). h X w X s The Kendall coefficients between each pair of nodes are used; then, the maximum column sum method is used to select the best root node, and the weight matrix is ​​constructed as follows:

[0031]

[0032] Where, τ xy Kendall coefficients;

[0033] (2) Calculate column sums to select the root node: The dependencies between the root node variables and other node variables form the first layer of C-vine Copula's structure. The second layer structure is to construct the dependencies between the remaining variables based on the first layer root node variables.

[0034] The optimal pair Copula function is selected for each edge of the vine structure. The hybrid Copula function is then combined, and the weight parameters and dependency coefficients of the hybrid Copula are estimated using the EM algorithm. The constructed hybrid Copula form is as follows:

[0035]

[0036] In the formula: u, v are the marginal distributions of random variables; λ n ∈[0,1] represents the weight parameters of the model, and These are the dependency parameters of the model;

[0037] (3) The information criterion method is used to select the optimal cutoff level. First, different cutoff levels are assumed, and the overall likelihood estimate of the C-vine Copula model under the corresponding cutoff level is calculated. The formula is as follows:

[0038]

[0039] Where: L i Let m be the likelihood estimate of the i-th layer of the C-vine Copula model, m be the total number of layers in the C-vine Copula model, and L be the overall likelihood estimate.

[0040] Based on the overall likelihood estimates obtained above, the AIC and BIC are calculated to determine the optimal cutoff level, as shown in the following formula:

[0041] AIC = 2k - xln(L) (22);

[0042] BIC = kln(n) - 2ln(L) (23);

[0043] Where k is the number of parameters, and L is the likelihood function. n It refers to sample size; AIC is the information criterion, and BIC is the Bayesian information criterion;

[0044] Based on the principle of minimizing AIC and BIC, the optimal truncation level is selected as level 2 of the C-vine Copula model; the final determined C-vine Copula structure is as follows, assuming n-dimensional variables X = (x1, x2, ..., x...). n ), F(x1,x2,...,x n and f(x1,x2,...,x) n F represents its joint distribution function and joint probability density function, respectively. i (x i ) and f i (x i Let ) represent the cumulative distribution function and probability density function, respectively; then the joint probability distribution of the C-vine structure of X is expressed as follows:

[0045]

[0046] In the formula, c j,j+i|11...j-1 (·) represents the known x1, x2, ..., x j-1 Under the condition, variable x j and x j+i The Copula probability density function formed by the two; F(x) j |x1,…,x j-1 Given x1, x2, ..., x j-1 Under the condition that variable x j The distribution function.

[0047] Preferably, step 5 is specifically performed as follows:

[0048] Latin hypercube sampling is performed on the improved C-vine Copula model. Now, assume the sample size is N, and the Nth sample is denoted as s. n =[s nh ,s nw ,s ns The specific steps are as follows:

[0049] (1) Assume x wLet x be a root node random variable and x w ∈[x wd ,x wu ], x w The marginal distribution function is f w (x w ); The range [f w (x wd ),f w (x wu The interval is divided into N equal probability intervals, and q is randomly selected from all probability sub-intervals. i Thus, the uniform variable Z1 is obtained;

[0050] (2) Let U1, U2, U3 be three sets of variables to be determined. The uniform variable obtained from the above steps is equal to U1, that is, U1 = Z1. Z1 is regarded as the sampling point of the variable to be determined.

[0051] (3) Similarly, generate random numbers that follow a uniform distribution in the manner described in step (1) above, and define them as uniform variables Z2 and Z3. From the improved C-type model and the conditional distribution function of equation (11), it can be seen that the second set of variables to be determined, U2, can be obtained using... The calculation is performed, where Z2 and U1 are known quantities. This transforms the above univariate nonlinear equation into a linear equation, which is then solved using the bisection method. The solution obtained is the sample points of the variable U2 to be solved.

[0052] (4) Similarly, since Z2 and Z3 have already been defined, we can use:

[0053]

[0054] We obtain F(x3|x1), and... Solve the equation; the result is the sample of the variable U3 to be solved.

[0055] (5) Finally, inverse transformation sampling is performed on U1, U2, and U3 to obtain the water, wind, and light scene set S = [S h ,S w ,S s ].

[0056] Preferably, in step 6, the dataset is divided into K clusters using K-means clustering sampling distance as an evaluation metric for similarity; the steps for performing K-means clustering on the scene set obtained in step 5 are as follows:

[0057] (1) Initialization: Randomly select K sample points from the dataset as the initial centroids;

[0058] (2) Assign sample points to the nearest centroid: For each sample point, calculate its Euclidean distance to all centroids and assign it to the cluster containing the nearest centroid. The Euclidean distance formula is as follows:

[0059]

[0060] Among them, C m With the center of mass, X i For sample points, d k (X i C m () represents the Euclidean distance from the sample point to the centroid;

[0061] (3) Update centroid: Calculate the average value of all sample points in the cluster and use it as the new centroid;

[0062] (4) Repeat steps (2) and (3): continuously assign sample points and update centroids until the centroids no longer change or the preset number of iterations is reached;

[0063] (5) Assign weights and count the number of scenes in each cluster. If a cluster contains n scenes, then... i If there are several scenarios, then the weight p of the representative scenario in that cluster is... i for:

[0064]

[0065] Where, n i p is the total number of sample points contained in the i-th cluster. i For the corresponding weight parameters;

[0066] (6) Output the centroid and corresponding weight parameters for each cluster.

[0067] Preferably, in step 7, the generated typical scene is evaluated by taking into account temporal correlation, spatial correlation and randomness;

[0068] The autocorrelation coefficient is used as an indicator of time correlation; its formula is as follows:

[0069]

[0070] In the formula, r t This represents the value for month t; k represents the time interval. ρ represents the mean. k This represents the autocorrelation coefficient with a time delay of k;

[0071] The average absolute error of the Kendall correlation coefficient is selected to evaluate the representation quality of the spatial correlation of the generated typical scenes; the formula is as follows:

[0072]

[0073] In the formula: τ i Let be the Kendall correlation coefficient between the i-th group of typical scenarios; p represents the Kendall correlation coefficient between historical time-series data. i Let be the weight parameters for the i-th group of typical scenarios;

[0074] Randomness refers to the deviation between the actual runoff and wind and solar power output scenarios and the generated scenario set over a time span. Scenario evaluation is based on distance, and the formula is as follows:

[0075]

[0076] In the formula: R i Let R be the i-th typical scene sequence; R be the historical actual sequence; ||·||2 is the L2 norm of the difference between the two sequences; E represents the average Euclidean distance between the typical scene set and the historical actual sequence.

[0077] A system for generating typical watershed waterscape scenes based on an improved C-vine Copula is provided, which employs the aforementioned method for generating typical watershed waterscape scenes based on an improved C-vine Copula.

[0078] The present invention can achieve the following beneficial effects:

[0079] 1) This invention innovatively utilizes an improved C-vine Copula model to characterize the spatiotemporal correlation between random variables. This model decomposes the high-dimensional joint distribution into a series of low-dimensional pair-Copula functions by hierarchically modeling the dependency structure of all random variables, while truncating at appropriate levels, thereby simplifying the calculation and modeling process and effectively balancing the complexity and fitting accuracy of the model.

[0080] 2) This invention employs the Latin hypercube sampling method to perform stratified sampling of the joint distribution of multidimensional random variables. The sample space for each dimension is uniformly divided into several sub-intervals, and sample points are randomly selected within each sub-interval. This not only ensures the uniformity of sampling but also improves sampling efficiency and coverage. Based on this, the K-means clustering algorithm is combined with Euclidean distance as an evaluation index for scene similarity to further analyze and optimize the water-wind-light system. Attached Figure Description

[0081] The present invention will be further described below with reference to the accompanying drawings and embodiments:

[0082] Figure 1 This is the overall flowchart of the present invention;

[0083] Figure 2This is a flowchart of the improved C-vine Copula model of this invention;

[0084] Figure 3 This is a structural diagram of the improved C-vine Copula model of this invention;

[0085] Figure 4 This is a flowchart of the Latin hypercube algorithm of this invention;

[0086] Figure 5 This is a flowchart of the K-means clustering algorithm of this invention. Detailed Implementation

[0087] Preferred solutions include Figures 1 to 5 As shown, a method for generating typical watershed hydro-wind-solar complementary scenarios based on an improved C-vine Copula is proposed. First, long-term historical inflow and wind / solar output data are used as model inputs, and non-parametric kernel density estimation is employed to fit the marginal distributions of each random variable. Second, based on the improved C-Vine Copula function, a hierarchical sampling modeling approach is used to construct the spatiotemporal correlations between the random variables, with truncation performed at certain levels. Next, Latin hypercube sampling is combined to hierarchically sample the joint distribution of the multivariate coupled system, generating a scenario set for the hydro-wind-solar multi-energy complementary system. Then, K-means clustering is used to cluster a large number of samples into a few representative scenarios. Finally, the obtained representative scenarios are effectively evaluated.

[0088] The following section describes the overall process. Figure 1 The detailed explanation mainly includes the following steps:

[0089] Step 1: Data selection. Based on the long-term scale demand for multi-energy complementarity of water, wind and solar power in this study, the required historical data are selected. The required historical data includes long-term historical runoff data of a large hydropower station and power output data of wind and solar power stations in the watershed of the station.

[0090] Step 2: Data preprocessing. Select the required historical data according to the method in Step 1 and preprocess the data, including filling missing values ​​and removing outliers.

[0091] Step 3: Let runoff and solar power output be random variables X. h ,X w ,X s Nonparametric kernel density estimation is used to smooth the data distribution to reflect the true distribution of the data, and X is estimated respectively. h ,X w ,X s The probability density function is obtained, and then the marginal distribution function of each random variable is obtained.

[0092] Step 4: Based on the marginal distribution functions of water, scenery, and landscape obtained in Step 3, the spatiotemporal correlation among them is characterized using the C-vine Copula model. First, the maximum column sum method is used to hierarchically construct the dependencies between the random variables. Then, the information criterion method is used to select the optimal truncation level to determine the final improved C-vine Copula model to characterize the dependencies among the three.

[0093] Step 5: Based on the joint distribution of water, wind, and solar energy from Step 4, stratified sampling is performed from the multidimensional distribution using the Latin hypercube sampling method. This achieves more uniform sampling by dividing the sample space of each dimension into uniform sub-intervals. A more uniform random sample set is obtained by dividing the sample space of each dimension into uniform sub-intervals.

[0094] Step 6: Perform cluster analysis on the random sample set obtained in Step 5, and use the K-means clustering method to cluster all samples into several representative typical scenarios.

[0095] Step 7: Based on the typical scenarios obtained in Step 6, evaluate the effectiveness of the generated typical scenarios to ensure their representativeness and practicality.

[0096] Step 1, as shown, specifically considers the inflow of a large hydropower station and assumes that new energy sources such as wind and solar power within the reservoir basin are connected to the grid in their respective clusters. The inflow data is provided by the relevant entities, and the wind and solar power output data is estimated based on historical meteorological data using a corresponding model developed by the National Renewable Energy Laboratory, resulting in the monthly average historical wind and solar power output over a long period of time.

[0097] Considering the inflow of hydropower stations into the reservoir and assuming that new energy sources such as wind and solar power within the reservoir basin are clustered and connected to the power grid according to energy type, monthly average wind power and solar power output sequence data of no less than 30 years are selected.

[0098] Specifically, step 2 involves preprocessing the historical data selected in step 1, including removing outliers and filling missing values ​​using interpolation. For the existing dataset, mathematical calculations are used to obtain the monthly average values ​​of inflow runoff and wind and solar power output data.

[0099] Specifically, in step 3 shown, let runoff and solar power output be random variables X. h ,X w ,X s With wind power output X w For example, if x 1w ,x 2w x nw For n independent and identically distributed wind power output sample points F, let their probability density function be f, and the kernel density estimate is as follows:

[0100]

[0101] In the formula, h > 0 is a smoothing parameter, called bandwidth; K(·) is the kernel function. Commonly used kernel functions include the Epanechikov function and the Gaussian function. This invention chooses the Gaussian kernel function, the expression of which is as follows:

[0102]

[0103] By using the data and bandwidth of each data point as parameters of the kernel function, N kernel functions are obtained. These are then linearly superimposed to form the kernel density estimation function, which, after normalization, becomes the kernel density probability density function. This leads to the distribution function of each random variable.

[0104] Specifically, step 4 involves first constructing a C-vine Copula model for inflow runoff and solar power output, and then selecting an appropriate correlation coefficient to characterize the relationship between the variables. Here, the Kendall coefficient is used. Assuming any pair of n-dimensional random variables (X, Y), the Kendall correlation coefficient is calculated as follows:

[0105] τ=P[(X-X')(Y-Y')>0]-P[(X-X')(Y-Y')<0] (33)

[0106] Where P(·) is the probability density function, and the random vectors (X,Y) and (X',Y') follow the same distribution.

[0107] The three random variables X of water, wind, and light are calculated using the above formulas. h X w X s The Kendall coefficients between each pair of nodes. Then, the maximum column sum method is used to select the optimal root node, constructing the weight matrix as follows:

[0108]

[0109] Where, τ xy This is the Kendall coefficient.

[0110] The root node is selected by calculating the column sum. The dependencies between the root node variables and other node variables form the first layer structure of the C-vine Copula. The second layer structure is constructed by using the root node variables of the first layer as conditions to build the dependencies between the remaining variables. Regarding the selection of the optimal pair Copula function for each edge of the vine structure, this invention introduces a hybrid Copula function and uses the EM algorithm to estimate the weight parameters and dependency coefficients of the hybrid Copula. The constructed hybrid Copula form is as follows:

[0111]

[0112] In the formula: u, v are the marginal distributions of random variables; λ n ∈[0,1] represents the weight parameters of the model, and These are the dependency parameters of the model.

[0113] Next, the information criterion method is used to select the optimal cutoff level. First, assuming different cutoff levels, the overall likelihood estimate of the C-vine Copula model at the corresponding cutoff level is calculated, as shown in the following formula:

[0114]

[0115] Where: L i Let m be the likelihood estimate of the i-th layer of the C-vine Copula model, m be the total number of layers in the C-vine Copula model, and L be the overall likelihood estimate.

[0116] Based on the overall likelihood estimate obtained above, the AIC (Akaike Information Criterion) and BIC (Bayesian Information Criterion) are calculated to determine the optimal cutoff level, as shown in the following formula:

[0117] AIC=2k-xln(L) (37)

[0118] BIC=kln(n)-2ln(L) (38)

[0119] Where k is the number of parameters, and L is the likelihood function. n That is the sample size.

[0120] Based on the principle of minimizing AIC and BIC, the optimal truncation level is chosen to be the second layer of the C-vine Copula model. The final determined C-vine Copula structure is as follows, assuming there are n-dimensional variables X = (x1, x2, ..., x...). n ), F(x1,x2,...,x n and f(x1,x2,...,x) n F represents its joint distribution function and joint probability density function, respectively. i (x i ) and f i (x i Let represent the cumulative distribution function and the probability density function, respectively. Then the joint probability distribution of the C-vine structure of X is expressed as follows:

[0121]

[0122] In the formula, c j,j+i|11...j-1(·) represents the known x1, x2, ..., x j-1 Under the condition, variable x j and x j+i The Copula probability density function formed by the two; F(x) j |x1,…,x j-1 Given x1, x2, ..., x j-1 Under the condition that variable x j The distribution function.

[0123] Specifically, step 5 involves performing Latin hypercube sampling on the improved C-vine Copula model. Assuming the sample size is N, the Nth sample can be represented as s. n =[s nh ,s nw ,s ns The specific steps are as follows:

[0124] (1) Assume x w Let x be a root node random variable and x w ∈[x wd ,x wu ], x w The marginal distribution function is f w (x w ). The range [f w (x wd ),f w (x wu The interval is divided into N equal probability intervals, and q is randomly selected from all probability sub-intervals. i Thus, the uniform variable Z1 is obtained.

[0125] (2) Let U1, U2, U3 be three sets of variables to be determined. The uniform variable obtained from the above steps is equal to U1, that is, U1 = Z1. Z1 can be regarded as the sampling point of the variable to be determined.

[0126] (3) Similarly, generate random numbers that follow a uniform distribution in the manner described in step 1 above, and define them as uniform variables Z2 and Z3. From the improved C-type model and the conditional distribution function of equation (11), it can be seen that the second set of variables to be determined, U2, can be obtained using... The calculation is performed, where Z2 and U1 are known quantities. This transforms the above univariate nonlinear equation into a linear equation, which is then solved using the bisection method. The solution obtained is the sample points of the variable U2 to be solved.

[0127] (4) Similarly, since Z2 and Z3 have already been defined, they can be derived from...

[0128]

[0129] We obtain F(x3|x1), and... Solve the equation. The result is a sample of the variable U3 to be solved.

[0130] (5) Finally, inverse transformation sampling is performed on U1, U2, and U3 to obtain the water, wind, and light scene set S = [S h ,S w ,S s ].

[0131] Step 6 specifically involves dividing the dataset into K clusters using K-means clustering sampling distance as a similarity evaluation metric. The following are the steps for K-means clustering based on the scene set obtained in step (5).

[0132] (1) Initialization: Randomly select K sample points from the dataset as the initial centroids.

[0133] (2) Assign sample points to the nearest centroid: For each sample point, calculate its Euclidean distance to all centroids and assign it to the cluster containing the nearest centroid. The Euclidean distance formula is as follows:

[0134]

[0135] Among them, C m With the center of mass, X i For sample points, d k (X i C m ) represents the Euclidean distance from the sample point to the centroid.

[0136] (3) Update centroid: Calculate the average value of all sample points in the cluster and use it as the new centroid.

[0137] (4) Repeat steps (2) and (3): continuously assign sample points and update centroids until the centroids no longer change or the preset number of iterations is reached.

[0138] (5) Assign weights and count the number of scenes in each cluster. If a cluster contains n scenes, then... i If there are several scenarios, then the weight p of the representative scenario in that cluster is... i for:

[0139]

[0140] Where, n i p is the total number of sample points contained in the i-th cluster. i For the corresponding weight parameters.

[0141] (6) Output the centroid and corresponding weight parameters for each cluster.

[0142] In step (7), specifically, the present invention evaluates the generated typical scenarios by taking into account temporal correlation, spatial correlation and randomness.

[0143] The autocorrelation coefficient is used as an indicator of time correlation. Its formula is as follows:

[0144]

[0145] In the formula, r t Represents the value in month t; k represents the time interval; r represents the mean; ρ k This represents the autocorrelation coefficient with a time delay of k.

[0146] The average absolute error of the Kendall correlation coefficient is used to evaluate the representation quality of the spatial correlation of the generated typical scenes. The formula is as follows:

[0147]

[0148] In the formula: τ i Let be the Kendall correlation coefficient between the i-th group of typical scenarios; p represents the Kendall correlation coefficient between historical time-series data. i Let be the weight parameters for the i-th group of typical scenarios;

[0149] Randomness refers to the deviation between the actual runoff and wind and solar power output scenarios and the generated scenario set over a time span. Scenario evaluation is based on distance, and the formula is as follows:

[0150]

[0151] In the formula: R i Let R be the i-th typical scene sequence; R be the historical actual sequence; ||·||2 is the L2 norm of the difference between the two sequences; E represents the average Euclidean distance between the typical scene set and the historical actual sequence.

[0152] The above embodiments are merely preferred technical solutions of the present invention and should not be considered as limitations on the present invention. The scope of protection of the present invention should be limited to the technical solutions described in the claims, including equivalent substitutions of the technical features described in the claims. That is, equivalent substitutions and improvements within this scope are also within the scope of protection of the present invention.

Claims

1. A method for generating typical watershed waterscape scenes based on an improved C-vine Copula, characterized in that... Includes the following steps: Step 1: Data selection: Select the required historical data based on the long-term scale demand for multi-energy complementarity of water, wind and solar power. The required historical data includes the long-term historical runoff data of a power station and the power output data of wind and solar power stations in the watershed of the power station. Step 2: Data preprocessing: Preprocess the required historical data, including filling missing values ​​and removing outliers; Step 3: Let historical runoff and solar power output be random variables. Nonparametric kernel density estimation is used to smooth the data distribution, thereby reflecting the true distribution of the data. The kernel density is estimated separately. The probability density function is obtained, and then the marginal distribution function of each random variable is obtained; Step 4: Based on the marginal distribution function of water, scenery and light obtained in Step 3, the spatiotemporal correlation among water, scenery and light is characterized by the improved C-vine Copula model, and the joint probability distribution of water, scenery and light is obtained. Step 5: Based on the joint probability distribution of water, wind and light in Step 4, the Latin hypercube sampling method is used to perform stratified sampling from the multidimensional distribution. By dividing the sample space of each dimension into uniform sub-intervals, a more uniform random sample set is obtained. Step 6: Perform cluster analysis on the random sample set obtained in Step 5, and use the K-means clustering method to cluster all samples into several representative typical scenarios; Step 7: Based on the typical scenarios obtained in Step 6, evaluate the effectiveness of the generated typical scenarios to ensure their representativeness and practicality; In step 4, the maximum column sum method is used to construct the dependencies between the random variables in a hierarchical manner; the information criterion method is used to select the optimal truncation level to determine the final improved C-vine Copula model to characterize the dependencies between water, scenery and landscape. The specific method for step 4 is as follows: (1) Construct a C-vine Copula model for inflow runoff and wind and solar power output, and select the Kendall coefficient to characterize the correlation between variables; Now assume any pair 3D random variables The Kendall correlation coefficient is calculated using the following formula: (3); in, Let be the probability density function and the random vector. and They follow the same distribution; The three random variables of water, wind, and light are calculated using formula (3). The Kendall coefficients between each pair of nodes are used; then, the maximum column sum method is used to select the best root node, and the weight matrix is ​​constructed as follows: (4); in, Kendall coefficients; (2) Calculate the column sum to select the root node: The dependency relationship between the root node variable and other node variables forms the first layer structure of C-vineCopula. The second layer structure is to construct the dependency relationship between the remaining variables based on the first layer root node variable. The optimal pair Copula function is selected for each edge of the vine structure. The hybrid Copula function is then combined, and the weight parameters and dependency coefficients of the hybrid Copula are estimated using the EM algorithm. The constructed hybrid Copula form is as follows: (5); In the formula: It is a marginal distribution of a random variable; These are the weight parameters of the model, and ; These are the dependency parameters of the model; (3) The information criterion method is used to select the optimal cutoff level. First, different cutoff levels are assumed, and the overall likelihood estimate of the C-vine Copula model under the corresponding cutoff level is calculated. The formula is as follows: (6); in: For the C-vine Copula model Likelihood estimate of layer This represents the total number of layers in the C-vine Copula model. This is the overall likelihood estimate; Based on the overall likelihood estimates obtained above, the AIC and BIC are calculated to determine the optimal cutoff level, as shown in the following formula: (7); (8); in, It is the number of parameters. It is the likelihood function. It refers to sample size; AIC is the information criterion, and BIC is the Bayesian information criterion; Based on the principle of minimizing AIC and BIC, the optimal truncation level is selected as layer 2 of the C-vine Copula model; the final determined C-vine Copula structure is as follows, assuming the following... Dimensional variables , as well as These are their joint distribution function and joint probability density function, respectively. and Let these represent the cumulative distribution function and the probability density function, respectively; then... The joint probability distribution of the C-vine structure is expressed as follows: (9); In the formula, Known Under the condition, variables and The Copula probability density function formed by the two; Known Under the condition, variables The distribution function.

2. The method for generating typical watershed waterscape scenes based on the improved C-vine Copula according to claim 1, characterized in that: In step 1, considering the inflow of hydropower stations into the reservoir and assuming that wind and solar new energy sources in the reservoir basin are clustered and connected to the power grid according to energy type, monthly average wind power and photovoltaic output sequence data of no less than 30 years are selected.

3. The method for generating typical watershed waterscape scenes based on the improved C-vine Copula according to claim 1, characterized in that: In step 2, preprocessing includes removing outliers from the required historical data and using interpolation to fill missing values. For the existing dataset, mathematical calculations are used to obtain the monthly average values ​​of inflow runoff and wind and solar power output data.

4. The method for generating typical watershed waterscape scenes based on the improved C-vine Copula according to claim 1, characterized in that: In step 3, runoff and solar power output are set as variables. , (1) The process of solving the distribution function is as follows: like Independent and identically distributed of Let there be a sample of wind power output points, and its probability density function be... The nuclear density is estimated as follows: (1); In the formula, This is a smoothing parameter, called bandwidth; For the kernel function; we choose the Gaussian kernel function, with the following expression: (2); By using the data and bandwidth of each data point as parameters of the kernel function, N kernel functions are obtained. These are then linearly superimposed to form the kernel density estimation function. After normalization, the kernel density probability density function is obtained; thus, the distribution function of each random variable is obtained. (2) The solution process and same.

5. The method for generating typical watershed waterscape scenes based on the improved C-vine Copula according to claim 1, characterized in that: Step 5 is as follows: Latin hypercube sampling was performed on the improved C-vine Copula model, assuming the sample size was... , No. Each sample is represented as The specific steps are as follows: (1) Assumption The root node random variable and , The marginal distribution function is ; to range Divided into There are several equally probable intervals, and each interval is randomly selected from all its sub-intervals. To obtain uniform variables ; (2) Let Given three sets of variables to be determined, the uniform variable obtained from the above steps is equal to... ,Right now , These are considered as sampling points for the variable to be determined. (3) Similarly, generate random numbers that follow a uniform distribution in the manner described in step (1) above, and define them as uniform variables. From the improved C-vine model and the conditional distribution function of the Euclidean distance formula, it can be seen that the second group of variables to be solved... use Calculation, where and All of these are known quantities, therefore the formula can be derived from them. Transform the equation into a linear equation and solve it using the bisection method. The solution to the resulting equation is the variable to be solved. Sample points; (4) Similarly, because It has already been defined by: (10); Seeking ,again Solve the equation; the result is the variable to be solved. ; (5) Finally, for Inverse transformation sampling is performed to obtain a set of water, wind, and light scenes that takes into account multiple uncertainties. .

6. The method for generating typical watershed waterscape scenes based on the improved C-vine Copula according to claim 5, characterized in that: In step 6, the dataset is divided into K clusters using K-means clustering sampling distance as a similarity evaluation metric. The steps for K-means clustering based on the scene set obtained in step 5 are as follows: (1) Initialization: Randomly select K sample points from the dataset as the initial centroids; (2) Assign sample points to the nearest centroid: For each sample point, calculate its Euclidean distance to all centroids and assign it to the cluster containing the nearest centroid. The Euclidean distance formula is as follows: (11); in, Center of mass, For sample points, Let Euclidean distance be the distance from the sample point to the centroid. (3) Update centroid: Calculate the average value of all sample points in the cluster and use it as the new centroid; (4) Repeat steps (2) and (3): continuously assign sample points and update centroids until the centroids no longer change or the preset number of iterations is reached; (5) Assign weights and count the number of scenes in each cluster. If a cluster contains... If there are several scenarios, then the weight of the representative scenario in that cluster is... for: (12); in, For the first The total number of sample points contained in each cluster For the corresponding weight parameters; (6) Output the centroid and corresponding weight parameters for each cluster.

7. The method for generating typical watershed waterscape scenes based on the improved C-vine Copula according to claim 1, characterized in that: In step 7, the generated typical scenarios are evaluated by taking into account temporal correlation, spatial correlation, and randomness. The autocorrelation coefficient is used as an indicator of time correlation; its formula is as follows: (13); In the formula, Indicates the first The value of the month; Indicates a time interval; This represents the mean; Indicates a time delay of The autocorrelation coefficient; The average absolute error of the Kendall correlation coefficient is selected to evaluate the representation quality of the spatial correlation of the generated typical scenes; the formula is as follows: (14); In the formula: For the first Kendall correlation coefficients between typical scenarios; Kendall correlation coefficient between historical time series data; For the first Weight parameters for a group of typical scenarios; Randomness refers to the deviation between the actual runoff and wind and solar power output scenarios and the generated scenario set over a time span. Scenario evaluation is based on distance, and the formula is as follows: (15); In the formula: For the first A sequence of typical scenarios; This is the actual historical sequence; The difference between two sequences Norm; This represents the average Euclidean distance between the typical scene set and the actual historical sequence.

8. A system for generating typical watershed waterscape scenes based on an improved C-vine Copula, characterized by: The method for generating typical watershed waterscape scenes based on the improved C-vine Copula, as described in claim 1, was adopted.

Citation Information

Patent Citations

  • Water-wind-light long-term scene generation method considering space-time correlation

    CN117609586A

  • Wind-solar combined output uncertainty modeling method and system

    CN118520686A