New energy joint output scene generation system and method based on data driving
By constructing a joint probability distribution model of wind power and photovoltaic output and generating a set of random output scenarios with time correlation, the problem of difficulty in characterizing the correlation between wind power and photovoltaic output in traditional power systems is solved, achieving more accurate output prediction and stability optimization of the power system.
Patent Information
- Application Number
- CN202510661438.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-09-09
AI Technical Summary
Traditional power systems find it difficult to effectively characterize the nonlinear dependence of wind power and photovoltaic output, resulting in large deviations between the generation of new energy scenarios and the actual operating characteristics of the power grid, and are unable to adapt to the needs of large-scale new energy access.
Through a data-driven approach, a joint probability distribution model of wind power and photovoltaic output is constructed. The Kendall rank correlation coefficient and tail dependence coefficient are used to analyze their correlation. The Markov Chain Monte Carlo algorithm is combined to generate a set of random output scenarios with time correlation, and representative scenarios are screened through a clustering algorithm.
It improves the accuracy of renewable energy output forecasts, reduces power system operation risks, optimizes operation strategies, reduces backup capacity requirements, reduces additional costs, and enhances the stability and economy of the power system.
Smart Images

Figure CN120613796A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of new energy power system planning and operation, and specifically to a data-driven new energy joint output scenario generation system and method. Background Art
[0002] With the global pursuit of a low-carbon economy and sustainable development, the proportion of wind power and photovoltaic power as new energy sources in the power system continues to increase. However, wind power output and photovoltaic power output are significantly intermittent and uncertain, and their power output is strongly affected by meteorological conditions. This uncertainty poses a huge challenge to the planning, operation, and scheduling of the power system. The traditional power system operation mode is difficult to adapt to the needs of large-scale access to new energy. The current new energy power system faces technical bottlenecks such as insufficient modeling accuracy of wind power output and photovoltaic output and lack of physical constraints. Traditional linear correlation methods cannot characterize the nonlinear dependence of wind power output and photovoltaic output. Traditional scenario generation methods mostly only consider the single time series characteristics of renewable energy output, while ignoring the temporal correlation between wind power output and photovoltaic output. It is difficult to effectively characterize the dynamic dependence of the time series of wind power output and photovoltaic output, resulting in a large deviation between the scenario and the actual grid operation characteristics. Summary of the Invention
[0003] The purpose of the present invention is to provide a data-driven new energy joint output scenario generation system and method, which improves the stability of the power system and reduces operating costs.
[0004] To achieve this goal, the present invention designs a data-driven new energy joint output scenario generation system, which includes:
[0005] The data preprocessing module is used to preprocess the historical output data of wind power output and photovoltaic output;
[0006] The joint probability distribution model construction module is used to fit the pre-processed historical output data of wind power output and photovoltaic output respectively through the probability density function to obtain the wind power output marginal distribution and the photovoltaic output marginal distribution, and to perform the correlation analysis between the wind power output and the photovoltaic output on the pre-processed historical output data of wind power output and photovoltaic output through the Kendall rank correlation coefficient and the tail dependence coefficient, and obtain the Kendall rank correlation coefficient matching error and the tail dependence coefficient matching error of each set connection function and the pre-processed historical output data of wind power output and photovoltaic output. The Euclidean distance between each set connection function and the pre-processed historical output data of wind power output and photovoltaic output is calculated, and the optimal connection function is selected in combination with the Kendall rank correlation coefficient matching error and the tail dependence coefficient matching error. Based on the optimal connection function and in combination with the wind power output marginal distribution and the photovoltaic output marginal distribution, a joint probability distribution model of wind power output and photovoltaic output is established;
[0007] The scenario generation module is used to use the Markov chain in combination with the joint probability distribution model to form a joint probability distribution model with time correlation, and to generate a random scenario set of new energy joint output from the joint probability distribution model with time correlation by sampling through the Markov chain Monte Carlo algorithm.
[0008] Preferably, the scenario reduction and evaluation module is used to cluster the random scenario set of new energy joint output using a clustering algorithm to obtain representative scenarios, and evaluate the degree of matching between the representative scenarios and the preprocessed historical output data of wind power output and photovoltaic output through KL divergence.
[0009] Preferably, the specific method for preprocessing the historical output data of wind power output and photovoltaic output includes:
[0010] The historical output data of wind power and photovoltaic power are normalized and preprocessed. The minimum and maximum normalization is used to linearly map the historical output data of wind power and photovoltaic power to the [0,1] per-unit value interval. The preprocessed historical output data of wind power and photovoltaic power are obtained. The formula is:
[0011]
[0012] Among them, m is the historical output data of wind power output and photovoltaic output, m min is the minimum value of the historical output data of wind power output and photovoltaic output, m max is the maximum value of the historical output data of wind power output and photovoltaic output, and m′ is the historical output data of wind power output and photovoltaic output after preprocessing.
[0013] Preferably, the specific process of fitting the pre-processed historical output data of wind power output and photovoltaic output respectively to obtain the marginal distribution of wind power output and photovoltaic output by using the probability density function is as follows:
[0014] The Weibull distribution probability density function of the marginal distribution of wind power output is obtained as follows:
[0015]
[0016] Where v is the wind speed; λ is the scale parameter of the Weibull distribution; k is the shape parameter of the Weibull distribution; is the scaling factor; Reflects the power relationship of wind speed v relative to scale parameter λ; is the index term;
[0017] The beta distribution probability density function of the photovoltaic output marginal distribution is obtained as follows:
[0018]
[0019] Where P is the photovoltaic output-related variable; α is the first shape parameter of the Beta distribution; β is the second shape parameter of the Beta distribution; P α-1 Indicates the impact of the value of α on the probability density; (1-P) β-1 represents the probability density characteristics of photovoltaic output; B(α,β) is the beta function.
[0020] Preferably, the correlation analysis between wind power output and photovoltaic output is performed on the pre-processed historical output data of wind power output and photovoltaic output by using the Kendall rank correlation coefficient and the tail dependence coefficient, and the specific process of obtaining the matching error between each set connection function and the Kendall rank correlation coefficient and the tail dependence coefficient of the pre-processed historical output data of wind power output and photovoltaic output is as follows:
[0021] The Kendall rank correlation coefficient τ is calculated by statistically analyzing the synergistic change ratio of historical output data of wind power output X and photovoltaic output Y. The calculation formula is as follows:
[0022]
[0023] Among them, N c is the number of same-direction logarithms at all moments in the historical output data. i >x j ) and (y i >y j ), or (x i <x j ) and (y i <y j ),(x i ,y i ) and (x j ,y j ) is the same direction logarithm; N d is the number of reverse logarithms of all moments in the historical output data. i >x j ) and (y i <y j ), or (x i <x j ) and (y i >y j ),(x i ,y i ) and (x j ,y j ) is the reverse logarithm; n represents the total number of moments in the historical output data; x i represents the wind power output at time i, x j represents the wind power output at time j, y irepresents the photovoltaic output at time i, y j represents the photovoltaic output at time j, and τ represents the Kendall rank correlation coefficient, which is used to measure the intensity of the trend of simultaneous increase or decrease of wind power output X and photovoltaic output Y;
[0024] Calculate the tail dependence coefficient λ of the preprocessed historical output data of wind power output and photovoltaic output. The tail dependence coefficient λ includes the upper tail dependence coefficient λ u and the lower tail dependence coefficient λ l , upper tail dependence coefficient λ u It represents the probability of wind power output and photovoltaic output reaching maximum value at the same time, and the lower tail dependence coefficient λ l Indicates the probability that wind power output and photovoltaic output reach their minimum values at the same time;
[0025] The upper tail dependence coefficient calculation formula is:
[0026]
[0027] Among them, λ u represents the upper tail dependence coefficient, q represents the quantile threshold, represents the inverse cumulative distribution function of the wind power output X variable, represents the inverse cumulative distribution function of the photovoltaic output Y variable, Wind power output And photovoltaic output probability;
[0028] The formula for calculating the lower tail dependence coefficient is:
[0029]
[0030] Among them, λ l represents the lower tail dependence coefficient, Wind power output And photovoltaic output probability;
[0031] Based on the correlation analysis between wind power output and photovoltaic output, the matching errors of each set connection function are obtained, including the Kendall rank correlation coefficient matching error, the upper tail dependence coefficient matching error, and the lower tail dependence coefficient matching error;
[0032] The Kendall rank correlation coefficient matching error calculation formula is:
[0033]
[0034] where τ 理论 is the theoretical Kendall rank correlation coefficient of each link function obtained by setting the parameters of each link function; τ 样本is the Kendall rank correlation coefficient τ of the preprocessed historical output data of wind power output and photovoltaic output;
[0035] The upper tail dependency coefficient matching error calculation formula is:
[0036]
[0037] where λ u,理论 is the theoretical tail dependence coefficient of each link function obtained by setting the parameters of each link function; τ 样本 is the upper tail dependence coefficient λ of the preprocessed historical output data of wind power output and photovoltaic output u ;
[0038] The formula for calculating the matching error of the lower tail dependence coefficient is:
[0039]
[0040] where λ I,理论 is the theoretical tail dependence coefficient of each link function obtained by setting the parameters of each link function; I,样本 is the lower tail dependence coefficient λ of the preprocessed historical output data of wind power output and photovoltaic output l .
[0041] Preferably, the Euclidean distance between each set connection function and the pre-processed historical output data of wind power output and photovoltaic output is calculated, and the optimal connection function is selected in combination with the matching error. The specific process of establishing the joint probability distribution model of wind power output and photovoltaic output based on the optimal connection function and in combination with the marginal distribution of wind power output and the marginal distribution of photovoltaic output is as follows:
[0042] Calculate the Euclidean distance between each copula function and the preprocessed historical output data of wind power output and photovoltaic output. The calculation formula of the Euclidean distance matching method is as follows:
[0043]
[0044] Where n represents the total number of moments in the historical output data, u i represents the historical output data of wind power after preprocessing, v i Represents the historical output data of photovoltaic power after preprocessing, C e represents the real dependency value directly calculated based on the pre-processed historical output data of wind power output and photovoltaic output, C t Indicates the theoretical dependence value of each set connection function Copula;
[0045] Selecting a connection function Copula with the smallest Euclidean distance from the preprocessed historical output data of wind power output and photovoltaic output, obtaining the Kendall rank correlation coefficient matching error, the upper tail dependence coefficient matching error, and the lower tail dependence coefficient matching error of the connection function Copula with the smallest Euclidean distance from the matching errors calculated above, and determining whether the Kendall rank correlation coefficient matching error, the upper tail dependence coefficient matching error, and the lower tail dependence coefficient matching error of the connection function Copula with the smallest Euclidean distance are all less than preset values; if so, the connection function is used as the optimal connection function Copula;
[0046] The optimal connection function Copula is combined with the marginal distribution of wind power output and photovoltaic output to generate the joint probability distribution model P of wind power output and photovoltaic output. X,Y (x,y):
[0047] P X,Y (x,y)=C θ (F X (x),F Y (y))
[0048] Among them, C θ is the optimal connection function Copula, x represents the value of wind power output in the [0,1] per unit value interval, F X (x) represents the probability value of wind power output ≤ x, y represents the value of photovoltaic output in the per-unit value interval [0,1], F Y (y) represents the probability value of photovoltaic output ≤ y.
[0049] Preferably, the specific process of generating the random scenario set of new energy joint output is:
[0050] The Markov chain is used to define the state space discretization rule, and the output interval S of the joint probability distribution model of wind power output and photovoltaic output is divided into n1 states according to the preset interval:
[0051]
[0052] Among them, s i Indicates the i-th output state interval, a i Indicates the lower limit of the output state range, b i Shows the upper limit of the force state range;
[0053] Construct the state transition probability matrix of wind power output and photovoltaic output:
[0054]
[0055] Among them, N i→j Represents the i-th output state interval s in the historical output datai To the jth output state interval s i Number of jumps, N i→k Represents the range from the i-th output state interval to the k-th output state interval s in the historical output data k Number of jumps, P ij Representing the jump probability, the state transition probability matrix of wind power output and photovoltaic output is used to count the state jump frequency of adjacent time periods in the historical output data, and combined with the joint probability distribution model to form a joint probability distribution model with time correlation;
[0056] The state transition probability matrix is used as the target distribution Π(i), which represents the ideal state probability distribution; the Gaussian distribution is used as the proposed distribution Q(j|i), which represents the proposed jump direction probability;
[0057] Extract candidate output state interval s from the proposal distribution Q(j|i) * , the acceptance probability α of state transition is calculated by Markov chain Monte Carlo algorithm, and the calculation formula is as follows:
[0058]
[0059] Among them, Π(s * ) represents the candidate output state interval s * The probability of being in the target distribution Π(i), Π(s t ) represents the current output state interval s t The probability of being in the target distribution Π(i), Q(s t |s * ) represents the candidate output state interval s * To the current output state interval s t The proposed probability density, Q(s * |s t ) indicates the interval from the current output state s t To the candidate output state interval s * The proposed probability density of ;
[0060] The output state interval is updated according to the Markov Chain Monte Carlo criterion:
[0061]
[0062] Among them, u represents the random variable that controls the probabilistic state transition, α represents the acceptance probability of the state transition, and s t+1 The updated output state interval, t+1 represents the next moment after time t;
[0063] Get the updated output state interval s t+1, repeat the updating process of the output state interval and generate scenarios, and finally generate a set of random scenarios of new energy joint output.
[0064] Preferably, a clustering algorithm is used to cluster the random scenarios of the new energy joint output to obtain representative scenarios, and the specific process of evaluating the matching degree between the representative scenarios and the pre-processed historical output data of wind power output and photovoltaic output by KL divergence is as follows:
[0065] Construct the average wind power output μ w , average photovoltaic output μ p , wind power output standard deviation σ w , photovoltaic output standard deviation σ p , wind power output skew w , photovoltaic output skew p , maximum wind power ramp rate Photovoltaic maximum ramp rate and the proportion of zero output period f zero The multidimensional feature vector of each scene in the random scene set is obtained i as follows:
[0066]
[0067] in, Represents a 9-dimensional feature vector;
[0068] In the eigenvector, the average wind power output μ w The calculation formula is as follows:
[0069]
[0070] Among them, X t is the per-unit value of wind power output at time t, where T represents T moments;
[0071] Average photovoltaic output μ p The calculation formula is as follows:
[0072]
[0073] Among them, Y t is the per-unit photovoltaic output at time t;
[0074] Wind power output standard deviation σ w The calculation formula is as follows:
[0075]
[0076] Among them, μ w The average output of wind power;
[0077] Photovoltaic output standard deviation σp The calculation formula is as follows:
[0078]
[0079] Among them, μ p is the average photovoltaic output;
[0080] Wind power output skew w The calculation formula is as follows:
[0081]
[0082] Among them, σ w is the standard deviation of wind power output;
[0083] Photovoltaic output skew p The calculation formula is as follows:
[0084]
[0085] Among them, σ p is the standard deviation of photovoltaic output;
[0086] Maximum wind power ramp rate The calculation formula is as follows:
[0087]
[0088] Among them, X t+1 is the per-unit wind power output at time t+1, and Δt is the time interval between time t and time t+1;
[0089] Photovoltaic maximum ramp rate The calculation formula is as follows:
[0090]
[0091] Among them, Y t+1 The per-unit value of photovoltaic output at time t+1;
[0092] The proportion of zero output period f zero The calculation formula is as follows:
[0093]
[0094] Among them, Count(X t <0.05 or Y t <0.05) indicates that the statistical X t <0.05 or Y t Number of moments <0.05;
[0095] The k-means++ algorithm is used to cluster the random scene set of new energy joint output to obtain representative scenes, and a scene is randomly selected from the random scene set of new energy joint output as the initial cluster center c1, and subsequent scenes are selected as cluster centers c i The probability (f i ) is calculated as follows:
[0096]
[0097] Among them, D(f i ) is the distance from scene i to the nearest cluster center, Entropy(f i ) is the statistical entropy of scene i, λ is the entropy weight coefficient;
[0098] For each scene i, assign it to the corresponding cluster k:
[0099]
[0100] Among them, the weighted distance The calculation formula is as follows:
[0101]
[0102] Among them, f i is the feature vector of the i-th scene, c k is the center vector of the kth cluster, W is the diagonal weight matrix, (f i -c k ) represents the difference vector between the feature vector of the i-th scene and the center vector of the k-th cluster in the feature space, (f i -c k ) T represents the transpose of the difference vector between the feature vector of the i-th scene and the center vector of the k-th cluster in the feature space;
[0103] Climbing rate penalty The calculation formula is as follows:
[0104]
[0105] Among them, R i is the climbing rate vector of the current scene i, is the historical average climbing rate of cluster k, is the wind power ramp rate of the current scenario i, is the historical wind power ramp rate of cluster k, is the photovoltaic ramp rate of the current scenario i, is the historical PV ramp rate of cluster k;
[0106] Perform clustering iteration on cluster k to obtain k representative scenes;
[0107] The KL divergence is used to evaluate the matching degree between the representative scenarios and the pre-processed historical output data of wind power output and photovoltaic output. KL (P hist ||P gen ) is calculated as follows:
[0108]
[0109] Among them, P hist is the true probability distribution of the historical output data of wind power output and photovoltaic output after preprocessing, P gen is the probability distribution of the generated representative scenes, P hist (i) is the true probability distribution of the historical output data of wind power output and photovoltaic output after preprocessing in the i-th discretization interval, P gen (i) is the probability distribution of the representative scene generated by the i-th discretization interval, i is the discretization interval number, N is the total number of intervals, and || is used to convert P hist With P gen Separated, indicating the comparison object of KL divergence;
[0110] According to D KL (P hist ||P gen ) meets the preset threshold range to evaluate the matching degree between the representative scenario and the pre-processed historical output data of wind power output and photovoltaic output.
[0111] A data-driven method for generating a new energy joint output scenario includes the following steps:
[0112] Preprocess the historical output data of wind power and photovoltaic power;
[0113] The wind power output marginal distribution and the photovoltaic output marginal distribution are obtained by fitting the pre-processed historical output data of wind power output and photovoltaic output respectively through the probability density function, and the correlation between wind power output and photovoltaic output is analyzed on the pre-processed historical output data of wind power output and photovoltaic output through the Kendall rank correlation coefficient and the tail dependence coefficient, and the Kendall rank correlation coefficient matching error and the tail dependence coefficient matching error of each set connection function and the pre-processed historical output data of wind power output and photovoltaic output are obtained. The Euclidean distance between each set connection function and the pre-processed historical output data of wind power output and photovoltaic output is calculated, and the optimal connection function is selected in combination with the Kendall rank correlation coefficient matching error and the tail dependence coefficient matching error. Based on the optimal connection function and in combination with the wind power output marginal distribution and the photovoltaic output marginal distribution, a joint probability distribution model of wind power output and photovoltaic output is established;
[0114] A Markov chain is used in combination with the joint probability distribution model to form a joint probability distribution model with time correlation, and a random scenario set of new energy joint output is generated by sampling from the joint probability distribution model with time correlation through the Markov chain Monte Carlo algorithm.
[0115] A computer program product comprises a computer program, wherein when the computer program is executed by a processor, the steps of the above method are implemented.
[0116] Beneficial effects of the present invention:
[0117] The present invention can more accurately predict the fluctuation characteristics of wind power output and photovoltaic output, provide a more reliable decision-making basis for the scheduling and operation of the power system, thereby effectively reducing the power system operation risk caused by the intermittency and uncertainty of new energy, and enhancing the stability of the power system; the present invention helps to optimize the operation strategy of the power system through precise output scenario generation, reduce the demand for backup capacity, improve equipment utilization, and reduce the additional costs caused by frequent adjustments and backup power startup, thereby realizing the economic operation of the power system; the present invention takes into account the time correlation of wind power output and photovoltaic output, generates a random scenario set with time consistency, more realistically reflects the actual characteristics of new energy output, and provides more realistic simulation results for the planning and operation of the power system. BRIEF DESCRIPTION OF THE DRAWINGS
[0118] Figure 1 It is a structural schematic diagram of the present invention;
[0119] Figure 2 Flowchart of the present invention. DETAILED DESCRIPTION
[0120] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:
[0121] Example 1
[0122] A data-driven new energy joint output scenario generation system, such as Figure 1 As shown, it includes:
[0123] The data acquisition module is used to collect historical wind power and photovoltaic output data (including wind speed, irradiance, and temperature) through the data acquisition and monitoring control system (SCADA system) and preprocess the historical wind power and photovoltaic output data (normalization preprocessing). This design preprocesses the historical wind power and photovoltaic output data to improve the quality and availability of the data, providing a rich data foundation for subsequent modeling and analysis.
[0124] The joint probability distribution model construction module is used to fit the historical output data of pre-processed wind power output and photovoltaic output respectively through the probability density function to obtain the marginal distribution of wind power output and the marginal distribution of photovoltaic output, and to analyze the correlation between wind power output and photovoltaic output through the Kendall rank correlation coefficient and the tail dependence coefficient of the historical output data of pre-processed wind power output and photovoltaic output, and obtain the Kendall rank correlation coefficient matching error and the tail dependence coefficient matching error of each set connection function (the connection function is the Copula function, and the set connection functions include Gaussian, Frank and Gumbel) and the historical output data of pre-processed wind power output and photovoltaic output. The Euclidean distance between the data is calculated, and the optimal connection function is selected from the set connection functions in combination with the Kendall rank correlation coefficient matching error and the tail dependence coefficient matching error. Based on the optimal connection function and combined with the wind power output edge distribution and the photovoltaic output edge distribution, a joint probability distribution model of wind power output and photovoltaic output is established. This design can describe the probability characteristics of wind power output and photovoltaic output separately by fitting the wind power output edge distribution and the photovoltaic output edge distribution. The correlation analysis is performed through the Kendall rank correlation coefficient and the tail dependence coefficient, which can comprehensively and accurately characterize the correlation between wind power output and photovoltaic output, including the correlation in extreme cases. A joint probability distribution model of wind power output and photovoltaic output is established, which comprehensively considers the edge distribution and correlation of wind power output and photovoltaic output, and is closer to the actual situation.
[0125] The scenario generation module is used to use the Markov chain in combination with the joint probability distribution model to form a joint probability distribution model with time correlation, and to sample and generate a random scenario set of new energy joint output from the joint probability distribution model with time correlation through the Markov chain Monte Carlo algorithm (the random scenario set is a collection of a large number of possible output trajectories generated by the joint probability distribution model, and each scenario represents the joint output data of wind power and photovoltaic energy under a complete time series). By generating a joint probability distribution model with time correlation, this design can be closer to the actual wind power output and photovoltaic output change process. By using the Markov chain Monte Carlo algorithm to sample and generate random scenario sets, a large number of random scenarios that conform to the characteristics of wind power output and photovoltaic output can be generated, providing rich scenario support for the scheduling, planning and reliability analysis of the power system.
[0126] In the above technical solution, the scenario reduction and evaluation module is used to cluster the random scenario set of new energy joint output using a clustering algorithm to obtain representative scenarios, and evaluate the degree of matching between the representative scenarios and the pre-processed historical output data of wind power output and photovoltaic output through KL divergence (i.e., information divergence, full name Kullback-Leibler divergence). The above design clusters the random scenario set through a clustering algorithm, which can reduce the random scenario set to representative scenarios. The representative scenarios can effectively reflect the typical characteristics of wind power output and photovoltaic output, and significantly reduce the complexity of subsequent calculations; the KL divergence is used to evaluate the degree of matching between the representative scenarios and the pre-processed historical output data of wind power output and photovoltaic output, which can ensure that the representative scenarios are highly consistent with the pre-processed historical output data of wind power output and photovoltaic output in statistical characteristics, and ensure that the generated representative scenarios can be used for power system planning, scheduling and reliability analysis.
[0127] In the above technical solution, the specific method for preprocessing the historical output data of wind power output and photovoltaic output includes:
[0128] The historical output data of wind power and photovoltaic power are normalized and preprocessed. Minimum and maximum normalization is used to linearly map the historical output data of wind power and photovoltaic power to the [0,1] per-unit value interval (per-unit value is a dimensionless relative value used to represent the size of a physical quantity relative to a reference value. For example, in wind power output, the per-unit value represents the ratio of the actual wind power output to the installed capacity of the wind farm). The preprocessed historical output data of wind power and photovoltaic power are obtained using the following formula:
[0129]
[0130] Among them, m is the historical output data of wind power output and photovoltaic output, m min is the minimum value of the historical output data of wind power output and photovoltaic output, m max is the maximum value of the historical output data of wind power output and photovoltaic output, and m′ is the historical output data of wind power output and photovoltaic output after preprocessing; the above design performs normalization preprocessing on the historical output data of wind power output and photovoltaic output, which can unify the data range, eliminate the dimensionality effect, improve the calculation stability, and provide a high-quality data foundation for the subsequent joint probability distribution model establishment and correlation analysis.
[0131] In the above technical solution, the specific process of fitting the pre-processed historical output data of wind power output and photovoltaic output respectively by probability density function to obtain the marginal distribution of wind power output and photovoltaic output is as follows:
[0132] The Weibull distribution probability density function of the marginal distribution of wind power output is obtained as follows:
[0133]
[0134] Where v is the wind speed; λ is the scale parameter of the Weibull distribution; k is the shape parameter of the Weibull distribution; is the scaling factor; Reflects the power relationship of wind speed v relative to scale parameter λ; is the index term;
[0135] The beta distribution probability density function of the photovoltaic output marginal distribution is obtained as follows:
[0136]
[0137] Where P is the photovoltaic output-related variable; α is the first shape parameter of the Beta distribution; β is the second shape parameter of the Beta distribution; P α-1 Indicates the impact of the value of α on the probability density; (1-P) β-1 represents the probability density characteristics of photovoltaic output; B(α, β) is the beta function; the above design fits the marginal distribution of wind power output and the marginal distribution of photovoltaic output through the Weibull distribution probability density function and the beta distribution probability density function, respectively. It can accurately describe the output characteristics of wind power output and photovoltaic output, improve the adaptability and accuracy of the joint probability distribution model, and simplify the calculation and analysis process, providing a reliable foundation for the subsequent generation of joint output scenarios.
[0138] In the above technical solution, the correlation between wind power output and photovoltaic output is analyzed by using the Kendall rank correlation coefficient and the tail dependence coefficient on the preprocessed historical output data of wind power output and photovoltaic output, and the specific process of obtaining the matching error between each set connection function and the Kendall rank correlation coefficient and the tail dependence coefficient of the preprocessed historical output data of wind power output and photovoltaic output is as follows:
[0139] The Kendall rank correlation coefficient τ is calculated by statistically analyzing the proportion of the coordinated changes in the historical output data of wind power output X and photovoltaic output Y. The calculation formula is as follows:
[0140]
[0141] Among them, N c is the number of same-direction logarithms at all moments in the historical output data. i >x j ) and (y i >y j ), or (x i <x j ) and (y i <y j ),(x i ,y i) and (x j ,y j ) is the same direction logarithm; N d is the number of reverse logarithms of all moments in the historical output data. i >x j ) and (y i <y j ), or (x i <x j ) and (y i >y j ),(x i ,y i ) and (x j ,y j ) is the reverse logarithm; n represents the total number of moments in the historical output data; x i represents the wind power output at time i, x j represents the wind power output at time j, y i represents the photovoltaic output at time i, y j represents the photovoltaic output at time j, and τ represents the Kendall rank correlation coefficient, which is used to measure the intensity of the trend of simultaneous increase or decrease of wind power output X and photovoltaic output Y (with a value between -1 and 1).
[0142] Calculate the tail dependence coefficient λ of the preprocessed historical output data of wind power output and photovoltaic output. The tail dependence coefficient λ includes the upper tail dependence coefficient λ u and the lower tail dependence coefficient λ l , upper tail dependence coefficient λ u It represents the probability of wind power output and photovoltaic output reaching maximum value at the same time, and the lower tail dependence coefficient λ l Indicates the probability that wind power output and photovoltaic output reach their minimum values at the same time;
[0143] The upper tail dependence coefficient calculation formula is:
[0144]
[0145] Among them, λ u represents the upper tail dependence coefficient, q represents the quantile threshold, represents the inverse cumulative distribution function of the wind power output X variable, represents the inverse cumulative distribution function of the photovoltaic output Y variable, Wind power output And photovoltaic output probability;
[0146] The symbolic explanation and specific examples of the upper tail dependence coefficient calculation formula are shown in Table 1:
[0147] Table 1 Example of upper tail dependence coefficient
[0148]
[0149] Here is an example:
[0150] Assume that the historical output data (per unit) of wind power and photovoltaic power output in a certain place are as shown in Table 2:
[0151] Table 2 Historical output data of wind power and photovoltaic power output in a certain place
[0152] time Wind power (X) Photovoltaic (Y) 12 noon 0.92 0.88 2 p.m. 0.95 0.82 ... ... ...
[0153] First, set the threshold, take q = 0.9 (i.e., the extreme value of the top 10% of the study), and calculate the quantile:
[0154] (90% quantile of wind power output)
[0155] (90% percentile of photovoltaic output)
[0156] For extreme events, the total sample size is 1000 moments, 100 of which meet X>0.9 (top 10%), and 35 of which meet Y>0.87.
[0157] Calculate the upper tail dependence coefficient λ u :
[0158]
[0159] where λ u =0.35 means that when wind power output is very high (in the top 10%), there is a 35% probability that photovoltaic output will also be in a high output state simultaneously;
[0160] The formula for calculating the lower tail dependence coefficient is:
[0161]
[0162] Among them, λ l represents the lower tail dependence coefficient, Wind power output And photovoltaic output probability;
[0163] Similarly, if λ l =0.4 means that when wind power output is extremely low, there is a 40% probability that photovoltaic output will also be extremely low;
[0164] Upper tail dependence coefficient λ u and the lower tail dependence coefficient λ l It can be used to assess the extreme risk of "all wind and solar power outage", as shown in Table 3:
[0165] Table 3 Extreme risk assessment
[0166] Coefficient value range Wind and solar power output relationship Power system response measures <![CDATA[λ u >0.5]]> Very easy to achieve high output at the same time It is necessary to add frequency modulation units to prevent overvoltage <![CDATA[λ l >0.3]]> At the same time, the risk of low output is significant Configure more energy storage or backup power <![CDATA[λ u ≈0]]> There is no correlation between extreme wind and solar output values Reduce redundant capacity configuration
[0167] Based on the correlation analysis between wind power output and photovoltaic output, the matching errors of each set connection function are obtained, including the Kendall rank correlation coefficient matching error, the upper tail dependence coefficient matching error, and the lower tail dependence coefficient matching error (where the Kendall rank correlation coefficient matching error, the upper tail dependence coefficient matching error, and the lower tail dependence coefficient matching error are all less than the preset value of 5%, the connection function meets the requirements);
[0168] The Kendall rank correlation coefficient matching error calculation formula is:
[0169]
[0170] where τ 理论 is the theoretical Kendall rank correlation coefficient of each link function obtained by setting the parameters of each link function; τ 样本 is the Kendall rank correlation coefficient τ of the preprocessed historical output data of wind power output and photovoltaic output;
[0171] The upper tail dependency coefficient matching error calculation formula is:
[0172]
[0173] where λ u,理论 is the theoretical tail dependence coefficient of each link function obtained by setting the parameters of each link function; τ 样本 is the upper tail dependence coefficient λ of the preprocessed historical output data of wind power output and photovoltaic output u ;
[0174] The formula for calculating the matching error of the lower tail dependence coefficient is:
[0175]
[0176] where λ I,理论 is the theoretical tail dependence coefficient of each link function obtained by setting the parameters of each link function; I,样本 is the lower tail dependence coefficient λ of the preprocessed historical output data of wind power output and photovoltaic output l ;
[0177] The above design can effectively measure the nonlinear correlation between wind power output and photovoltaic output through the Kendall rank correlation coefficient, making the analysis results closer to the actual physical process; the tail dependence coefficient is used to further analyze the correlation between wind power output and photovoltaic output under extreme values (maximum or minimum); by calculating the matching error of each connection function, the goodness of fit of different connection functions to the joint probability distribution model can be quantified, thereby improving the accuracy and reliability of the joint probability distribution model.
[0178] In the above technical solution, the Euclidean distance between each set connection function and the preprocessed historical output data of wind power output and photovoltaic output is calculated, and the optimal connection function is selected based on the matching error. The specific process of establishing the joint probability distribution model of wind power output and photovoltaic output based on the optimal connection function and the marginal distribution of wind power output and photovoltaic output is as follows:
[0179] Calculate the Euclidean distance between each copula function and the preprocessed historical output data of wind power and photovoltaic output. The calculation formula of the Euclidean distance matching method is as follows:
[0180]
[0181] Where n represents the total number of moments in the historical output data, u i represents the historical output data of wind power after preprocessing, v i Represents the historical output data of photovoltaic power after preprocessing, C e represents the real dependency value directly calculated based on the pre-processed historical output data of wind power output and photovoltaic output, C t Indicates the theoretical dependence value of each set connection function Copula;
[0182] The calculation example of the Euclidean distance matching method is as follows:
[0183] Assume there are three sets of data, as shown in Table 4:
[0184] Table 4. Data related to the Euclidean distance matching method
[0185] Point Number <![CDATA[u i ]]> <![CDATA[v i ]]> <![CDATA[C e ]]> <![CDATA[Gaussian C t ]]> <![CDATA[Frank C t ]]> 1 0.3 0.5 0.10 0.12 0.09 2 0.6 0.8 0.25 0.20 0.26 3 0.9 0.2 0.05 0.08 0.04
[0186] Gaussian-Copula score:
[0187]
[0188] Frank-Copula score:
[0189]
[0190] Since d1>d2, and combined with the matching errors of the various connection functions, Frank-Copula is selected as the optimal connection function Copula;
[0191] Select the connection function Copula with the smallest Euclidean distance from the preprocessed historical output data of wind power output and photovoltaic output, obtain the Kendall rank correlation coefficient matching error, upper tail dependence coefficient matching error and lower tail dependence coefficient matching error of the connection function Copula with the smallest Euclidean distance from the matching error calculated above, and determine whether the Kendall rank correlation coefficient matching error, upper tail dependence coefficient matching error and lower tail dependence coefficient matching error of the connection function Copula with the smallest Euclidean distance are all less than preset values. If so, the connection function is used as the optimal connection function Copula. Otherwise, determine whether the matching errors of the connection function Copula with the second smallest Euclidean distance are all less than the preset value, and so on until the optimal connection function Copula is determined;
[0192] The optimal connection function Copula is combined with the marginal distribution of wind power output and photovoltaic output to generate the joint probability distribution model P of wind power output and photovoltaic output. X,Y (x,y):
[0193] P X,Y (x,y)=C θ (F X (x),F Y (y))
[0194] Among them, C θ is the optimal connection function Copula, x represents the value of wind power output in the [0,1] per unit value interval, F X (x) represents the probability value of wind power output ≤ x, y represents the value of photovoltaic output in the per-unit value interval [0,1], F Y (y) represents the probability value of PV output ≤ y;
[0195] An example of combining the optimal link function Copula (Frank-Copula) is as follows:
[0196] P X,Y (wind power < 0.4, photovoltaic < 0.6) = C Frank (F X (0.4),F Y (0.6)
[0197] If the calculated result is 0.15, it means there is a 15% probability that the wind power output will be lower than 40% and the photovoltaic output will be lower than 60%;
[0198] The above design can quantify the goodness of fit between different Copula functions and actual data (historical output data of preprocessed wind power and photovoltaic output) by calculating the Euclidean distance. The smaller the Euclidean distance, the closer the theoretical value of the Copula function is to the dependency structure of the actual data. Further optimization and selection based on the matching error can more comprehensively evaluate the applicability of the Copula function and avoid the deviation caused by a single indicator. Selecting the optimal connection function Copula combined with the marginal distribution of wind power and photovoltaic can more accurately describe the complex correlation between wind power output and photovoltaic output, including nonlinear relationships and tail dependence, and can ensure that the optimal connection function Copula maintains a good fitting effect under different conditions, thereby improving the adaptability and reliability of the joint probability distribution.
[0199] In the above technical solution, the specific process of generating the random scenario set of new energy joint output is as follows:
[0200] The Markov chain is used to define the state space discretization rule, and the output interval S of the joint probability distribution model of wind power output and photovoltaic output is divided into n1 states according to the preset interval:
[0201]
[0202] Among them, s i represents the i-th output state interval (such as [0.4,0.5]), a i Indicates the lower limit of the output state range, b i Shows the upper limit of the force state range;
[0203] Construct the state transition probability matrix of wind power output and photovoltaic output:
[0204]
[0205] Among them, N i→j Represents the i-th output state interval s in the historical output data i To the jth output state interval s i Number of jumps, N i→k Represents the range from the i-th output state interval to the k-th output state interval s in the historical output data k Number of jumps, P ij Representing the jump probability, the state transition probability matrix of wind power output and photovoltaic output is used to count the state jump frequency of adjacent time periods in the historical output data, and combined with the joint probability distribution model to form a joint probability distribution model with time correlation;
[0206] The state transition probability matrix is used as the target distribution Π(i), which represents the ideal state probability distribution (the target distribution Π(i) is derived from the state transition probability matrix constructed by preprocessing the historical output data of wind power output and photovoltaic output); the Gaussian distribution is used as the proposed distribution Q(j|i), which represents the probability of the recommended jump direction;
[0207] Extract candidate output state interval s from the proposal distribution Q(j|i) * (As currently i =[0.3,0.4], proposal s * =[0.35,0.45]), the acceptance probability α of the state transition is calculated by the Markov Chain Monte Carlo algorithm, and the calculation formula is as follows:
[0208]
[0209] Among them, Π(s * ) represents the candidate output state interval s * The probability of being in the target distribution Π(i), Π(s t ) represents the current output state interval s t The probability of being in the target distribution Π(i), Q(s t |s * ) represents the candidate output state interval s * To the current output state interval s t The proposed probability density, Q(s * |s t ) indicates the interval from the current output state s t To the candidate output state interval s * The proposed probability density of ;
[0210] The output state interval is updated according to the Markov Chain Monte Carlo criterion:
[0211]
[0212] Among them, u represents the random variable that controls the probabilistic state transition, α represents the acceptance probability of the state transition, and s t+1 The updated output state interval, t+1 represents the next moment after time t;
[0213] Input: current state s t , candidate state s * , acceptance probability α;
[0214] Assume that the current wind power output is 60%, and the system recommends jumping to 65%. According to historical data, 60%→65% has occurred 200 times, and 60% has remained unchanged for 800 times. The acceptance probability α is calculated to be 200 / 800=25%. A random number between 0 and 1 (for example, 0.4) is generated. Since 0.4>0.25 (acceptance probability α), the state s at the next moment is t+1 The wind power output remains unchanged at 60%;
[0215] Output: next moment state s t+1 ;
[0216] Get the updated output state interval s t+1 , repeat the updating process of the output state interval and generate scenarios, and finally generate a set of random scenarios of new energy joint output (five thousand scenarios can be generated in the random scenario set); the above design can convert the pre-processed historical output data of wind power output and photovoltaic output into a Markov chain model with time series characteristics by defining the state space discretization rules and constructing the state transition probability matrix. Combined with the joint probability distribution model, it can generate output scenarios with time correlation, which are closer to the actual wind power output and photovoltaic output change process; the Markov chain Monte Carlo algorithm can ensure that the generated scenarios are highly consistent with the historical output data of wind power output and photovoltaic output in terms of statistical characteristics through the calculation of acceptance probability and the introduction of random variables, and can generate representative random scenarios, providing rich scenario support for the scheduling, planning and reliability analysis of the power system.
[0217] In the above technical solution, a clustering algorithm is used to cluster the random scenarios of renewable energy combined output to obtain representative scenarios. The KL divergence is used to evaluate the matching degree between the representative scenarios and the pre-processed historical output data of wind power output and photovoltaic output. The specific process is as follows:
[0218] Construct the average wind power output μ w , average photovoltaic output μ p , wind power output standard deviation σ w , photovoltaic output standard deviation, wind power output skew w , photovoltaic output skew p , maximum wind power ramp rate Photovoltaic maximum ramp rate and the proportion of zero output period f zero The multidimensional feature vector of each scene in the random scene set is obtained i as follows:
[0219]
[0220] in, Represents a 9-dimensional feature vector;
[0221] In the eigenvector, the average wind power output μ w The calculation formula is as follows:
[0222]
[0223] Among them, X t is the per-unit value of wind power output at time t, where T represents T moments;
[0224] Average photovoltaic output μ p The calculation formula is as follows:
[0225]
[0226] Among them, Y t is the per-unit photovoltaic output at time t;
[0227] Wind power output standard deviation σ w The calculation formula is as follows:
[0228]
[0229] Among them, μ w The average output of wind power;
[0230] Photovoltaic output standard deviation σ p The calculation formula is as follows:
[0231]
[0232] Among them, μ p is the average photovoltaic output;
[0233] Wind power output skew w The calculation formula is as follows:
[0234]
[0235] Among them, σ w is the standard deviation of wind power output;
[0236] Photovoltaic output skew p The calculation formula is as follows:
[0237]
[0238] Among them, σ p is the standard deviation of photovoltaic output;
[0239] Maximum wind power ramp rate The calculation formula is as follows:
[0240]
[0241] Among them, X t+1 is the per-unit wind power output at time t+1, and Δt is the time interval between time t and time t+1;
[0242] Photovoltaic maximum ramp rate The calculation formula is as follows:
[0243]
[0244] Among them, Y t+1 The per-unit value of photovoltaic output at time t+1;
[0245] The proportion of zero output period f zero The calculation formula is as follows:
[0246]
[0247] Among them, Count(X t <0.05 or Y t <0.05) indicates that the statistical X t <0.05 or Y t Number of moments <0.05;
[0248] An example is as follows: 24-hour data for a certain scenario is shown in Table 5:
[0249] Table 5 24-hour data of a scene
[0250] time <![CDATA[Wind power X t > <![CDATA[Photovoltaic Y t > 00:00 0.62 0.01 ... ... ... 23:45 0.58 0.02
[0251] The calculation results are:
[0252]
[0253] f i In this scenario, the photovoltaic power output accounts for 25% and the wind power skewness is 0.3, which is right-skewed.
[0254] Input feature matrix:
[0255]
[0256] Among them, M represents 5000 random scenes;
[0257] Normalize the feature matrix to get the weight matrix:
[0258] W=diag(0.3,0.3,0.2,0.2,0.1,0.1,0.5,0.5,1.0)
[0259] Among them, the zero output frequency (9th dimension) is given the highest weight of 1.0;
[0260] The k-means++ algorithm is used to cluster the random scene set of new energy joint output to obtain representative scenes, and a scene is randomly selected from the random scene set of new energy joint output as the initial cluster center c1, and subsequent scenes are selected as cluster centers c i The probability (f i ) is calculated as follows:
[0261]
[0262] Among them, D(f i ) is the distance from scene i to the nearest cluster center, Entropy(f i ) is the statistical entropy of scene i, λ is the entropy weight coefficient;
[0263] For each scene i, assign it to the corresponding cluster k:
[0264]
[0265] Among them, the weighted distance The calculation formula is as follows:
[0266]
[0267] Among them, f i is the feature vector of the i-th scene, c k is the center vector of the kth cluster, W is the diagonal weight matrix, (f i -c k ) represents the difference vector between the feature vector of the i-th scene and the center vector of the k-th cluster in the feature space, (f i -c k ) T represents the transpose of the difference vector between the feature vector of the i-th scene and the center vector of the k-th cluster in the feature space;
[0268] Climbing rate penalty The calculation formula is as follows:
[0269]
[0270] Among them, R i is the climbing rate vector of the current scene i, is the historical average climbing rate of cluster k, is the wind power ramp rate of the current scenario i, is the historical wind power ramp rate of cluster k, is the photovoltaic ramp rate of the current scenario i, is the historical PV ramp rate of cluster k;
[0271] Perform clustering iteration on cluster k to obtain k representative scenes;
[0272] The matching degree between the representative scenario and the pre-processed historical output data of wind power output and photovoltaic output is evaluated by KL divergence (the smaller the KL divergence, the higher the matching degree between the representative scenario and the pre-processed historical output data of wind power output and photovoltaic output). KL divergence D KL (P hist ||P gen ) is calculated as follows:
[0273]
[0274] Among them, P hist is the true probability distribution of the historical output data of wind power output and photovoltaic output after preprocessing, P gen is the probability distribution of the generated representative scenes, P hist (i) is the true probability distribution of the historical output data of wind power output and photovoltaic output after preprocessing in the i-th discretization interval, P gen (i) is the probability distribution of the representative scene generated by the i-th discretization interval, i is the discretization interval number, N is the total number of intervals, and || is used to convert P hist With P gen Separated, indicating the comparison object of KL divergence;
[0275] The calculation example is as follows:
[0276] The wind power output range [0, 1.05] is evenly divided into 20 segments (each segment is 0.0525), and the photovoltaic output range [0, 1.05] is evenly divided into 20 segments (each segment is 0.0525). The wind power output range and the photovoltaic output range form 400 joint output intervals (one joint output interval is one grid) (for example, wind power segment 3 + photovoltaic segment 5: X∈[0.105, 0.1575) and Y∈[0.21, 0.2625);
[0277] For the pre-processed historical output data of wind power and photovoltaic power, the number of times each combined output interval grid appears is counted; for the data generated by the representative scenario, the number of times each combined output interval grid appears is also counted;
[0278] Compute the probability distribution:
[0279]
[0280] Calculate the KL component grid by grid and calculate for each grid:
[0281]
[0282] The final KL value is summarized as follows:
[0283]
[0284] According to D KL (P hist ||P gen ) meets the preset threshold range to evaluate the matching degree between the representative scenario and the pre-processed historical output data of wind power output and photovoltaic output, as shown in Table 6:
[0285] Table 6 KL value evaluation
[0286] KL value range Degree of Match Countermeasures <0.05 excellent Ready to use 0.05-0.1 qualified Check the tail interval >0.1 Unqualified Re-adjust the model
[0287] The final KL value is calculated, and the degree of match between the representative scenario and the pre-processed historical output data of wind power and photovoltaic output is evaluated based on whether the KL value meets the preset threshold range. The above design can comprehensively reflect the statistical characteristics of wind power output and photovoltaic output by constructing a multi-dimensional feature vector. The K-means++ clustering algorithm can reduce the number of iterations and improve clustering efficiency. By calculating the climbing rate penalty, the clustering results are further optimized to ensure that the representative scenario also matches the pre-processed historical output data of wind power output and photovoltaic output in terms of dynamic characteristics. Combining the clustering algorithm and KL divergence evaluation, it is possible to generate scenarios that are both representative and highly matched with the pre-processed historical output data of wind power output and photovoltaic output, providing reliable data support for the optimized operation of the power system.
[0288] Example 2
[0289] A data-driven method for generating new energy joint output scenarios, such as Figure 2 As shown in the figure, historical output data of wind power and photovoltaic power are collected and preprocessed; the marginal distribution of wind power and photovoltaic power is obtained by fitting the probability density function, the correlation analysis of wind power and photovoltaic power is performed by the Kendall rank correlation coefficient and the tail dependence coefficient, the matching error of each connection function is obtained, the Euclidean distance between each connection function and the preprocessed historical output data of wind power and photovoltaic power is calculated, and the optimal connection function is obtained by combining the matching error, thereby establishing a joint probability distribution model of wind power and photovoltaic power; the Markov chain is used to form a joint probability distribution model with time correlation, and the random scene set is generated by sampling through the Markov chain Monte Carlo algorithm.
[0290] The specific method for generating a new energy joint output scenario includes the following steps:
[0291] Preprocess the historical output data of wind power and photovoltaic power;
[0292] The wind power output marginal distribution and the photovoltaic output marginal distribution are obtained by fitting the pre-processed historical output data of wind power output and photovoltaic output respectively through the probability density function, and the correlation between wind power output and photovoltaic output is analyzed on the pre-processed historical output data of wind power output and photovoltaic output through the Kendall rank correlation coefficient and the tail dependence coefficient, and the Kendall rank correlation coefficient matching error and the tail dependence coefficient matching error of each set connection function and the pre-processed historical output data of wind power output and photovoltaic output are obtained. The Euclidean distance between each set connection function and the pre-processed historical output data of wind power output and photovoltaic output is calculated, and the optimal connection function is selected in combination with the Kendall rank correlation coefficient matching error and the tail dependence coefficient matching error. Based on the optimal connection function and in combination with the wind power output marginal distribution and the photovoltaic output marginal distribution, a joint probability distribution model of wind power output and photovoltaic output is established;
[0293] A Markov chain is used in combination with the joint probability distribution model to form a joint probability distribution model with time correlation, and a random scenario set of new energy joint output is generated by sampling from the joint probability distribution model with time correlation through the Markov chain Monte Carlo algorithm.
[0294] Example 3
[0295] A computer program product includes a computer program, characterized in that when the computer program is executed by a processor, the steps of the method described in Example 2 are implemented.
[0296] The contents not described in detail in this specification belong to the prior art known to professional and technical personnel in this field.
Claims
1. A data-driven new energy joint output scenario generation system, characterized by: include: The data preprocessing module is used to preprocess the historical output data of wind power output and photovoltaic output; The joint probability distribution model construction module is used to fit the pre-processed historical output data of wind power output and photovoltaic output respectively through the probability density function to obtain the wind power output marginal distribution and the photovoltaic output marginal distribution, and to perform the correlation analysis between the wind power output and the photovoltaic output on the pre-processed historical output data of wind power output and photovoltaic output through the Kendall rank correlation coefficient and the tail dependence coefficient, and obtain the Kendall rank correlation coefficient matching error and the tail dependence coefficient matching error of each set connection function and the pre-processed historical output data of wind power output and photovoltaic output. The Euclidean distance between each set connection function and the pre-processed historical output data of wind power output and photovoltaic output is calculated, and the optimal connection function is selected in combination with the Kendall rank correlation coefficient matching error and the tail dependence coefficient matching error. Based on the optimal connection function and in combination with the wind power output marginal distribution and the photovoltaic output marginal distribution, a joint probability distribution model of wind power output and photovoltaic output is established; The scenario generation module is used to use the Markov chain in combination with the joint probability distribution model to form a joint probability distribution model with time correlation, and to generate a random scenario set of new energy joint output from the joint probability distribution model with time correlation by sampling through the Markov chain Monte Carlo algorithm.
2. The data-driven new energy joint output scenario generation system according to claim 1 is characterized in that: It also includes: The scenario reduction and evaluation module is used to cluster the random scenario set of renewable energy joint output using a clustering algorithm to obtain representative scenarios, and to evaluate the matching degree between the representative scenarios and the preprocessed historical output data of wind power output and photovoltaic output through KL divergence.
3. The data-driven new energy joint output scenario generation system according to claim 1 is characterized by: The specific methods for preprocessing historical output data of wind power output and photovoltaic output include: The historical output data of wind power and photovoltaic power are normalized and preprocessed. The minimum and maximum normalization is used to linearly map the historical output data of wind power and photovoltaic power to the [0,1] per-unit value interval. The preprocessed historical output data of wind power and photovoltaic power are obtained. The formula is: Among them, m is the historical output data of wind power output and photovoltaic output, m min is the minimum value of the historical output data of wind power output and photovoltaic output, m max is the maximum value of the historical output data of wind power output and photovoltaic output, and m′ is the historical output data of wind power output and photovoltaic output after preprocessing.
4. The data-driven new energy joint output scenario generation system according to claim 1 is characterized by: The specific process of fitting the pre-processed historical output data of wind power output and photovoltaic output through probability density function to obtain the marginal distribution of wind power output and photovoltaic output is as follows: The Weibull distribution probability density function of the marginal distribution of wind power output is obtained as follows: Where v is the wind speed; λ is the scale parameter of the Weibull distribution; k is the shape parameter of the Weibull distribution; is the scaling factor; Reflects the power relationship of wind speed v relative to scale parameter λ; is the index term; The beta distribution probability density function of the photovoltaic output marginal distribution is obtained as follows: Where P is the photovoltaic output-related variable; α is the first shape parameter of the Beta distribution; β is the second shape parameter of the Beta distribution; P α-1 Indicates the impact of the value of α on the probability density; (1-P) β-1 represents the probability density characteristics of photovoltaic output; B(α,β) is the beta function.
5. The data-driven new energy joint output scenario generation system according to claim 1 is characterized in that: The correlation between wind power output and photovoltaic output is analyzed by using the Kendall rank correlation coefficient and tail dependence coefficient on the pre-processed historical output data of wind power output and photovoltaic output. The specific process of obtaining the matching error between the set connection functions and the Kendall rank correlation coefficient and tail dependence coefficient of the pre-processed historical output data of wind power output and photovoltaic output is as follows: The Kendall rank correlation coefficient τ is calculated by statistically analyzing the synergistic change ratio of historical output data of wind power output X and photovoltaic output Y. The calculation formula is as follows: Among them, N c is the number of same-direction logarithms at all moments in the historical output data. i >x j ) and (y i >y j ), or (x i <x j ) and (y i <y j ),(x i ,y i ) and (x j ,y j ) is the same direction logarithm; N d is the number of reverse logarithms of all moments in the historical output data. i >x j ) and (y i <y j ), or (x i <x j ) and (y i >y j ),(x i ,y i ) and (x j ,y j ) is the reverse logarithm; n represents the total number of moments in the historical output data; x i represents the wind power output at time i, x j represents the wind power output at time j, y i represents the photovoltaic output at time i, y j represents the photovoltaic output at time j, and τ represents the Kendall rank correlation coefficient, which is used to measure the intensity of the trend of simultaneous increase or decrease of wind power output X and photovoltaic output Y; Calculate the tail dependence coefficient λ of the preprocessed historical output data of wind power output and photovoltaic output. The tail dependence coefficient λ includes the upper tail dependence coefficient λ u and the lower tail dependence coefficient λ I , upper tail dependence coefficient λ u It represents the probability of wind power output and photovoltaic output reaching maximum value at the same time, and the lower tail dependence coefficient λ l Indicates the probability that wind power output and photovoltaic output reach their minimum values at the same time; The upper tail dependence coefficient calculation formula is: Among them, λ u represents the upper tail dependence coefficient, q represents the quantile threshold, represents the inverse cumulative distribution function of the wind power output X variable, represents the inverse cumulative distribution function of the photovoltaic output Y variable, Wind power output And photovoltaic output probability; The formula for calculating the lower tail dependence coefficient is: Among them, λ1 represents the lower tail dependence coefficient, Wind power output And photovoltaic output probability; Based on the correlation analysis between wind power output and photovoltaic output, the matching errors of each set connection function are obtained, including the Kendall rank correlation coefficient matching error, the upper tail dependence coefficient matching error, and the lower tail dependence coefficient matching error; The Kendall rank correlation coefficient matching error calculation formula is: where τ 理论 is the theoretical Kendall rank correlation coefficient of each link function obtained by setting the parameters of each link function; τ 样本 is the Kendall rank correlation coefficient τ of the preprocessed historical output data of wind power output and photovoltaic output; The upper tail dependency coefficient matching error calculation formula is: where λ u,理论 is the theoretical tail dependence coefficient of each link function obtained by setting the parameters of each link function; τ 样本 is the upper tail dependence coefficient λ of the preprocessed historical output data of wind power output and photovoltaic output u ; The formula for calculating the matching error of the lower tail dependence coefficient is: where λ I,理论 is the theoretical tail dependence coefficient of each link function obtained by setting the parameters of each link function; I,样本 is the lower tail dependence coefficient λ of the preprocessed historical output data of wind power output and photovoltaic output l .
6. The data-driven new energy joint output scenario generation system according to claims 4 and 5 is characterized in that: By calculating the Euclidean distance between each set connection function and the preprocessed historical output data of wind power output and photovoltaic output, and combining the matching error to select the optimal connection function, the specific process of establishing the joint probability distribution model of wind power output and photovoltaic output based on the optimal connection function and the marginal distribution of wind power output and photovoltaic output is as follows: Calculate the Euclidean distance between each copula function and the preprocessed historical output data of wind power output and photovoltaic output. The calculation formula of the Euclidean distance matching method is as follows: Where n represents the total number of moments in the historical output data, u i represents the historical output data of wind power after preprocessing, v i Represents the historical output data of photovoltaic power after preprocessing, C e represents the real dependency value directly calculated based on the pre-processed historical output data of wind power output and photovoltaic output, C t Indicates the theoretical dependence value of each set connection function Copula; Selecting a connection function Copula with the smallest Euclidean distance from the preprocessed historical output data of wind power output and photovoltaic output, obtaining the Kendall rank correlation coefficient matching error, the upper tail dependence coefficient matching error, and the lower tail dependence coefficient matching error of the connection function Copula with the smallest Euclidean distance from the matching errors calculated above, and determining whether the Kendall rank correlation coefficient matching error, the upper tail dependence coefficient matching error, and the lower tail dependence coefficient matching error of the connection function Copula with the smallest Euclidean distance are all less than preset values; if so, the connection function is used as the optimal connection function Copula; The optimal connection function Copula is combined with the marginal distribution of wind power output and photovoltaic output to generate the joint probability distribution model P of wind power output and photovoltaic output. X,Y (x,y): P X,Y (x,y)=C θ (F X (x),F Y (y)) Among them, C θ is the optimal connection function Copula, x represents the value of wind power output in the [0,1] per unit value interval, F X (x) represents the probability value of wind power output ≤ x, y represents the value of photovoltaic output in the per-unit value interval [0,1], F Y (y) represents the probability value of photovoltaic output ≤ y.
7. The data-driven new energy joint output scenario generation system according to claim 6 is characterized by: The specific process of generating the random scenario set of new energy joint output is as follows: The Markov chain is used to define the state space discretization rule, and the output interval S of the joint probability distribution model of wind power output and photovoltaic output is divided into n1 states according to the preset interval: Among them, s i Indicates the i-th output state interval, a i Indicates the lower limit of the output state range, b i Shows the upper limit of the force state range; Construct the state transition probability matrix of wind power output and photovoltaic output: Among them, N i→j Represents the i-th output state interval s in the historical output data i To the jth output state interval s i Number of jumps, N i→k Represents the range from the i-th output state interval to the k-th output state interval s in the historical output data k Number of jumps, P ij Representing the jump probability, the state transition probability matrix of wind power output and photovoltaic output is used to count the state jump frequency of adjacent time periods in the historical output data, and combined with the joint probability distribution model to form a joint probability distribution model with time correlation; The state transition probability matrix is used as the target distribution Π(i), which represents the ideal state probability distribution; the Gaussian distribution is used as the proposed distribution Q(j|i), which represents the proposed jump direction probability; Extract candidate output state interval s from the proposal distribution Q(j|i) * , the acceptance probability a of state transition is calculated by Markov chain Monte Carlo algorithm, and the calculation formula is as follows: Among them, Π(s * ) represents the candidate output state interval s * The probability of being in the target distribution Π(i), Π(s t ) represents the current output state interval s t The probability of being in the target distribution Π(i), Q(s t |s * ) represents the candidate output state interval s * To the current output state interval s t The proposed probability density, Q(s * |s t ) indicates the interval from the current output state s t To the candidate output state interval s * The proposed probability density of ; The output state interval is updated according to the Markov Chain Monte Carlo criterion: Among them, u represents the random variable that controls the probabilistic state transition, α represents the acceptance probability of the state transition, and s t+1 The updated output state interval, t+1 represents the next moment after time t; Get the updated output state interval s t+1 , repeat the updating process of the output state interval and generate scenarios, and finally generate a set of random scenarios of new energy joint output.
8. The data-driven new energy joint output scenario generation system according to claims 2 and 7 is characterized in that: The clustering algorithm is used to cluster the random scenarios of renewable energy combined output to obtain representative scenarios. The KL divergence is used to evaluate the matching degree between the representative scenarios and the pre-processed historical output data of wind power output and photovoltaic output. The specific process is as follows: Construct the average wind power output μ w , average photovoltaic output μ p , wind power output standard deviation σ w , photovoltaic output standard deviation σ p , wind power output skew w , photovoltaic output skew p , maximum wind power ramp rate Photovoltaic maximum ramp rate and the proportion of zero output period f zero The multidimensional feature vector of each scene in the random scene set is obtained i as follows: in, Represents a 9-dimensional feature vector; In the eigenvector, the average wind power output μ w The calculation formula is as follows: Among them, X t is the per-unit value of wind power output at time t, where T represents T moments; Average photovoltaic output μ p The calculation formula is as follows: Among them, Y t is the per-unit photovoltaic output at time t; Wind power output standard deviation σ w The calculation formula is as follows: Among them, μ w The average output of wind power; Photovoltaic output standard deviation σ p The calculation formula is as follows: Among them, μ p is the average photovoltaic output; Wind power output skew w The calculation formula is as follows: Among them, σ w is the standard deviation of wind power output; Photovoltaic output skew p The calculation formula is as follows: Among them, σ p is the standard deviation of photovoltaic output; Maximum wind power ramp rate The calculation formula is as follows: Among them, X t+1 is the per-unit wind power output at time t+1, and Δt is the time interval between time t and time t+1; Photovoltaic maximum ramp rate The calculation formula is as follows: Among them, Y t+1 The per-unit value of photovoltaic output at time t+1; The proportion of zero output period f zero The calculation formula is as follows: Among them, Count(X t <0.05 or Y t <0.05) indicates that the statistical X t <0.05 or Y t Number of moments <0.05; The k-means++ algorithm is used to cluster the random scene set of new energy joint output to obtain representative scenes, and a scene is randomly selected from the random scene set of new energy joint output as the initial cluster center c1, and subsequent scenes are selected as cluster centers c i The probability (f i ) is calculated as follows: Among them, D(f i ) is the distance from scene i to the nearest cluster center, Entropy(f i ) is the statistical entropy of scene i, λ is the entropy weight coefficient; For each scene i, assign it to the corresponding cluster k: Among them, the weighted distance The calculation formula is as follows: Among them, f i is the feature vector of the i-th scene, c k is the center vector of the kth cluster, W is the diagonal weight matrix, (f i -c k ) represents the difference vector between the feature vector of the i-th scene and the center vector of the k-th cluster in the feature space, (f i -c k ) T represents the transpose of the difference vector between the feature vector of the i-th scene and the center vector of the k-th cluster in the feature space; Climbing rate penalty The calculation formula is as follows: Among them, R i is the climbing rate vector of the current scene i, is the historical average climbing rate of cluster k, is the wind power ramp rate of the current scenario i, is the historical wind power ramp rate of cluster k, is the photovoltaic ramp rate of the current scenario i, is the historical PV ramp rate of cluster k; Perform clustering iteration on cluster k to obtain k representative scenes; The KL divergence is used to evaluate the matching degree between the representative scenarios and the pre-processed historical output data of wind power output and photovoltaic output. KL (P hist ||P gen ) is calculated as follows: Among them, P hist is the true probability distribution of the historical output data of wind power output and photovoltaic output after preprocessing, P gen is the probability distribution of the generated representative scenes, P hist (i) is the true probability distribution of the historical output data of wind power output and photovoltaic output after preprocessing in the i-th discretization interval, P gen (i) is the probability distribution of the representative scene generated by the i-th discretization interval, i is the discretization interval number, N is the total number of intervals, and || is used to convert P hist With P gen Separated, indicating the comparison object of KL divergence; According to D KL (P hist ||P gen ) meets the preset threshold range to evaluate the matching degree between the representative scenario and the pre-processed historical output data of wind power output and photovoltaic output.
9. A data-driven method for generating a new energy joint output scenario, characterized in that: It includes the following steps: Preprocess the historical output data of wind power and photovoltaic power; The wind power output marginal distribution and the photovoltaic output marginal distribution are obtained by fitting the pre-processed historical output data of wind power output and photovoltaic output respectively through the probability density function, and the correlation between wind power output and photovoltaic output is analyzed on the pre-processed historical output data of wind power output and photovoltaic output through the Kendall rank correlation coefficient and the tail dependence coefficient, and the Kendall rank correlation coefficient matching error and the tail dependence coefficient matching error of each set connection function and the pre-processed historical output data of wind power output and photovoltaic output are obtained. The Euclidean distance between each set connection function and the pre-processed historical output data of wind power output and photovoltaic output is calculated, and the optimal connection function is selected in combination with the Kendall rank correlation coefficient matching error and the tail dependence coefficient matching error. Based on the optimal connection function and in combination with the wind power output marginal distribution and the photovoltaic output marginal distribution, a joint probability distribution model of wind power output and photovoltaic output is established; A Markov chain is used in combination with the joint probability distribution model to form a joint probability distribution model with time correlation, and a random scenario set of new energy joint output is generated by sampling from the joint probability distribution model with time correlation through the Markov chain Monte Carlo algorithm.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to claim 9 are implemented.
Citation Information
Cited By
Distribution line load prediction and optimal scheduling method and system
CN121124057A
A power distribution line load forecasting and optimal scheduling method and system thereof
CN121124057B