Probability modeling-based large-scale new energy station output convergence method
Through probabilistic modeling and Monte Carlo simulation, the output timing sequence of new energy stations is generated, which solves the problems of lack of data and low power supply reliability in the output aggregation of new energy stations in remote areas, and realizes the accurate simulation and aggregation of new energy cluster output.
Patent Information
- Application Number
- CN202510307064.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-15
- Publication Date
- 2025-07-18
AI Technical Summary
In the modeling of the output probability of new energy stations in remote areas, the existing technology lacks accurate assessment of meteorological data, and fails to effectively consider the situation of maintenance and shutdowns and faults, resulting in low power supply reliability and difficulty in achieving the output of new energy clusters.
Through a probabilistic modeling method, the Markov chain and Monte Carlo simulation are used to generate the output timing sequence of the new energy station, simulate the output characteristics of the surrounding stations, and gather the total output of the new energy cluster while taking into account maintenance and failures.
It has achieved accurate output simulation and convergence in new energy bases in remote areas, improved the economic and reliability of system operation, and solved the problem of low power supply reliability in remote areas.
Smart Images

Figure CN120341980A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of new energy power station planning, and specifically to a method for aggregating the output of large-scale new energy power stations based on probability modeling. Background Technique
[0002] The probability modeling of the output of large-scale new energy power stations (such as wind power and photovoltaic power) plays a crucial role in the planning and management of renewable energy systems. This probability modeling not only forms the core of optimizing the system operation strategy, but also provides the basic basis for energy market trading and energy storage planning. By accurately evaluating the probability distribution of the output of new energy power stations, the economic efficiency and reliability of system operation can be effectively improved, the effective integration of renewable energy and the sustainability of system operation can be realized.
[0003] Currently, the probability modeling of the output of new energy power stations mainly relies on historical data and does not consider the situation of missing meteorological data in remote areas. In this context, how to accurately grasp the spatio-temporal characteristics of renewable resources and generate comprehensive and feasible output scenarios has become an urgent problem to be deeply studied. To solve this problem, a large amount of output data needs to be generated to simulate the scenarios of new energy clusters and provide data support for the aggregation of the output of new energy clusters.
[0004] At the same time, the current research mainly focuses on urban microgrid systems, ignoring the special situations in remote areas such as fragile infrastructure, poor connection with the main grid, low power supply reliability, and frequent off-grid situations due to faults. Therefore, it is necessary to consider the maintenance outage and fault outage situations of new energy power stations, and at the same time aggregate the output of all new energy power stations on the same power generation side port of the DC transmission channel, and obtain the total output sequence of the new energy cluster connected to this port as the basis for subsequent research. Summary of the Invention
[0005] In view of the deficiencies of the prior art, the present invention proposes a method for aggregating the output of large-scale energy station clusters based on probability modeling that takes into account the outage status of new energy power stations. The method includes:
[0006] Step S1: Based on the normalized output data of the historical time series of the output of existing new energy power stations, statistically analyze its probability characteristics and the annual utilization hour index;
[0007] Step S2: Through discretization, extract the state transition matrix of the normalized output sequence of the new energy power station, establish the transition relationship between each state through a Markov chain, and thus generate a large number of new energy output time series with similar probability characteristics and annual utilization hour indexes to simulate the output characteristics of several surrounding power stations of the new energy power station;
[0008] Step S3: Based on the historical operation logs, maintenance records, or real-time monitoring data of existing new energy power stations, obtain the equipment failure and maintenance information, establish the time series of the operation states of all power stations, and simulate the operation states such as normal, maintenance, and failure of the power stations.
[0009] Step S4: Considering the maintenance and failure of new energy power stations, aggregate the power outputs of all new energy power stations at the same power generation side port of the DC transmission channel to obtain the total power output series of the new energy cluster connected to this port.
[0010] In the method of the present invention, aiming at the situation of limited existing reference data in the planned area of the Shagehuang new energy base, a large number of output time series of surrounding similar power stations of existing new energy power stations are simulated and generated, so as to simulate the scenarios of new energy clusters, complete the probability and time series modeling of the power outputs of large-scale new energy power stations in the Shagehuang new energy base, and propose a method for aggregating the power outputs of new energy clusters considering the maintenance outages and failure outages of new energy power stations, so as to solve the problem of simulating the power output characteristics of new energy clusters in the capacity configuration planning of new energy bases. Description of the Drawings
[0011] Figure 1 It is a flowchart of the steps of the method for aggregating the power outputs of large-scale new energy power stations based on probability modeling in the embodiment of the present invention.
[0012] Figure 2 It is a schematic diagram of the operation and outage cycle process of repairable components in the embodiment of the present invention.
[0013] Figure 3 It is a schematic diagram of the state space of repairable components in the embodiment of the present invention.
[0014] Figure 4 It is a schematic diagram of the state space of forced and planned outages in the embodiment of the present invention.
[0015] Figure 5 It is a schematic diagram of the state space of a separate planned outage in the embodiment of the present invention.
[0016] Figure 6 It is a schematic diagram of the sequential state transition process of components in the embodiment of the present invention.
[0017] Figure 7 It is a schematic diagram of the sequential state transition process of the system in the embodiment of the present invention.
[0018] Figure 8 It is a schematic diagram of the generation process of the component operation state sequence based on sequential Monte Carlo in the embodiment of the present invention.
[0019] Figure 9In the embodiment of the present invention, for the known output sequence of a single new energy power station, through a probability modeling method that takes into account its probability and time series characteristics, the schematic diagram of the implementation process of simulating the time series output sequences of multiple new energy power stations around it is shown.
[0020] Figure 10 The schematic diagram of the implementation process of the large-scale new energy power station output aggregation method based on probability modeling in the embodiment of the present invention is shown. Specific implementation manner
[0021] In this embodiment, for the new energy power stations around the Shagehuang New Energy Base, the large-scale new energy power station output aggregation method based on probability modeling has steps basically as shown in Figure 1, including:
[0022] Step S1: Based on the normalized output data of the historical time series of the existing new energy power station output, statistically analyze its probability characteristics and the annual utilization hour index;
[0023] Step S2: Through discretization, extract the state transition matrix of the normalized output sequence of the new energy power station, establish the transition relationship between each state through a Markov chain, so as to generate a large number of new energy output time series with similar probability characteristics and annual utilization hour indexes, and simulate the output characteristics of several surrounding power stations of the new energy power station;
[0024] Step S3: Based on the historical operation logs, maintenance records or real-time monitoring data of the existing new energy power stations, obtain the fault and maintenance information of the equipment, and establish the time series of the operation states of all power stations to simulate the operation states such as normal, maintenance and fault of the power stations;
[0025] Step S4: Considering the maintenance and faults of the new energy power stations, aggregate the outputs of all new energy power stations at the same power generation side port of the DC transmission channel to obtain the total output sequence of the new energy cluster connected to this port.
[0026] In the application of the Shagehuang New Energy Base plan, for a single new energy power station with an installed capacity of G rated , the time series output sequence set can be expressed as:
[0027] G Gen = {G(t)|t = 1, 2,..., N t} (1)
[0028] Among them, G Gen represents the time series output set of a single new energy power station, t is the time scale, and N t is the number of elements in the output set. In this embodiment, the new energy power station output data used are all hourly data, so the actual time interval between adjacent time scales is 1 hour.
[0029] Some data indicators involved in this example are specifically as follows:
[0030] (1) Average output and annual utilization hours of a single new energy power station
[0031] The average output of the new energy power station can be represented by its mean value. According to the relevant statistical theory, for the output sequence of this new energy power station, its output mean value can be calculated by the following formula.
[0032]
[0033] In the formula, E[] represents taking the mean value of the variable inside the square brackets.
[0034] In practical applications, the power generation capacity of a new energy power station can also be measured by the annual utilization hours of its power generation. Select the output sequence of one year in the time-series output set G of a single new energy power station Gen in the output sequence of one year, so that the number of elements N in the time-series output set t should at least satisfy N t ≥ 8760, then the annual utilization hours of the power generation of this new energy power station can be calculated by the following formula (assuming starting from time t = 1):
[0035]
[0036] In the formula, h Gen represents the annual utilization hours of the power generation of this new energy power station.
[0037] Observing the above formula, it can be seen that when the time step of the time-series output sequence is 1 hour and the number of elements N in the time-series output sequence set t = 8760, the mean value of the output of a single new energy power station and its annual utilization hours satisfy the following relationship:
[0038] h Gen = 8760E[G] (4)
[0039] According to the above formula, it can be seen that if the given time-series output of a single new energy power station is exactly one full year of data, its annual utilization hours and its average output can be directly converted through the coefficient 8760. At this time, the two indicators of annual utilization hours and average output are equivalent and can both represent the power generation capacity of a single new energy power station.
[0040] (2) Output variance of a single new energy power station
[0041] The average output and annual utilization hours of a single new energy power station are indicators used to describe the power generation capacity of this power station, while the overall degree of deviation of the output of this power station from its average output can be represented by the variance of its output sequence. The variance of the output of a single new energy power station can be calculated by the following formula:
[0042] D[G] = E[G(t) 2-E[G(t)] 2 (5)
[0043] Where D[] represents taking the variance of the variable within the square brackets. If the variance of the output sequence of a certain new energy power station is larger, it indicates that the overall output situation deviates more from its average output; if the variance of this output sequence is smaller, the degree to which its overall output situation deviates from its average output is lower.
[0044] In addition, the standard deviation of the output of the new energy power station can be calculated by the following formula:
[0045]
[0046] Where S[] represents taking the standard deviation of the variable within the square brackets. Obviously, the standard deviation of the output sequence of the new energy power station is the arithmetic mean of its variance. This standard deviation index can be interpreted in a similar way to the variance.
[0047] (3) Autocorrelation function of the time-series output of a single new energy power station
[0048] The autocorrelation function (ACF) is an important tool in time-series analysis, used to measure the correlation between data values of the same time series at different time points. In the calculation of the ACF, the length of the lag time used is gradually increased, so that the ACF graph of this time-series data can be drawn, and thus the characteristics of the time-series output sequence of a single new energy power station over time can be represented by its ACF graph.
[0049] Let k be the length of the lag time, k = 1, 2, …, N t -t. For the output sequence G(t) of the new energy power station, t = 1, 2, … N t , the ACF of the output at time t and the output at time t + k can be defined by its Pearson correlation coefficient:
[0050]
[0051] Where ACF(k) represents the autocorrelation function value of this time-series output sequence when the lag time length is k, also called the autocorrelation coefficient value. Cov[] represents taking the covariance of the two variables within the square brackets. Among them, the value range of the autocorrelation coefficient is [-1, 1]. When the autocorrelation coefficient value is 1, it indicates that this time-series output sequence is completely positively correlated at this lag time length; when the autocorrelation coefficient value is -1, it indicates that this time-series output sequence is completely negatively correlated at this lag time length; when the autocorrelation coefficient is 0, it indicates that this time-series output sequence has no correlation at this lag time length.
[0052] Specifically, the ACF graph of the time-series output of a single new energy power station can be calculated through the following steps.
[0053] Step 1: Obtain the average output of the new energy power station.
[0054] Step 2: Let the lag time length k = 1 as the initial value for calculating the ACF.
[0055] Step 3: Obtain the deviation of the data at each time point relative to the average output, as shown in the following formula.
[0056]
[0057] Where, ΔG(t) and ΔG(t + k) are the deviation magnitudes of the output at time t and t + k relative to the average output, respectively.
[0058] Step 4: Obtain the covariance between the output at the current time and the output after k time points under the current lag time length, which is calculated by the following formula.
[0059]
[0060] Step 5: Divide the covariance in the above formula by the variance of the time series output sequence of the new energy power station to obtain the autocorrelation coefficient at the current lag time length.
[0061] Step 6: Judge the value of the lag time length. If k = N t - t, the calculation ends, and the results of the ACF at all lag time lengths are obtained, and the drawing of the ACF graph is completed; if k < N t - t, then jump back to Step 3.
[0062] In this example, the specific implementation process of Step S1 is as follows:
[0063] S1-1: Obtain the normalization of the time series output sequence of a single new energy power station
[0064] For the convenience of calculation, the time series output sequence of the new energy power station can be first normalized, and the normalization method is shown in the following formula.
[0065]
[0066] Where, G * (t) is the normalized output of the new energy power station at time t, and its value is equal to the ratio of the actual output at that time to the station capacity. After normalization, the output value of the new energy power station at any time is between 0 and 1, that is, G * (t) ∈ [0, 1].
[0067] S1-2: Obtain the statistical characteristics of the output subsequence of the new energy power station
[0068] Analyze the historical output data of new energy power stations, divide the normalized time-series output subsequences according to the types of new energy power stations, and obtain the statistical characteristics of each subsequence; if the new energy power station is a wind farm, divide its normalized time-series output into 12 groups according to months, and statistically calculate the maximum value G max , the minimum value G min , and the average value G mean and other statistical characteristics. However, in addition to the differences in different months, the characteristics of PV output have obvious monotonically increasing and decreasing intervals over time during the day. Therefore, for PV, first remove the nighttime periods when the PV output is zero, and perform corresponding statistics by month and morning / afternoon periods, resulting in a total of 24 groups of subsequences.
[0069] In this example, the specific process of step S2 includes the following content:
[0070] S2-1: State discretization of the normalized output sequence
[0071] The state transition matrix is established based on the output state, and the Markov chain is also established based on the state transition matrix. Therefore, it is necessary to discretize the normalized output sequence to determine the output state at each moment. For each subsequence divided in S1-2, the same number of discrete states can be used.
[0072] S2-2: Obtaining the state transition matrix
[0073] Let the number of output states after division be n. Then, the output state corresponding to the normalized output value at time t can be determined by the following formula.
[0074]
[0075] In the formula, s i represents the i-th output state of the new energy power station; the above formula indicates that if the normalized output of the new energy power station at time t is within the interval [(i - 1) / n, i / n], it is considered to be in the i-th output state.
[0076] For two different output states s i and s j (i ≠ j), statistically calculate the output state situation of the normalized output sequence from time t to time t + 1. After traversing the entire sequence, if the current output state is s i , and the number of times the next output state is s j is N (ij) , and the number of times its output state is s i is N (i) , then for the entire output sequence, the new energy output transfers from state s i to state s jThe probability can be calculated by the following formula:
[0077]
[0078] where p ij represents the probability that the new energy output transfers from state s i to state s j , and Pr ob{} represents the probability of the event in the brackets occurring.
[0079] Based on the entire normalized output sequence, the state transition matrix of the new energy power station output can be obtained:
[0080]
[0081] In addition, if a certain output state does not exist in the normalized output sequence, the corresponding rows and columns of the state transition matrix are recorded as NAN (Not a number). Obviously, excluding the NAN elements in the state transition matrix, the sum of the state transition probabilities of each row or each column is equal to 1.
[0082] S2-3: Determine the normalized output value at the initial moment of the surrounding power station output subsequence to be generated
[0083] Based on the maximum and minimum values of the output of each subsequence obtained by statistics in S1-2 and Taking as the lower bound of the uniform distribution and taking as the upper bound of the uniform distribution, a normalized output value at the initial moment is randomly generated as the output value at the initial moment of the corresponding subsequence of the surrounding new energy power station. In addition, since the divided subsequences have an obvious order, such as the month order of the wind power sequence and the month and morning / afternoon order of the photovoltaic sequence, the output at the end of the previous sequence after simulation can be used as the output at the initial moment of the subsequent sequence.
[0084] Let the current moment t = 1 of the first sequence, the output value be represented as G * (t), and the length of the subsequence be L.
[0085] S2-4: Simulation generation of the normalized output subsequence of the surrounding new energy power station
[0086] Determine the output state s(t) = s * corresponding to the current output value G i (t). If the current output state is set as state s i , and the next output state is state s j , and its cumulative probability can be calculated by the following formula.
[0087]
[0088] where p sum,ij i.e., the cumulative probability of reaching state s i under the condition of s j .
[0089] Then, sample the next state s i according to the current state s j . Generate a random number u uniformly distributed in the interval [0, 1] using simple random sampling, compare u with the cumulative probability distribution. When the following conditions are met, select the corresponding state j as the next state and convert it to the corresponding power value according to the following formula.
[0090] p sum,i(j-1) <u ≤ p sum,ij (15)
[0091] G * (t + 1) ~ U[G * j,min , G * j,max (16)
[0092] where G * j,min and G * j,max are the upper and lower limits of the power interval of state j, and G * (t + 1) is the normalized output at time t + 1. If the magnitude distribution of the normalized output corresponding to this output subsequence is between [0, 1] and the number of discrete states divided is n, then G * j,min and G * j,max satisfy the following formula.
[0093]
[0094] S2 - 5: Check whether the generation of the current subsequence is completed
[0095] Let t = t + 1. If t > L, it indicates that the samples of this sequence have been completely generated, and go to S2 - 8; otherwise, return to S2 - 4.
[0096] S2 - 6: Check whether samples have been completely generated for all subsequences
[0097] If the number of generated subsequences is less than 12 (for wind power) or 24 (for photovoltaic), let G * (1) = G * (t - 1), set L to the length of the next sequence, set the time scale t to 1, and return to S2 - 4; if new normalized output subsequences have been generated for all sub - output sequences of the existing new - energy power stations, then go to S2 - 7.
[0098] S2-7: Concatenate the newly generated normalized output subsequences
[0099] Concatenate all the above-generated normalized output subsequences in the chronological order of month and morning / afternoon to obtain a set of peripheral station output simulation data with a length of 8760; if more annual output sequences need to be generated, return to S2-3, so as to realize the generation of annual output sequences of multiple stations around the current new energy station.
[0100] Based on the above steps, wind power generation and photovoltaic power generation data with randomness and chronology and any sequence length can be generated. In the research, the operation status of new energy generation equipment is often simulated in units of years. Therefore, the time span of the output sequence generated by each wind farm or photovoltaic power station is 8760 hours. Among them, the normalized sequence of the output of the m-th power station generated is denoted as G * m .
[0101] In this example, the specific process of step S3 includes the following contents:
[0102] S3-1: Establish a system component outage model, and the following factors need to be considered:
[0103] (1) Repairable forced failure
[0104] Repairable forced failure can be simulated through a steady-state "operation-outage-operation" cycle process. Figure 1 and Figure 2 are the cycle process diagram and the state transition diagram respectively. The average unavailability rate in the long-term cycle process can be expressed in mathematical form by one of the following three definitions:
[0105]
[0106] where λ is the failure rate (number of failures / year); μ is the repair rate (number of repairs / year); MTTR is the mean time to repair (hours); MTTF is the mean time to failure (hours); f is the average failure frequency (number of failures / year).
[0107] These three definitions are essentially the same. Only two of the parameters in equations (19) to (21) are independent. In other words, as long as two of them are known, the remaining parameters can be deduced. The following relationships can be obtained from equations (19) to (21).
[0108] Let: d = MTTF / 8760 and r = MTTR / 8760, then d and r are MTTF and MTTR in units of years, and we can get
[0109]
[0110]
[0111] U = fr(25)
[0112]
[0113] It should be emphasized that f and λ are two completely different parameters, which can be clearly seen from Eqs. (26) and (27). In most cases, the values of f (or λ) and r are very small, so f and λ are numerically close. However, in some special cases, the value of r is very large, such as the repair time of underwater cables. In this case, substituting f for λ or λ for f will cause a rather large error.
[0114] (2) Scheduled outage
[0115] Due to various reasons such as maintenance, replacement, refurbishment, or certain operating requirements, it may be necessary to schedule a planned outage. The planned outage can be simulated in two ways. The first method is to assume that the planned outage and restoration times follow a given distribution, and the parameters of this distribution can be estimated from the statistical data of the planned outage events; the second method is to regard the planned outage as an event scheduled in a predetermined time period. In the first method, the planned outage is regarded as a random event. Figure 4 Give the state space diagrams of forced failure and planned outage. Applying the Markov method to this state space diagram, the following results can be obtained:
[0116]
[0117] where P up , P fo and P po are the probabilities of the operating, forced outage, and planned outage states respectively; f p and f are the frequencies (number of outages / year) of the planned and forced outage states respectively; λ p and λ are the transition rates of the planned and forced outage states respectively; μ p and μ are the repair rates (number of repairs / year) for restoration from the planned and forced outage states respectively.
[0118] In most data acquisition systems, the outage rates (λ p and λ) are not directly collected, but the outage frequencies and repair times of components are statistically counted. Therefore, strictly speaking, before calculating the state probabilities using Eqs. (28) to (30), λ p and λ should be calculated using Eqs. (31) and (32) first, which leads to the complexity of the input data preprocessing. From the perspective of approximate calculation, it can be assumed that the planned outage and forced outage of components are not mutually exclusive. Based on this assumption, such as Figure 5As shown, forced outages and scheduled outages are each represented separately by a two-state model. Thus, the following model can be used for scheduled outages:
[0119]
[0120] where λ p , μ p and f p are defined as above; MTTR p = 8760 / λ p is the average time before scheduled outage (hours); MTTR p = 8760 / μ p is the average repair time for scheduled outage (hours); U p is the unavailability rate caused by scheduled outages, which has the same meaning as P po in Equation (30).
[0121] All parameters are based on the average values of scheduled outage statistical data. The above assumptions can greatly reduce the data preparation and calculation workload for risk assessment in most cases, and will not cause large errors. Applying Figure 4 the combined model shown, the total outage probability is:
[0122]
[0123] Applying the model with forced and scheduled outages separated, the total outage probability is:
[0124]
[0125] Comparing (36) and (37), it can be seen that the error in Equation (37) is only caused by a term λ p λ that exists in both the denominator and the numerator. Since it is known that λ p << μ p and λ << μ, this second-order error can be ignored. It should be noted that when the duration of the scheduled outage is very long, the error caused by this assumption will increase, and this assumption should be used with caution at this time.
[0126] Each of the two methods has its own advantages and disadvantages. The first method cannot accurately reflect the planned schedule and is applicable to the long-term planning of systems with unknown planned schedules. The second method can be applied to short-term operation planning, where either the start and end times of the scheduled outage are known, or the start and end times of the optimal scheduled outage are sought through risk assessment methods.
[0127] S3-2: Monte Carlo simulation.
[0128] (1) Sequential Monte Carlo simulation method
[0129] The sequential Monte Carlo method is a simulation carried out over a time span in chronological order. There are different methods for establishing the virtual system state transition cycle process. The most common one is the so-called state duration sampling method discussed here.
[0130] The state duration sampling method is based on sampling the probability distribution of the state duration of components, and it is divided into the following steps:
[0131] Step 1: Specify the initial states of all components. Usually, it is assumed that all components start in the operating state.
[0132] Step 2: Sample the duration of each component staying in the current state. The probability distribution of the state duration should be set. For different states, such as the operating or repair process, different probability distributions of the state duration can be assumed. For example, the following formula gives the sampled value of the state duration of the exponential distribution:
[0133]
[0134] where R i is a random number uniformly distributed in the [0,1] interval corresponding to the i-th component. If the current state is the operating state, then λ i is the failure rate of the i-th component; and if the current state is the outage state, λ i is the repair rate of the i-th component.
[0135] Step 3: Repeat Step 2 within the studied time span (a large number of sampled years), and record the sampled values of the duration of each state of all components. Then, the chronological state transition process of each component within the given time span can be obtained, as Figure 6 shown.
[0136] Step 4: Combine the state transition processes of all components to establish the system chronological state transition cycle process, as shown in the following figure.
[0137] Step 5: Calculate the risk index function through the system analysis of each different system state. Since the occurrence, duration, and consequences of the system failure states can be clearly determined and recorded in the system state transition cycle process, the calculation of the system risk index is simple and intuitive.
[0138] It can be seen that the key of the sequential Monte Carlo method lies in the generation of the system state transition process. Once this step is completed, the index calculation is relatively simple. The essence of this method is to establish a virtual transfer cycle process of system operation and failure.
[0139] For each component, the process of generating its operating state sequence using the sequential Monte Carlo simulation method is as Figure 8 shown. Figure 8Among them, n is the number of years of the i-th generated operation status sequence, and s(t) is the operation status of the i-th component at the t-th hour. If s(t)=0, it indicates that the component is in the outage state at time t; if s(t)=1, it indicates that the component is in the normal operation state at time t; round() represents the operation of rounding the value in the parentheses to an integer; the expression s(t:t+T-1)=1 means that all the states at all times between time t and time t+T-1 are set to 1.
[0140] (2) Convergence Criterion of Monte Carlo Simulation
[0141] The basic idea of Monte Carlo simulation is to use a random number sequence to generate a series of experimental samples. When the number of samples is large enough, according to the central limit theorem or the law of large numbers, the sample mean can be used as an unbiased estimate of the mathematical expectation. The variance of the sample mean is an indicator of the estimation accuracy.
[0142] When calculating the reliability or other technical indicators of the system by Monte Carlo simulation, the calculation result of the i-th Monte Carlo sample is x i , then x i 's sample mean can be calculated by the following formula:
[0143]
[0144] In the formula, N is the number of samples of Monte Carlo sampling.
[0145] Its sample variance is defined as
[0146]
[0147] The estimate of the sample mean given in formula (39) is also a random variable, which depends on the number of samples and the sampling process. The uncertainty of the estimate can be measured by the variance of the sample mean, and its definition is
[0148]
[0149] It can be seen that: the sample variance V(x) given by formula (40) and the variance of the sample mean given by formula (41) are two different concepts and cannot be confused with each other.
[0150] The standard deviation of the sample mean is σ
[0151]
[0152] Equation (42) shows that there are two measures to reduce the standard deviation of the estimator in Monte Carlo simulation: either increasing the number of samples or reducing the sample variance. There are many variance reduction techniques to improve the efficiency of Monte Carlo simulation. It is important to note that in any case, the variance cannot be reduced to zero, so a reasonable and sufficiently large number of samples always needs to be considered.
[0153] Monte Carlo simulation produces a convergent process with fluctuations, and it cannot be guaranteed that a smaller error will definitely be obtained by increasing a small number of samples. However, it is clear that the upper and lower bounds or confidence ranges of the error will decrease as the number of samples increases. The accuracy level achieved by Monte Carlo simulation can be measured by the coefficient of variation η, which is defined as the standard deviation of the estimator divided by the estimator, as follows:
[0154]
[0155] The coefficient of variation is commonly used as a convergence criterion for calculating reliability indicators in reliability assessment, and generally takes values in the range of 0.05 - 0.1.
[0156] In step S4 of this example, for a certain port on the power supply side of the transmission channel, if the normalized sequence of the output of the m-th new energy power station is represented as G * m , then the output of the new energy power station cluster at this port can be calculated by the following formula.
[0157]
[0158] In the formula, G Terminal (t) is the output of the new energy power station cluster at this port at time t, G * m (t) is the normalized output of the m-th new energy power station at time t, and G m,rated is the installed capacity of the m-th new energy power station.
[0159] Combining the previous steps, according to the known output sequence of a single new energy power station, through a probability modeling method that takes into account its probability and time series characteristics, the specific implementation process of simulating the time series output sequences of multiple surrounding new energy power stations is as Figure 9 shown.
[0160] Note that Figure 9 The left half of is the modeling process of the time series output of existing new energy power stations, which includes two steps: subsequence division of the time series processing sequence of existing power stations and establishment of the subsequence state transition matrix. To accurately take into account the time series characteristics and probability characteristics of the output of existing new energy power stations, this time series output sequence can include data for multiple years. And Figure 9The right half is the generation process of the annual output sequence of surrounding power stations in the existing power station. For application in subsequent production simulations (with an annual simulation time length), the length of the generated time series output sequence should be 8760 hours (step size: 1 hour). In the application, for the historical output data of the same new energy power station, by executing Figure 9 the process shown, the annual output curves of multiple similar power stations around this power station can be simulated, so as to simulate the output scenario of the new energy power station cluster.
[0161] However, in the actual operation process of new energy power stations, there will be situations of outage due to faults or planned maintenance. During the outage period, normal power generation cannot be carried out, which further reduces the annual utilization hours of the DC transmission channel. Therefore, if the annual utilization hours of the DC transmission channel are used as one of the technical indicators to evaluate its performance, then when planning the capacity configuration of the power sources in the Shagehuang new energy base, the impact brought by the outage must be considered, otherwise there will be an overestimation of the annual utilization hours.
[0162] In order to consider the faults and maintenance of power generation equipment in the operation simulation and thus be used in the power source configuration process of the Shagehuang new energy base, in this embodiment, each new energy power station is taken as an object, and its probability models of faults and maintenance are respectively established. At the same time, time series samples of the operation states of each new energy power station are generated, and faults and maintenance are considered in units of individual new energy power stations.
[0163] For a single new energy power station in the simulation, given its normalized time series output sample G * m , if the time series operation state sequence of this power station is represented as S OS , in the production simulation, its actual normalized output can be obtained through the Hadamard product of G * m and S OS :
[0164] G * m,act = G * m ⊙ S OS (45)
[0165] In the formula, G * m,act is the actual normalized output sequence considering the faults and maintenance of the new energy power station. The symbol ⊙ represents the Hadamard product of vectors. The time series operation state sequence S OS is of the same length as G * mThe same, a one-dimensional vector whose elements are only composed of 0 and 1. If the corresponding state value at time t is 0, it indicates that the new energy power station is in a shutdown state at this moment, and the power output of the power station is 0 at this moment; if the corresponding state value at time t is 1, it indicates that the new energy power station is in a normal operation state at this moment.
[0166] If G * m,act (t), G * m (t) and S OS (t) respectively represent the values at the t-th moment in the time series G * m,act , G * m and S OS . Then the Hadamard product operation rule is:
[0167] G * m,act (t) = G * m (t)S OS (t) (46)
[0168] Obviously, according to the above formula, the key to obtaining the normalized output sequence of the new energy power station considering shutdown is to generate its time series operation state sequence S OS , and this process is implemented in step S3. The specific steps are as follows.
[0169] Step ①: Obtain the relevant data of the historical operation state of the source power station (i.e., the existing power station used to generate multiple surrounding power stations), and count its annual average maintenance and fault shutdown hours as h1 and h2 respectively; calculate its failure rate and repair rate through formulas (47) and (48).
[0170]
[0171] In the formula, λ and μ respectively represent the failure rate and repair rate of this power station.
[0172] Step ②: Initialize the time scale t = 1. Initialize the initial state of this new energy power station. Assume that the initial state of this power station in the annual operation state is the normal operation state, that is, s(t) = 1.
[0173] Step ③: Sample the state duration. If this new energy power station is currently in the normal operation state, that is, s(t) = 1, then calculate the number of hours of continuous normal operation through formula (49); if this new energy power station is currently in the shutdown state, that is, s(t) = 0, then calculate the number of hours of continuous shutdown through formula (50).
[0174]
[0175] In the formula, T represents the number of hours of the state duration, round() represents the operation of rounding the value in the parentheses to the nearest integer, and R is a uniformly distributed random number within the range of [0, 1].
[0176] Step ④: Update the operation status sequence according to the sampling result. If s(t) = 1, then set all operation statuses from time t to time t + T - 1 to 1, that is, s(t:t + T - 1) = 1, and change the status at the next moment s(t + T) = 0; if s(t) = 0, then set all operation statuses from time t to time t + T - 1 to 0, that is, s(t:t + T - 1) = 0, and change the status at the next moment s(t + T) = 1.
[0177] Step ⑤: Update the time scale and judge the end condition of the step. Update the time scale t = t + T. If the current t ≥ 8760, then the generated status sequence has reached one year of data. Let the annual operation status sequence S of the new energy power station OS = s(1:8760), and complete the generation of its time-series operation status sequence; otherwise, jump back to step ③.
[0178] Based on the above steps, combined with the normalized time-series output sequence of the new energy power station, the normalized time-series output sequence considering the outage of the new energy power station can be obtained. Perform the above operations on all the simulated time-series outputs of the new energy power station, and the aggregation result of the output of the new energy power station cluster considering the outage of the new energy power station can be obtained. The specific process of this example can be specifically summarized in the order of execution as Figure 9 shown.
Claims
1. A large-scale new energy power station output aggregation method based on probability modeling, characterized in that Including: Step S1: Based on the normalized output historical time series data of existing new energy power stations, statistically analyze its probability characteristics and annual utilization hours index; Step S2: Through discretization, extract the state transition matrix of the normalized output sequence of the new energy power station, establish the transition relationship between each state through the Markov chain, so as to generate a large number of new energy output time series with similar probability characteristics and annual utilization hours index, and simulate the output characteristics of several surrounding power stations of the new energy power station; Step S3: Based on the historical operation logs, maintenance records or real-time monitoring data of existing new energy power stations, obtain the fault and maintenance information of the equipment, establish the operation state time series of all power stations, and simulate the operation states such as normal, maintenance and fault of the power stations; Step S4: Considering the maintenance and faults of new energy power stations, aggregate the outputs of all new energy power stations at the same power generation side port of the DC transmission channel to obtain the total output sequence of the new energy cluster connected to the port.
2. The method according to claim 1, wherein Step S1 includes: S1-1: Obtain the normalization of the time series output sequence of a single new energy power station: Normalize the time series output sequence of the new energy power station, and the normalization method is shown in the following formula: Among them, G * (t) is the normalized output of the new energy power station at time t, and its value is equal to the ratio of the actual output at that time to the capacity of the power station; S1-2: Obtain the statistical characteristics of the output subsequence of the new energy power station: Divide the normalized time series output subsequence according to the new energy power station type to obtain the statistical characteristics of each subsequence; If the new energy power station is a wind farm, divide its normalized time series output into 12 groups according to months, and statistically calculate the maximum value G of each subsequence max , the minimum value G min and the average value G mean ; If the new energy power station is a photovoltaic power station, first eliminate the night time periods when the photovoltaic output is zero, and conduct corresponding statistics by month and morning / afternoon time periods, with a total of 24 sub-sequences. Statistically calculate the maximum value G of each sub-sequence max , the minimum value G min , and the average value G mean .
3. The method according to claim 2, wherein Step S2 includes: S2-1: State discretization of the normalized output sequence For each subsequence divided in S1-2, perform state discretization with exactly the same number of discrete states; S2-2: Obtain the state transition matrix Let the number of output states after division be n, then the output state corresponding to the normalized output magnitude at time t can be determined by the following formula: where s i represents the i-th output state of the new energy power station; the above formula indicates that if the normalized output of the new energy power station at time t is within the interval [(i - 1) / n, i / n], it is considered to be in the i-th output state; For two different output states s i and s j , i≠j, count the output state conditions of the normalized output sequence from time t to time t+1; after traversing the entire sequence, if under the condition that the current output state is s i , the number of times the next output state is s j is N (ij) , and the number of times its output state is s i is N (i) , then for the entire output sequence, the probability that the new energy output transfers from state s i to state s j can be calculated by the following formula: where p ij represents the probability that the new energy output transfers from state s i to state s j , and Pr ob{} represents obtaining the probability of the event within the brackets; Based on the entire normalized output sequence, obtain the state transition matrix of the output of the new energy power station as follows: If a certain output state does not exist in the normalized output sequence, the corresponding rows and columns of the state transition matrix are recorded as NAN; S2-3: Determine the normalized output magnitude at the initial moment of the output subsequence of the surrounding power stations to be generated The maximum and minimum values of the output of each subsequence obtained based on the statistics in S1-2 and Taking as the lower bound of the uniform distribution and as the upper bound of the uniform distribution, randomly generate the normalized output value at the initial moment, which is used as the output size of the corresponding new energy power station in the neighborhood at the initial moment of this subsequence; the output at the end after the simulation generation of the previous sequence is used as the output at the initial moment of the subsequent sequence; assume that the current moment of the first sequence is t = 1, and the output size is expressed as G * (t), and the length of the subsequence is L; S2-4: Simulate the generation of the normalized output subsequence of the surrounding new energy power stations Determine the output magnitude G at the current moment * (t) corresponds to the output state s(t)=s i , if the current output state is set to state s i , and the output state at the next moment is state s j , and its cumulative probability can be calculated by the following formula: where p sum,ij i.e., the cumulative probability of reaching state s i under condition s j ; According to the current state s i Sample the state s at the next moment j , generate a random number u uniformly distributed in the interval [0, 1] using simple random sampling, compare u with the cumulative probability distribution, and when the following conditions are met, select the corresponding state j as the state at the next moment and convert it to the corresponding power value according to the following formula. p sum,i(j-1) <u ≤ p sum,ij (15) G * (t + 1) ~ U[G * j,min ,G * j,max (16) where G * j,min and G * j,max are the upper and lower limits of the power interval in state j, and G * (t + 1) is the normalized output at time t + 1. If the magnitude distribution of the normalized output corresponding to this output subsequence is between [0, 1] and the number of discrete states divided is n, then G * j,min and G * j,max satisfy the following equation. S2-5: Check whether the generation of the current subsequence is completed Let t = t + 1. If t > L, it means that the samples of this sequence have been completely generated, and go to S2-8; otherwise, return to S2-4. S2-6: Check whether samples have been completely generated for all subsequences If the number of generated subsequences is less than 12 in the case of a wind farm or less than 24 in the case of a PV farm, let G * (1) = G * (t - 1), set L to the length of the next sequence, set the time stamp t to 1, and return to S2-4; if new normalized output subsequences have been generated for all sub-output sequences of the existing new energy power stations, then perform S2-7; S2-7: Stitch the newly generated normalized output subsequences Stitch all the above-generated normalized output subsequences in the time sequence order of month, morning / afternoon to obtain a set of simulated output data of the surrounding power stations with a length of 8760; if more annual output sequences need to be generated, return to S2-3 to realize the generation of the annual output sequences of multiple surrounding power stations of the current new energy power station.
4. The method according to claim 3, wherein In step S3, the process of establishing the operation state time series of all power stations includes: Step ①: Obtain the relevant data of the historical operation states of the established power stations used to generate multiple surrounding power stations, and statistically obtain the annual average maintenance and fault outage hours as h1 and h2 respectively; obtain its failure rate and repair rate through formulas (47) and (48): Where λ and μ represent the failure rate and repair rate of the station, respectively; Step ②: Initialize the time scale t = 1, and initialize the initial state of the new energy station. Assume that the initial state of the station in the annual operation state is the normal operation state, that is, s(t) = 1; Step ③: Sample the state duration; If the new energy station is currently in the normal operation state, that is, s(t) = 1, then the number of hours of continuous normal operation is obtained through Equation (49); if the new energy station is currently in the outage state, that is, s(t) = 0, then the number of hours of continuous outage state is obtained through Equation (50): Where T represents the number of hours of state duration, round() represents the operation of rounding the value in the parentheses to the nearest integer, and R is a uniformly distributed random number in the range of [0, 1]; Step ④: Update the operation state sequence according to the sampling result; If s(t) = 1, then set all operation states from time t to time t + T - 1 to 1, that is, s(t:t + T - 1) = 1, and change the state at the next moment s(t + T) = 0; if s(t) = 0, then set all operation states from time t to time t + T - 1 to 0, that is, s(t:t + T - 1) = 0, and change the state at the next moment s(t + T) = 1; Step ⑤: Update the time scale and judge the end condition of the step; When updating the time stamp \(t = t+T\), if the current \(t\geq8760\), the generated status sequence has reached one-year data. Let the annual operation status sequence \(S\) of the new energy power station OS \(=s(1:8760)\), and the generation of its time-series operation status sequence is completed; otherwise, jump back to step ③.
5. The method according to claim 4, wherein In Step ③, the state duration sampling method is used for sampling.
6. The method according to claim 5, wherein In step ③, for the system component outage model used for sampling, a model that separates forced outages and planned outages is applied, and the total outage probability U t is as follows: where, in the formula, λ p and λ are the transition rates of the planned and forced outage states respectively; μ p and μ are the repair rates for repairing from the planned and forced outage states respectively.
7. The method according to claim 5, characterized in that, In step ③, for the outage model of system components used for sampling, a combined model of forced and scheduled outages is applied, and the total outage probability U t is as follows: where, in the formula, λ p and λ are the transition rates of the planned and forced outage states respectively; μ p and μ are the repair rates for repair from the planned and forced outage states respectively.
8. The method according to claim 6 or 7, characterized in that, In step S4, for a certain port on the power source side of the power transmission channel, obtain the normalized sequence representation of the output of the m-th new energy power station as G * m , and calculate the actual normalized output sequence G considering faults and maintenance * m,act as follows: G * m,act (t) = G * m (t)S OS (t) The output of the new energy station cluster at this port can be calculated by the following formula: Where, G Terminal (t) is the aggregated output of the new energy power stations in the cluster at the port at time t, G * m (t) is the normalized output of the m-th new energy power station at time t, G m,rated is the installed capacity of the m-th new energy power station.