Method and device for generating output joint scene of multiple new energy bases
By generating a combined output scenario of multiple new energy bases using the SARIMA-Copula model and K-means clustering, the problem of inaccurate probability distribution characterization in traditional methods is solved, thereby improving the reliability of power system optimization decisions and the capacity for new energy absorption.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAZHONG UNIV OF SCI & TECH
- Filing Date
- 2026-01-23
- Publication Date
- 2026-05-08
AI Technical Summary
Traditional single-scenario generation methods are unable to accurately depict the joint probability distribution of multiple new energy bases, leading to a decrease in the reliability of power system optimization decisions.
A combined approach using the seasonal difference autoregressive moving average (SARIMA) model and the Copula model was employed to generate joint power output scenarios for multiple new energy bases. Power output was predicted using the SARIMA model, and a multivariate frequency histogram was plotted using the marginal probability distribution function. The optimal Copula function was selected to construct the SARIMA-Copula model, which was then simplified using K-means clustering to obtain a set of typical scenarios.
It improved the simulation accuracy of scenarios involving the combined output of multiple new energy bases, optimized the new energy absorption capacity and grid expansion plan, and enhanced the robustness of dispatch decisions.
Smart Images

Figure CN122000878A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of electrical engineering technology, and more specifically, relates to a method and apparatus for generating a scenario of combined power output from multiple new energy bases. Background Technology
[0002] As the global energy structure accelerates its transformation towards low-carbon and clean energy, new energy power generation technologies, represented by wind and solar photovoltaic power, are developing rapidly. Their installed capacity and power generation share in the power system are continuously and rapidly increasing, gradually becoming one of the main power sources. However, new energy power generation is highly dependent on meteorological resources, and its output exhibits significant intermittency, strong fluctuations, and inherent uncertainties. These inherent characteristics fundamentally change the traditional "source follows load" operation mode of the power system, bringing severe challenges to the system's long-term planning, medium-term dispatch, and real-time operation across all dimensions.
[0003] Especially in the context of regional power grid interconnection and inter-provincial power transmission, multiple new energy bases often need to operate jointly to achieve resource complementarity and optimized allocation. In this context, the power output characteristics of bases in different geographical locations are not independent but exhibit complex spatiotemporal correlations. For example, affected by the same large weather system, wind farms within a radius of hundreds of kilometers may simultaneously experience power output peaks or troughs, forming a positive spatial correlation and exacerbating power fluctuations across the entire grid. Meanwhile, wind and solar resources in different climate zones may exhibit temporal complementarity or negative correlation, presenting potential opportunities and coordination challenges for system balancing. This spatiotemporal correlation structure makes the joint power output behavior of multiple new energy bases far more complex than the superposition of individual sites, with its probability distribution exhibiting high-dimensional, nonlinear coupling characteristics.
[0004] However, traditional single-scenario generation methods are unable to accurately characterize their joint probability distribution, leading to a decrease in the reliability of power system optimization decisions. Summary of the Invention
[0005] In view of the above-mentioned defects or improvement needs of the existing technology, the present invention provides a method and apparatus for generating a joint scenario of multiple new energy bases. Its purpose is to solve the technical problem that the traditional single scenario generation method is difficult to accurately characterize its joint probability distribution, which leads to a decrease in the reliability of power system optimization decision-making.
[0006] To achieve the above objectives, according to one aspect of the present invention, a method for generating a multi-new energy base power output joint scenario is provided, comprising:
[0007] S1: Input the historical power output data of multiple new energy bases into the constructed seasonal differential autoregressive moving average (SARIMA) model so that the SARIMA model outputs power output prediction data. S2: Use the marginal probability distribution function corresponding to the power output prediction data to draw a multivariate frequency histogram; S3: Select the optimal Copula function based on the tail characteristics of the multivariate frequency histogram to construct the SARIMA-Copula model; S4: Use the SARIMA-Copula model to generate the random point output sequence of the multiple new energy bases at each time step, and perform inverse operation on the edge distribution function corresponding to the random point output sequence to obtain the joint output scenario set of the multiple new energy bases; S5: The set of power output joint scenarios of the multiple new energy bases is simplified to obtain a set of typical scenarios.
[0008] Further, the step of selecting the optimal Copula function based on the tail characteristics of the multivariate frequency histogram to construct the SARIMA-Copula model includes: If the tail characteristics of the multivariate frequency histogram of the power output prediction data are symmetrical tails and asymptotically independent tails, then the Gaussian Copula function is selected as the optimal Copula function. If the tail characteristics of the multivariate frequency histogram of the power output prediction data are symmetrical tails and have asymptotic tail correlation characteristics, then the t-Copula function is selected as the optimal Copula function. If the tail characteristics of the multivariate frequency histogram of the power output prediction data are asymmetric tails and sensitive to either the upper tail or the lower tail, then the Gumbel Copula function is selected as the optimal Copula function. If the tail characteristics of the multivariate frequency histogram of the power output prediction data are symmetrical tails with moderate correlation between the upper and lower tails, then the Frank Copula function is selected as the optimal Copula function.
[0009] Further, in step S5: determine the cluster centers and number of clusters using the K-means clustering method; and use the K-means clustering method to simplify the power output joint scenario set of the multiple new energy bases to obtain the target scenario set.
[0010] Furthermore, the process of determining the cluster centers in the K-means clustering method is as follows: if the number of cluster centers N is set to be 1, the average value of the output of all scenarios is taken as the cluster center; if the number of cluster centers N is set to be 2, the two scenarios with the greatest Euclidean distance are selected as the cluster centers; if the number of cluster centers N is set to be 3, the average value of the output of all scenarios and the two scenarios with the greatest Euclidean distance are selected as the initial three cluster centers.
[0011] Furthermore, the process of determining the cluster centers in the K-means clustering method is as follows: if the number of cluster centers N is set to be greater than 3, first select the scene with the average output of all scenes and the two scenes with the farthest Euclidean distance as the initial three cluster centers; calculate the Euclidean distance between each scene and the multiple predetermined cluster centers, and select the scene with the farthest sum of Euclidean distances to the multiple predetermined cluster centers as the next cluster center, and so on, until N cluster centers are found.
[0012] Furthermore, the process of determining the number of clusters in the K-means clustering method is as follows: the optimal number of clusters is selected when the sum of squared cluster errors decreases significantly with the increase of the number of clusters.
[0013] Furthermore, the process of determining the number of clusters in the K-means clustering method is as follows: calculate the sum of squared errors within each cluster (SSE) index and the cluster analysis CH index for each number of clusters within the search range; select the number of clusters with the largest CH index near the inflection point where the line graph corresponding to the SSE index tends to flatten out, and take this as the optimal number of clusters.
[0014] According to another aspect of the present invention, a device for generating a multi-new energy base power output joint scenario is provided, comprising: The prediction module is used to input the historical output data of multiple new energy bases into the constructed seasonal differential autoregressive moving average (SARIMA) model so that the SARIMA model outputs output prediction data. The conversion module is used to draw a multivariate frequency histogram using the marginal probability distribution function corresponding to the output prediction data; A construction module is used to select the optimal Copula function based on the tail characteristics of the multivariate frequency histogram to construct the SARIMA-Copula model; The processing module is used to generate the random point output sequence of the multiple new energy bases at each time step using the SARIMA-Copula model, and to perform inverse operation on the edge distribution function corresponding to the random point output sequence to obtain the joint output scenario set of the multiple new energy bases. The simplification module is used to simplify the power output joint scenario set of the multiple new energy bases to obtain a typical scenario set.
[0015] According to another aspect of the present invention, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method for generating the multi-new energy base power output joint scenario.
[0016] According to another aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method for generating the multi-new energy base power output joint scenario.
[0017] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects: (1) This invention provides a method for generating joint output scenarios of multiple new energy bases. Historical output data from multiple new energy bases are input into a SARIMA model to obtain output prediction data, which can accurately capture the temporal characteristics of new energy output. Furthermore, by utilizing the SARIMA-Copula model to obtain a set of joint output scenarios for multiple new energy bases, the spatial correlation between the outputs of multiple bases can be described. Combined with the former, temporal-spatial joint modeling can be achieved, thereby obtaining more accurate joint output scenarios of multiple new energy bases. This invention is applicable to new energy power system planning, can improve the simulation accuracy of joint output from multiple bases and the capacity for new energy absorption, optimize energy storage configuration and grid expansion schemes, provide more accurate output predictions for new energy power generators, and enhance the robustness of dispatch decisions.
[0018] (2) This scheme sets up multiple methods for setting cluster centers, and avoids fluctuations in clustering results by improving the initial cluster centers.
[0019] (3) This scheme takes into account the sum of squared errors of clusters and the Calinski-Harabas index to determine the optimal number of clusters, which improves the accuracy compared with the traditional elbow method which determines the optimal number of clusters by visually observing the inflection point. Attached Figure Description
[0020] Figure 1 This is a flowchart of a method for generating a multi-new energy base power output joint scenario according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the edge cumulative distribution function provided in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the specific implementation process of the inverse operation provided in an embodiment of the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0022] Example 1 This embodiment provides a method for generating a scenario of combined output from multiple new energy bases, such as... Figure 1 As shown, it includes: S1-S5.
[0023] S1: Input the historical power output data of multiple new energy bases into the constructed seasonal differential autoregressive moving average (SARIMA) model so that the SARIMA model outputs power output prediction data.
[0024] Regarding S1, the basic information of the new energy bases is first determined, including the number, geographical location (latitude and longitude, altitude), spatial correlation between the new energy bases, and installed capacity of each new energy base. As an optional implementation method, historical power output data of each new energy base is obtained to determine the temporal correlation of power output. By inputting the historical power output data of multiple new energy bases into the constructed seasonal differential autoregressive moving average (SARIMA) model, power output prediction data can be obtained.
[0025] Regarding the Seasonal Auto-regressive Tntegrated Moving Average (SARIMA) model, the autoregressive moving average model combines the autoregressive and moving average models. When the difference in data fit is not significant, it simplifies the model and does not require the partial autocorrelation function or the correlation function to be truncated. The SARIMA model, on the other hand, considers the significant seasonal fluctuations in wind and solar power output over long time scales, effectively addressing the uncertainty of seasonal wind and solar power output and significantly improving the accuracy of power output data prediction.
[0026] The primary condition for using the Autoregressive Moving Average (ARMA) model is that the time series used for prediction must be a stationary stochastic process, meaning all sample points are randomly distributed around a certain horizontal line. Therefore, non-stationary time series must be transformed into stationary ones through differencing. Applying the ARMA model to the differencing-transformed series is conventionally called the ARIMA model. The time unit for seasonality models is the corresponding period T. The general form: (1) In the formula, B is the time shift operator, i.e., the difference operator; A stationary time series that has undergone stationary processing, i.e. ,in The predicted value of contribution to new energy sources It is a difference operator of order d, i.e. , For a seasonal D-order difference operator with period T, i.e. ; ; These are the parameters for the autoregressive model; This is the seasonal moving average coefficient; The moving average coefficient; Let be the seasonal moving average coefficient. (2) Given a specific historical power output time series In this case, based on the above SARIMA model, a large number of independently distributed new energy power output forecast samples for the next year are generated. Based on the predicted output of new energy sources, the empirical average value for predicted new energy output can be derived. and experience variance They are respectively: (3) In this embodiment, the predicted value of new energy output is used. For example, its formula is: (4) In the formula, For the parameters of the k-th order autoregressive model, ; The coefficients of the i-th order moving average are... ; It is a difference operator, acting on Delay it by j bits, that is d represents the degree of difference; To accommodate random errors that conform to a normal distribution, in addition to using the SARIMA model to predict new energy output data, the input prediction sequence is also differencing to ensure sequence stationarity.
[0027] The SARIMA model, based on the ARIMA model, further incorporates seasonal autoregressive coefficients, seasonal difference coefficients, and seasonal moving average coefficients. Its mathematical model is as follows: (5) In the formula, for seasonal autoregressive coefficient, ; for seasonal moving average coefficient, D represents the seasonality difference frequency, and T represents the seasonal cycle length. D is typically determined using empirical rules: a clear seasonal pattern but "stable" seasonal amplitude and no obvious seasonal trend → often... Seasonality shifts over time, with seasonal points rising / falling year by year → often... (2) First determine the seasonal AR / MA order, then determine the coefficients. Empirically, the seasonal component is rarely taken to be too large: [The text abruptly ends here, likely due to an incomplete sentence or a formatting error.] The coefficients are not "calculated manually," but are obtained through parameter estimation, a common method being maximum likelihood estimation.
[0028] Combining equations (1)-(5) above, a SARIMA prediction model is constructed. By inputting the historical power output data of each new energy base, the predicted power output data of each new energy base can be generated, effectively addressing the uncertainty of seasonal power output of wind and solar power.
[0029] S2: Use the marginal probability distribution function corresponding to the output prediction data to draw a multivariate frequency histogram.
[0030] Specifically, the predicted power output data of each new energy source obtained from S1 can be used to determine the marginal probability distribution function of the power output of each new energy base. Due to the randomness and uncertainty of new energy power output, its power output probability distribution does not belong to common normal distributions, t-distributions, etc. Therefore, the cumulative probability distribution function (CDF) of its power output can be approximately fitted based on the empirical distribution function (EDF).
[0031] Forecast data sample of new energy output Its empirical distribution function is defined as: (6) in, This is an indicator function.
[0032] By smoothing the EDF, the continuous probability density function (PDF) is obtained, and then integrated to obtain the CDF, as follows: (7) (8) In the formula, Here, h is the kernel function, and h is the kernel parameter, such as the default value for the Gaussian kernel. The CDF obtained by integration is the cumulative probability distribution function (CDF).
[0033] The marginal probability distribution function of the power output of each new energy base can be obtained through equations (6)-(8). Figure 2 As shown in (a), when there is a large amount of new energy output data, the empirical distribution function can be approximated as its cumulative distribution function. In fact, it is an approximation of the cumulative distribution function. Although there is an error in the initial fitting stage, the difference is very small. If the accuracy allows, the empirical distribution function can be approximately considered to be its cumulative distribution function. Figure 2 As shown in (b), the kernel distribution estimate almost overlaps with the empirical distribution function, indicating that the kernel distribution can fit the kernel CDF well and can be used for subsequent parameter estimation.
[0034] S3: Select the optimal Copula function based on the tail characteristics of the multivariate frequency histogram to construct the SARIMA-Copula model.
[0035] Specifically, let's first introduce the relevant concepts and theorems of Copula. A Copula function is a function that connects the joint distribution function of multiple random variables with their respective marginal distribution functions; it can be understood as a connection function. The definition of a multivariate Copula function is as follows.
[0036] Definition 1 (Nelsen, 2006): An N-ary Copula function is a function that has the following properties. : a. Definition for ; b. It has a zero basal plane and is N-dimensionally increasing; c. marginal distribution And satisfy
[0037] in, Obviously, if the marginal distribution function Both are continuous univariate distribution functions, let ,but It is a subject to the edge The multivariate distribution function of a uniform distribution.
[0038] Sklar's theorem regarding multivariate distributions. Theorem 1: Let... For having a marginal distribution If the joint distribution function exists, then there must exist a Copula function. ,satisfy:
[0039] In the formula, Represents random variables The joint distribution function.
[0040] Regarding the construction method of the Copula model, based on the relevant definitions and theorems mentioned above, the steps for constructing a Copula function are as follows: (a) Determine the marginal distribution of the random variable; (b) Select an appropriate Copula function based on the correlation characteristics of the random variable. Commonly used Copula functions include the normal Copula function, t-Copula function, Gumbel Copula function, Clayton Copula function, and Frank Copula function. (c) Estimate the unknown parameters in the Copula model based on the selected Copula function. Several commonly used methods for estimating the correlation coefficient of the Copula function include the two-stage estimation method (IFM), the maximum likelihood estimation method based on the empirical distribution function (CML), and the maximum likelihood estimation method based on the nonparametric kernel density (MLK).
[0041] In the above steps, the marginal distribution in (a) is obtained by the method described in S2; the probability density expressions and tail characteristics of each type of Copula function in (b) are shown in Table 1; the estimation of unknown parameters in (c) adopts the CML method. Given a Copula function, its log-likelihood function is given by equation (9), and the parameter solution is given by equation (10).
[0042] (9) (10) Based on the above steps and solution methods, the copula function can be determined.
[0043] Table 1
[0044] As shown in the table, the specific method for determining the Copula function is as follows: Based on the marginal probability distribution function of the power output of each new energy base, draw a multivariate frequency histogram of its prediction data, and determine a suitable Copula function based on the tail characteristics of the histogram.
[0045] Specifically, if the tail characteristics of the multivariate frequency histogram of the predicted data are symmetrical tails with asymptotically independent tails, then the Gaussian Copula function is selected as the target Copula function; if the tail characteristics of the multivariate frequency histogram of the predicted data are symmetrical tails with asymptotically correlated tails, then the t-Copula function is selected as the target Copula function; if the tail characteristics of the multivariate frequency histogram of the predicted data are asymmetrical tails with sensitivity to the upper tail, then the Gumbel Copula function is selected as the target Copula function; if the tail characteristics of the multivariate frequency histogram of the predicted data are asymmetrical tails with sensitivity to the lower tail, then the Clayton Copula function is selected as the target Copula function; if the tail characteristics of the multivariate frequency histogram of the predicted data are symmetrical tails with moderate correlation between the upper and lower tails, then the Frank Copula function is selected as the target Copula function.
[0046] S4: Use the SARIMA-Copula model to generate the random point output sequence of the multiple new energy bases at each time step, and perform inverse operation on the edge distribution function corresponding to the random point output sequence to obtain the joint output scenario set of the multiple new energy bases.
[0047] Definition 2 Contextualization: Using a set of discrete probability distributions To approximate the continuous probability distribution function The process of this is called contextualization. For the quantile points of the scene, For the corresponding probability, This represents the total number of scenarios. The following describes the transformation relationship between Copula function scenarios and the original joint distribution function scenarios in S3: From Theorem 1, we know ,make The above equation becomes Discretized Copula functions are extensions of traditional continuous Copula functions for discrete random variables, used to describe the dependency structure between discrete variables. Suppose that the discretized Copula function yields the quantiles of a certain scenario. ,because Then it can be obtained through the inverse operation of the marginal distribution function. Obtain the scenario corresponding to the original joint distribution function .
[0048] Using a predetermined optimal Copula function, random points are generated probabilistically. If there are N new energy bases, the sequence of random points generated at time t is as follows: ,in The generated point's value represents the per-unit output of each new energy base at time t, allowing calculation of the actual output of each base. Since the points are generated randomly based on probability, the resulting large number of random points follow the joint output characteristics of these N new energy bases. It should be noted that, for ease of calculation, the generated random points are in the range [0, 1], corresponding to the per-unit output value, while the actual output equals the per-unit value. Rated capacity. For example, if the rated capacity is 100, and the calculated per-unit value is 0.82, then the actual output is 82. Here, actual output refers to the predicted actual output, not the actual output. The quantile for a given scenario is obtained from discretized Copula. ,because Then it can be obtained through the inverse operation of the marginal distribution function. Obtain the scenario corresponding to the original joint distribution function The specific process of the inverse operation is as follows: Figure 3As shown (taking wind power as an example). The data contained in parentheses () and curly braces ({}) are the same. {} represents the sequence of random points generated by each new energy base at time t. () represents the sequence of random points used for inverse operations.
[0049] S5: The set of power output joint scenarios of the multiple new energy bases is simplified to obtain a set of typical scenarios.
[0050] The traditional K-means clustering method for extracting typical scenarios follows these steps: 1. Data preprocessing: Screen core features, perform standardization / normalization, and remove outliers and missing values. 2. Parameter initialization: Determine the number of clusters K using experience or the elbow rule, and randomly select K samples as initial cluster centers. 3. Iterative sample allocation: Calculate the distance of each sample to the K cluster centers, and assign the samples to the nearest cluster, forming K clusters. 4. Update cluster center positions: Calculate the mean feature value of samples within each cluster, and use the mean point as the new cluster center. 5. Convergence output: Repeat the allocation and update steps until the cluster center change is less than a threshold or the maximum number of iterations is reached. The cluster center samples / average features represent the typical scenarios.
[0051] As an optional implementation, step S5 involves: determining the cluster centers and number of clusters using the K-means clustering method; and using the K-means clustering method to simplify the power output joint scenario set of the multiple new energy bases to obtain the target scenario set.
[0052] It should be noted that the initial cluster centers N are determined by the cluster centers K. You can define your own initial cluster centers, but they must be less than K to facilitate finding the final K cluster centers. The N described below is one that you determine beforehand. Then, depending on the different values of N, different methods are used to determine the initial cluster centers.
[0053] As an optional implementation, the process for determining cluster centers using the K-means clustering method is as follows: 1) When N=1, the average value of the output of all scenarios is used as the cluster center.
[0054] 2) When N=2, select the two scenes with the greatest Euclidean distance as cluster centers. The Euclidean distance is calculated as follows:
[0055] In the formula, Let be the Euclidean distance between scene i and scene j; These are elements of scene i and j at time t, respectively.
[0056] 3) When N≥3, first select the two scenes with the average output of all scenes and the longest Euclidean distance as the initial three clustering scenes. For the remaining scenes, calculate the Euclidean distance to each of the initial three clustering scenes, and select the scene with the longest sum of Euclidean distances to the three initial scenes as the fourth clustering scene. Then calculate the Euclidean distances between the remaining scenes and the selected four clustering scenes to obtain the fifth clustering scene, and so on, until N initial clustering scenes are found. This yields the initial cluster centers with a high degree of diversity.
[0057] The optimal number of clusters K is determined by comprehensively considering the sum of squared errors within clusters and the Calinski-Harabas index. An inappropriate selection of the number of clusters can lead to unrepresentative results. The traditional elbow method determines the optimal number of clusters by visually estimating inflection points, but this method has low accuracy. Therefore, relevant indices are introduced to assist in the determination.
[0058] As an optional implementation, the process for determining the number of clusters in the K-means clustering method is as follows: the optimal number of clusters is selected when the sum of squared cluster errors significantly decreases with increasing cluster size. The sum of squared cluster errors is defined as follows:
[0059] In the formula, For the first The value of scene n in the class at time t; For time t, the first The value of the cluster center; The number of clusters; denoted as the number of scenes in the k-th class.
[0060] As the number of clusters increases, the degree of cohesion within each cluster increases, and the SSE value gradually decreases. When the number of clusters is small, increasing the number of clusters will significantly increase the degree of cohesion within each cluster, and the SSE value will decrease significantly. When approaching the optimal number of clusters, the rate of decrease in the SSE value slows down, and this inflection point is the optimal number of clusters.
[0061] As an optional implementation method, visual estimation is commonly used to determine the inflection point, but its accuracy is low. Therefore, the clustering effectiveness index Calinski-Harabasz (CH) is introduced to help determine the optimal number of clusters. CH is calculated as follows:
[0062] In the formula, The trace of the intra-class scatter matrix measures intra-class compactness. The trace of the inter-class scatter matrix measures the inter-class separation. The number of clusters; This represents the total number of scenes.
[0063] In summary, the steps to determine the optimal number of clusters are as follows: calculate the SSE and CH indices for each number of clusters within the search range; find the number of clusters with the highest clustering effectiveness index (CH) near the inflection point where the SSE line graph tends to flatten out; this number is the optimal number of clusters. .
[0064] Example 2 This embodiment provides a device for generating a scenario of combined output from multiple new energy bases, including: a prediction module, a conversion module, a construction module, a processing module, and a simplification module.
[0065] The system comprises the following modules: a prediction module, which inputs historical power output data from multiple new energy bases into a constructed seasonal differential autoregressive moving average (SARIMA) model to output predicted power output data; a transformation module, which plots a multivariate frequency histogram using the marginal probability distribution function corresponding to the predicted power output data; a construction module, which selects the optimal Copula function based on the tail characteristics of the multivariate frequency histogram to construct a SARIMA-Copula model; a processing module, which uses the SARIMA-Copula model to generate random point power output sequences for each of the multiple new energy bases at each time step, performs inverse operations on the marginal distribution function corresponding to the random point power output sequences, and obtains a joint power output scenario set for the multiple new energy bases; and a reduction module, which reduces the joint power output scenario set for the multiple new energy bases to obtain a typical scenario set.
[0066] Example 3 This embodiment provides an electronic device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the method for generating the multi-new energy base power output joint scenario.
[0067] The electronic device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The memory can be used to store computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and by accessing data stored in the memory.
[0068] Example 4 This embodiment provides a computer-readable storage medium storing a computer program thereon, wherein when the computer program is executed by a processor, it implements the steps of the method for generating the multi-new energy base power output joint scenario.
[0069] Specifically, the memory may include high-speed random access memory, as well as non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart media cards (SMC), secure digital (SD) cards, flash cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.
[0070] Example 5 This invention provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of the method described in the above embodiments of this invention.
[0071] The technical features of the embodiments described above can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification. It should be noted that the terms "in one embodiment," "for example," and "again" in this invention are intended to illustrate the invention and are not intended to limit the invention.
[0072] The embodiments described above are merely examples of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.
Claims
1. A method for generating a scenario of combined power output from multiple new energy bases, characterized in that, include: S1: Input the historical power output data of multiple new energy bases into the constructed seasonal differential autoregressive moving average (SARIMA) model so that the SARIMA model outputs power output prediction data. S2: Use the marginal probability distribution function corresponding to the power output prediction data to draw a multivariate frequency histogram; S3: Select the optimal Copula function based on the tail characteristics of the multivariate frequency histogram to construct the SARIMA-Copula model; S4: Use the SARIMA-Copula model to generate the random point output sequence of the multiple new energy bases at each time step, and perform inverse operation on the edge distribution function corresponding to the random point output sequence to obtain the joint output scenario set of the multiple new energy bases; S5: The set of power output joint scenarios of the multiple new energy bases is simplified to obtain a set of typical scenarios.
2. The method for generating a multi-new energy base power output joint scenario as described in claim 1, characterized in that, The step of selecting the optimal Copula function based on the tail characteristics of the multivariate frequency histogram to construct the SARIMA-Copula model includes: If the tail characteristics of the multivariate frequency histogram of the power output prediction data are symmetrical tails and asymptotically independent tails, then the Gaussian Copula function is selected as the optimal Copula function. If the tail characteristics of the multivariate frequency histogram of the power output prediction data are symmetrical tails and have asymptotic tail correlation characteristics, then the t-Copula function is selected as the optimal Copula function. If the tail characteristics of the multivariate frequency histogram of the power output prediction data are asymmetric tails and sensitive to either the upper tail or the lower tail, then the Gumbel Copula function is selected as the optimal Copula function. If the tail characteristics of the multivariate frequency histogram of the power output prediction data are symmetrical tails with moderate correlation between the upper and lower tails, then the Frank Copula function is selected as the optimal Copula function.
3. The method for generating a multi-new energy base power output joint scenario as described in claim 1, characterized in that, S5: Determine the cluster centers and number of clusters using the K-means clustering method; use the K-means clustering method to simplify the output joint scenario set of the multiple new energy bases to obtain the target scenario set.
4. The method for generating a multi-new energy base power output joint scenario as described in claim 3, characterized in that, The process of determining cluster centers in the K-means clustering method is as follows: if the number of cluster centers N is set to be 1, the average value of the output of all scenes is used as the cluster center; if the number of cluster centers N is set to be 2, the two scenes with the greatest Euclidean distance are selected as cluster centers; if the number of cluster centers N is set to be 3, the average value of the output of all scenes and the two scenes with the greatest Euclidean distance are selected as the initial three cluster centers.
5. The method for generating a multi-new energy base power output joint scenario as described in claim 4, characterized in that, The process of determining cluster centers in the K-means clustering method is as follows: If the number of cluster centers N is set to be greater than 3, first select the scene with the average output of all scenes and the two scenes with the farthest Euclidean distance as the initial three cluster centers; calculate the Euclidean distance between each scene and the multiple predetermined cluster centers, and select the scene with the farthest sum of Euclidean distances to the multiple predetermined cluster centers as the next cluster center, and so on, until N cluster centers are found.
6. The method for generating a multi-new energy base power output joint scenario as described in claim 3, characterized in that, The process for determining the number of clusters in the K-means clustering method is as follows: the optimal number of clusters is selected when the sum of squared cluster errors decreases significantly with the increase of the number of clusters.
7. The method for generating a multi-new energy base power output joint scenario as described in claim 3, characterized in that, The process of determining the number of clusters in the K-means clustering method is as follows: calculate the sum of squared errors within each cluster (SSE) and the cluster analysis CH index for each number of clusters within the search range; select the number of clusters with the largest CH index near the inflection point where the line graph corresponding to the SSE index tends to flatten out, and take this as the optimal number of clusters.
8. A device for generating a scenario of combined power output from multiple new energy bases, characterized in that, include: The prediction module is used to input the historical output data of multiple new energy bases into the constructed seasonal differential autoregressive moving average (SARIMA) model so that the SARIMA model outputs output prediction data. The conversion module is used to draw a multivariate frequency histogram using the marginal probability distribution function corresponding to the output prediction data; A construction module is used to select the optimal Copula function based on the tail characteristics of the multivariate frequency histogram to construct the SARIMA-Copula model; The processing module is used to generate the random point output sequence of the multiple new energy bases at each time step using the SARIMA-Copula model, and to perform inverse operation on the edge distribution function corresponding to the random point output sequence to obtain the joint output scenario set of the multiple new energy bases. The simplification module is used to simplify the power output joint scenario set of the multiple new energy bases to obtain a typical scenario set.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.