Data simulation method and device of wind power plant, electronic equipment and storage medium
By constructing a joint distribution function of the marginal distribution model and the connection function model, the problem of missing data in wind farms is solved, and the simulation accuracy and prediction effect of wind power data are improved.
Patent Information
- Application Number
- CN202511019649.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-11-21
AI Technical Summary
Data gaps in wind farms due to human or environmental factors increase the difficulty of data management and analysis, affecting the accurate prediction of wind power output.
By constructing a target edge distribution model and a target connection function model, a target joint distribution function is obtained by fusing them, and wind power data is simulated to fill in the missing data.
It effectively characterizes the randomness and correlation of wind speed and active power, improves the accuracy of data simulation, and reduces the impact of missing data.
Smart Images

Figure CN120994979A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of data simulation, and in particular to a data simulation method and device for a wind farm, an electronic device, and a storage medium. BACKGROUND
[0002] Wind power data of each wind turbine in a wind farm is indispensable for analyzing fluctuation characteristics, power prediction, and optimal scheduling. In addition, the completeness of wind power data is of great significance for accurate prediction of wind power and utilization of wind energy. However, in the actual data collection process, data loss often occurs due to human or environmental factors, which increases the difficulty for researchers to manage and analyze data, and also causes deviation of analysis results, thereby adversely affecting the accurate prediction of wind power. Therefore, how to fill in the data to eliminate or reduce the impact of missing data has become a problem to be solved. SUMMARY
[0003] To solve the above technical problems or at least partially solve the above technical problems, embodiments of the present application provide a data simulation method and device for a wind farm, an electronic device, and a storage medium, to solve the problem of how to fill in the data to eliminate or reduce the impact of missing data.
[0004] To achieve the above purpose, the technical solutions provided by the embodiments of the present application are as follows:
[0005] In a first aspect, the present application provides a data simulation method for a wind farm, which comprises: obtaining measured wind power data of wind turbines in the wind farm, wherein the measured wind power data comprises wind speed and active power;
[0006] According to the measured wind power data, a target edge distribution model and a target connection function model are constructed, wherein the target edge distribution model is used to represent the randomness between the measured wind power data, and the target connection function model is used to represent the correlation between the measured wind power data;
[0007] The target edge distribution model and the target connection function model are fused to obtain a target joint distribution function;
[0008] The measured wind power data is simulated by using the target joint distribution function to obtain target simulation sample data.
[0009] As an optional implementation, in the first aspect of the embodiments of the present application, the target edge distribution model is constructed according to the measured wind power data, which comprises:
[0010] According to the measured wind power data, a plurality of initial edge distribution models are constructed, the initial edge distribution models comprising at least two of a truncated normal distribution model, a lognormal distribution model, a truncated extreme value distribution model and a Weibull distribution model;
[0011] According to a goodness-of-fit test method, the target edge distribution model is selected from the plurality of initial edge distribution models.
[0012] As an optional implementation, in the first aspect of the embodiment of the present application, the selecting the target edge distribution model from the plurality of initial edge distribution models according to the goodness-of-fit test method comprises:
[0013] According to the goodness-of-fit test method, a goodness-of-fit test value corresponding to each initial edge distribution model is calculated respectively;
[0014] The initial edge distribution model with the minimum goodness-of-fit test value is determined as the target edge distribution model.
[0015] As an optional implementation, in the first aspect of the embodiment of the present application, the constructing the target connection function model according to the measured wind power data comprises:
[0016] According to the measured wind power data, a plurality of initial connection function models are constructed, the initial connection function models comprising at least two of a Gaussian connection function model, a Plackett connection function model and an Archimedean connection function model;
[0017] According to a goodness-of-fit test method, the target connection function model is selected from the plurality of initial connection function models.
[0018] As an optional implementation, in the first aspect of the embodiment of the present application, the selecting the target connection function model from the plurality of initial connection function models according to the goodness-of-fit test method comprises:
[0019] According to the goodness-of-fit test method, a goodness-of-fit test value corresponding to each initial connection function model is calculated respectively;
[0020] The initial connection function model with the minimum goodness-of-fit test value is determined as the target connection function model.
[0021] As an optional implementation, in the first aspect of the embodiment of the present application, the simulating the measured wind power data by using the target joint distribution function to obtain target simulated sample data comprises:
[0022] The measured wind power data is simulated by using the target connection function model to obtain initial simulated sample data;
[0023] The initial simulation sample data is converted by an inverse function of the target edge distribution model to obtain the target simulation sample data.
[0024] In a second aspect, an embodiment of the present application provides a data simulation device for a wind farm, the data simulation device comprising: an acquisition module configured to acquire measured wind power data of wind turbines in the wind farm, the measured wind power data comprising: wind speed and active power;
[0025] a processing module configured to construct a target edge distribution model and a target connection function model according to the measured wind power data, the target edge distribution model being configured to represent randomness between the measured wind power data, and the target connection function model being configured to represent correlation between the measured wind power data;
[0026] The processing module is further configured to fuse the target edge distribution model and the target connection function model to obtain a target joint distribution function.
[0027] The processing module is further configured to simulate the measured wind power data by using the target joint distribution function to obtain target simulation sample data.
[0028] As an optional implementation, in the second aspect of the embodiment of the present application, the processing module is specifically configured to construct a plurality of initial edge distribution models according to the measured wind power data, the initial edge distribution models comprising at least two of the following: a truncated normal distribution model, a lognormal distribution model, a truncated extreme value distribution model, and a Weibull distribution model.
[0029] The processing module is specifically configured to select the target edge distribution model from the plurality of initial edge distribution models according to a goodness-of-fit test method.
[0030] As an optional implementation, in the second aspect of the embodiment of the present application, the processing module is specifically configured to calculate a pseudo test value corresponding to each initial edge distribution model according to the goodness-of-fit test method.
[0031] The processing module is specifically configured to determine the initial edge distribution model with the minimum pseudo test value as the target edge distribution model.
[0032] As an optional implementation, in the second aspect of the embodiment of the present application, the processing module is specifically configured to construct a plurality of initial connection function models according to the measured wind power data, the initial connection function models comprising at least two of the following: a Gaussian connection function model, a Plackett connection function model, and an Archimedean connection function model.
[0033] The processing module is specifically configured to select the target connection function model from the plurality of initial connection function models according to a goodness-of-fit test method.
[0034] As an optional implementation, in the second aspect of the embodiment of the present application, the processing module is specifically configured to calculate a quasi-test value corresponding to each initial connection function model respectively according to the goodness-of-fit test method.
[0035] The processing module is specifically configured to determine the initial connection function model with the minimum quasi-test value as the target connection function model.
[0036] As an optional implementation, in the second aspect of the embodiment of the present application, the processing module is specifically configured to simulate the actual wind power data by using the target connection function model to obtain initial simulation sample data.
[0037] The processing module is specifically configured to convert the initial simulation sample data by using an inverse function of the target edge distribution model to obtain the target simulation sample data.
[0038] In a third aspect, an electronic device is provided, and the electronic device comprises:
[0039] a memory storing executable program codes;
[0040] a processor coupled with the memory;
[0041] The processor invokes the executable program codes stored in the memory to execute the data simulation method for a wind farm in the first aspect of the embodiment of the present application.
[0042] In a fourth aspect, a computer readable storage medium storing a computer program is provided, and the computer program causes a computer to execute the data simulation method for a wind farm in the first aspect of the embodiment of the present application. The computer readable storage medium comprises ROM / RAM, a magnetic disk or an optical disk, etc.
[0043] In a fifth aspect, a computer program product is provided, and when the computer program product runs on a computer, the computer program product causes the computer to execute part or all steps of any one method in the first aspect.
[0044] In a sixth aspect, an application publishing platform is provided, and the application publishing platform is used to publish a computer program product, wherein when the computer program product runs on a computer, the computer program product causes the computer to execute part or all steps of any one method in the first aspect.
[0045] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0046] The embodiments of the present application provide a wind farm data simulation method and device, electronic equipment and storage medium. The measured wind power data of a wind turbine in a wind farm is obtained, the measured wind power data including wind speed and active power. A target edge distribution model and a target connection function model are constructed according to the measured wind power data. The target edge distribution model is used to represent the randomness between the measured wind power data, and the target connection function model is used to represent the correlation between the measured wind power data. The target edge distribution model and the target connection function model are fused to obtain a target joint distribution function. The measured wind power data is simulated through the target joint distribution function to obtain target simulation sample data. In this scheme, the edge distribution function can effectively represent the randomness characteristics of the active power and the wind speed, and the influence factors that cannot be predicted are considered from the perspective of probability theory. Since the relationship between the wind speed and the active power is not a simple linear relationship, the connection function can also effectively represent the correlation characteristics between the wind speed and the active power. Therefore, the combination of the edge distribution function and the connection function can more accurately and effectively realize the simulation and prediction of the missing data. BRIEF DESCRIPTION OF DRAWINGS
[0047] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present application and, together with the specification, serve to explain the principles of the application.
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0049] Figure 1 is a flowchart of a wind farm data simulation method provided by the embodiments of the present application Figure 1 ;
[0050] Figure 2 is a flowchart of a wind farm data simulation method provided by the embodiments of the present application Figure 2 ;
[0051] Figure 3 is an active power curve and frequency distribution histogram provided by the embodiments of the present application;
[0052] Figure 4 is a wind speed curve and frequency distribution histogram provided by the embodiments of the present application;
[0053] Figure 5 is a scatter plot of initial simulation sample data provided by the embodiments of the present application;
[0054] Figure 6 is a scatter plot corresponding to target simulation sample data provided by an embodiment of the present application;
[0055] Figure 7 is a structural schematic diagram of a data simulation device for a wind farm provided by an embodiment of the present application;
[0056] Figure 8 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0057] In order to more clearly understand the above objectives, features and advantages of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. It should be explained that, in the case of no conflict, the embodiments of the present application and the features in the embodiments can be combined with each other. Obviously, the described embodiments are some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0058] The terms "first" and "second" and the like in the specification and claims of the present application are used to distinguish different objects, and are not used to describe a specific order of the objects.
[0059] The terms "include" and "have" and any variations thereof in the embodiments of the present application are intended to cover the inclusions not exclusively, for example, the processes, methods, systems, products or devices containing a series of steps or units do not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0060] It should be explained that, in the embodiments of the present application, the words "exemplary" or "for example" are used to represent as an example, illustration or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the words "exemplary" or "for example" are intended to present the relevant concept in a specific manner.
[0061] The historical data of wind power is indispensable for analyzing the fluctuation characteristics of wind power, power prediction, and optimal scheduling. The completeness of wind power data is of great significance for accurate prediction of wind power and utilization of wind energy. However, due to human or physical factors, data loss often occurs, which increases the difficulty of data management and analysis for researchers, causes deviation of analysis results, and adversely affects the accurate prediction of wind power. Therefore, it is increasingly important to eliminate or reduce the impact of these missing data through data completion.
[0062] At present, the main methods for filling missing wind power data in China include the continuous method, the average interpolation method, and the artificial neural network (ANN). Due to the randomness and uncertainty of wind power, it is often treated as a random variable in research, ignoring the correlation factors of wind power.
[0063] To solve the above technical problems or all technical problems, the embodiments of the present application provide a data simulation method and device for a wind farm, an electronic device, and a storage medium. The method includes obtaining measured wind power data of wind turbines in the wind farm, the measured wind power data including wind speed and active power; constructing a target edge distribution model and a target connection function model according to the measured wind power data, the target edge distribution model being used to represent the randomness between the measured wind power data, and the target connection function model being used to represent the correlation between the measured wind power data; fusing the target edge distribution model and the target connection function model to obtain a target joint distribution function; and simulating the measured wind power data through the target joint distribution function to obtain target simulation sample data. In this scheme, the edge distribution function can effectively represent the randomness characteristics of the active power and the wind speed, considering the unpredictable influencing factors from the perspective of probability theory. Since the relationship between the wind speed and the active power is not simply linear, the connection function can also effectively represent the correlation characteristics between the wind speed and the active power. Therefore, combining the edge distribution function and the connection function to construct a joint distribution model can more accurately and effectively simulate and predict missing data.
[0064] As shown in Figure 1 , a flowchart of a data simulation method for a wind farm is provided in the embodiments of the present application. The method can include the following steps: Figure 1
[0065] 101. Obtain measured wind power data of wind turbines in the wind farm.
[0066] In the embodiments of the present application, the wind turbines of the wind farm are the core equipment of the wind power generation system, which converts wind energy into electric energy and is an important part of modern clean energy. The measured wind power data can specifically include wind speed and active power. The wind speed can be directly determined by an anemometer, and the active power can be obtained by a motor.
[0067] In some embodiments, in order to guarantee the accuracy of the data, after obtaining the measured wind power data, the measured wind power data can also be preprocessed, such as calculating the average value and the standard deviation of the data, and performing data cleaning and denoising.
[0068] 102. Constructing a target edge distribution model and a target connection function model according to the measured wind power data.
[0069] In the embodiments of the present application, since the wind speed and the wind power have randomness and uncertainty, and also have certain correlation, the above characteristics can be described by the edge distribution model and the connection function model respectively, wherein the target edge distribution model can be used to represent the randomness between the measured wind power data, and the target connection function model can be used to represent the correlation between the measured wind power data.
[0070] It should be noted that the edge distribution function (Marginal Distribution Function) is a function that describes the distribution characteristics of at least one random variable in the joint distribution function of multi-dimensional random variables. In the multi-dimensional probability space, the joint distribution function describes the probability law of all random variables taking values at the same time, and the edge distribution function extracts the distribution information of a specific random variable by "marginalizing" other variables. For discrete random variables, the edge distribution function is obtained by summing the edge probability mass function (PMF); for continuous random variables, the edge distribution function is obtained by integrating the edge probability density function (PDF).
[0071] It should be noted that the connection function can include a Copula function, and the Copula function involves the Copula theory.
[0072] Specifically, the Copula theory is an important tool for describing the dependence structure of variables. It decomposes the joint distribution of multi-dimensional random variables into two parts, the edge distribution and the Copula function, so that researchers can model the marginal behavior of variables and their correlation independently. The Copula function itself is a multivariate distribution function, and its edge distribution is uniform distribution, which can flexibly capture various complex dependence relationships, including nonlinearity, asymmetry and tail correlation. That is, the Copula function is a connection function that combines the multi-dimensional joint distribution with the corresponding one-dimensional edge distribution. The function can be imagined as a multi-dimensional cumulative distribution function on a unit cube with uniform edge distribution, and its density reflects the strength of the correlation between the edge distributions.
[0073] In some embodiments, the Copula function is invariant to monotonic transformations of the marginal distributions, which means that the dependence structure between variables is always described by the Copula function regardless of how the marginal distributions change. Unlike traditional linear correlation coefficients (such as Pearson correlation coefficient), the Copula function is able to capture the dependence relationship of variables in the tail of the distribution. For example, the Gumbel Copula is good at describing upper tail dependence, while the Clayton Copula is more sensitive to lower tail dependence. In addition, the Copula function can describe the asymmetric dependence relationship between variables, while the linear correlation coefficient assumes that the relationship between variables is symmetric.
[0074] 103. Fusing the target marginal distribution model and the target copula function model to obtain a target joint distribution function.
[0075] In the embodiments of the present application, after obtaining the target marginal distribution model and the target copula function model respectively, the target marginal distribution model and the target copula function model can be fused to obtain the final target joint distribution function based on the Copula theory. The target joint distribution function can simulate the measured wind power data from two aspects of the marginal distribution and the Copula function.
[0076] In some embodiments, the formula of the target joint distribution function is:
[0077] F(P, V) = C(F1(P), F2(V), θ) = C(u, v; P)
[0078] f(P, V) = f1(P)f2(V)D(F1(P), F2(V), θ)
[0079]
[0080] Where P is the active power, V is the speed, θ is the Copula function parameter, τ is the Kendall correlation coefficient, ρ is the Pearson correlation coefficient, C is the Copula function describing the correlation structure between variables, F1(P), F2(V), f1(P), f2(V) are the marginal distribution functions and marginal probability density functions of P and V respectively, and D(.) is the density function corresponding to the Copula function.
[0081] 104. Simulating the measured wind power data by using the target joint distribution function to obtain target simulation sample data.
[0082] In this embodiment of the application, since there may be missing data in the measured wind power data and the amount of data is small, in order to better analyze the wind power characteristics through a large amount of data, the measured wind power data can be simulated by the target joint distribution function to obtain a larger number of target simulation sample data. The target simulation sample data can be regarded as actual supplementary data based on the measured wind power data, thereby enriching the amount of wind power data.
[0083] This application provides a data simulation method for wind farms. The marginal distribution function can effectively characterize the randomness of active power and wind speed, and consider unpredictable influencing factors from a probabilistic perspective. Since the relationship between wind speed and active power is not a simple linear one, the connection function can also effectively characterize the correlation between wind speed and active power. Therefore, by combining the marginal distribution function and the connection function to construct a joint distribution model, the simulation and prediction of missing data can be achieved more accurately and effectively.
[0084] like Figure 2 As shown, Figure 2 A flowchart of a wind farm data simulation method provided in this application embodiment, the method may further include the following steps:
[0085] 201. Obtain measured wind power data of wind turbines in the wind farm.
[0086] In this embodiment, the description of step 201 is the same as the detailed description of step 101 in the above embodiments, and will not be repeated in this embodiment.
[0087] 202. Based on measured wind power data, construct multiple initial edge distribution models.
[0088] In this embodiment of the application, the target marginal distribution model can be considered as the optimal marginal distribution model, that is, the optimal model needs to be selected from multiple initial marginal distribution models. The marginal distribution model can include a variety of functions, specifically including at least two of the following: truncated normal distribution model, log-normal distribution model, truncated extreme value distribution model and Weibull distribution model.
[0089] The truncated normal distribution model is a statistical model that restricts the normal distribution. Its core idea is to cut off the portion of the normal distribution that does not meet specific conditions, thereby better reflecting the distribution characteristics of actual data. The truncated normal distribution refers to the truncation of the standard normal distribution or general normal distribution, retaining only the portion that meets specific conditions. The truncation can be one-sided (e.g., retaining only values greater than a certain threshold) or two-sided (e.g., retaining only values within a certain interval).
[0090] In some embodiments, the probability density function and cumulative distribution function of the truncated normal distribution model are as follows:
[0091]
[0092] p = μ
[0093] q = σ
[0094] where, for the log-normal distribution model, the log-normal distribution is a continuous probability distribution characterized by the fact that the logarithm of a random variable follows a normal distribution. If a random variable X follows a log-normal distribution, then ln(X) follows a normal distribution N(μ, σ2). The log-normal distribution is naturally skewed to the right, making it suitable for describing non-negative and right-skewed data. The product of multiple independent log-normal distribution variables still follows a log-normal distribution. The log-normal distribution is a generalization of the normal distribution under exponential transformation.
[0095] In some embodiments, the probability density function and cumulative distribution function of the log-normal distribution model are as follows:
[0096]
[0097] where, the truncated extreme value distribution model combines the theories of truncated distribution and extreme value distribution, aiming to handle extreme value data truncated within a certain range, to more accurately model and predict extreme events.
[0098] Truncated distribution refers to the range of random variable values being limited within a certain interval, and values exceeding the interval are truncated or ignored. Truncated distribution includes: lower truncation when variable values less than a certain threshold are truncated; upper truncation when variable values greater than a certain threshold are truncated; double-sided truncation when variable values between two thresholds are truncated.
[0099] Extreme value distribution is used to describe the asymptotic distribution of the maximum or minimum of a set of independent and identically distributed random variables. Extreme value distribution includes: Gumbel distribution (Type I) for light-tailed distribution; Frechet distribution (Type II) for heavy-tailed distribution; Weibull distribution (Type III) for distribution with finite upper or lower bounds, which is more suitable for truncated extreme value Type I distribution in the embodiments of the present application. In the context of truncated data, the extreme value distribution model is adjusted to accommodate the characteristics of data truncation, thereby more accurately estimating the probability distribution of extreme values.
[0100] In some embodiments, the probability density function and cumulative distribution function of the truncated extreme value distribution model are as follows:
[0101]
[0102] wherein the Weibull distribution model is a continuous probability distribution, the Weibull distribution can be fitted to different types of data by adjusting the parameters k and λ, the parameters k and λ can be estimated by maximizing the likelihood function, or the parameters can be estimated by minimizing the sum of squared residuals after linearizing the Weibull distribution.
[0103] In some embodiments, the probability density function and the cumulative distribution function of the Weibull distribution model are as follows:
[0104]
[0105] In some embodiments, through the above four edge distribution models, the active power PDF curve and the frequency distribution histogram as shown in Figure 3 , and the wind speed curve and the frequency distribution histogram as shown in Figure 4 , and Figure 3 , and Figure 4 all show the distribution of each function and the frequency histogram.
[0106] 203. According to the goodness-of-fit test method, the target edge distribution model is selected from the plurality of initial edge distribution models.
[0107] In the embodiments of the present application, in order to select the optimal edge distribution model from the above-mentioned plurality of initial edge distribution models, the goodness-of-fit test method can be used. The goodness-of-fit test (Goodness-of-Fit Test) is a statistical method for evaluating the degree of fitting between observed data and theoretical distribution or model. It helps to judge whether the observed data conforms to a certain assumed distribution, or whether the model can accurately describe the characteristics of the data.
[0108] In some embodiments, the goodness-of-fit test method can specifically include Chi-Square Goodness-of-Fit Test, Kolmogorov-Smirnov (K-S) test, Anderson-Darling (A-D) test, Cramér-von Mises test, AIC (Akaike Information Criterion) and BIC (Bayesian Information Criterion) and the like. Each initial edge distribution model can be calculated according to the above method, so as to obtain the optimal target edge distribution model.
[0109] In some embodiments, according to the goodness-of-fit test method, the target edge distribution model is selected from the plurality of initial edge distribution models, which can specifically include: according to the goodness-of-fit test method, calculating the chi-square test value corresponding to each initial edge distribution model respectively; determining the initial edge distribution model with the smallest chi-square test value as the target edge distribution model.
[0110] It should be noted that in order to accurately select the target edge distribution model, specific test values can be used to describe the characteristics of the edge distribution model, so as to directly select through numerical comparison.
[0111] In the embodiments of the present application, AIC criterion and BIC criterion can be used for selection.
[0112] In some embodiments, AIC and BIC can be used to compare the goodness of fit of different models, and select the optimal model. AIC and BIC balance the goodness of fit and complexity of the model to evaluate the model.
[0113] Specifically, the AIC criterion formula is:
[0114]
[0115] The BIC criterion formula is:
[0116]
[0117] Where (X i ; p, q) is the maximum likelihood function of the model, that is, the probability density function of the edge distribution function, X i is the measured wind power data, p and q are distribution parameters, in the embodiments of the present application, p and q are active power and wind speed, and k is the number of parameters of the model.
[0118] In the embodiments of the present application, the initial edge distribution model with the minimum AIC value and BIC value can be determined as the target edge distribution model.
[0119] 204、According to the measured wind power data, a plurality of initial connection function models are constructed.
[0120] In the embodiments of the present application, the target edge connection function model can be considered as the optimal connection function model, that is, the optimal model is selected from a plurality of initial connection function models. The connection function model can include a plurality of functions, specifically including at least two of the Gaussian connection function model, the Plackett connection function model and the Archimedean connection function model.
[0121] In some embodiments, the Gaussian Copula is a kind of connection function (Copula) based on multivariate normal distribution, which is used to describe the dependence structure between multiple random variables. The core idea is to convert the marginal distribution of the variable into a uniform distribution, and then use the multivariate normal distribution to capture the correlation between variables. Gaussian Copula assumes that the dependence relationship between variables is symmetric, that is, Corr(X, Y) = Corr(Y, X). Gaussian Copula has low tail correlation, that is, the probability of extreme events occurring simultaneously is underestimated. Since the CDF and inverse CDF of multivariate normal distribution have analytical expressions, the calculation of Gaussian Copula is relatively efficient.
[0122] wherein the Gaussian Copula function formula is:
[0123]
[0124] wherein f(·) and f -1 (·) are the standard normal distribution function and its inverse function, respectively; θ is the parameter of the Copula function, θ ∈ [-1, 1].
[0125] In some embodiments, the Plackett Copula is a kind of bivariate Copula function used to describe the dependence structure between random variables, which can flexibly depict the nonlinear correlation between variables, especially the symmetric or asymmetric tail correlation. Plackett Copula is symmetric in u and v, that is, C(u, v; θ) = C(v, u; θ). When θ > 1, the variable has positive tail correlation (i.e., the probability of extreme values occurring simultaneously is higher); when θ < 1, the variable has negative tail correlation (i.e., extreme values tend to appear alternately); when θ = 1, the variable is independent, without tail correlation.
[0126] wherein the Plackett Copula function formula is:
[0127]
[0128] wherein θ ∈ (0, ∞), representing the positive and negative correlation between parameters, and τ is the Kendall rank correlation coefficient.
[0129] In some embodiments, the Archimedean connection function can include a Frank Copula function, which is a probability distribution model describing the non-linear correlation between variables, is a binary joint distribution function based on Frank distribution, and has C(u, v) = C(v, u) for any u and v, that is, the function has symmetry. The marginal distribution of the Frank Copula function can be modeled by the Frank distribution, which is suitable for various variable types. It can also describe the non-linear correlation between variables, and the upper tail and lower tail correlation coefficients of the Frank Copula function are both zero, indicating that the variables are asymptotically independent in the tail of the distribution. Because the density distribution is in the shape of "U", it cannot capture the asymmetric correlation between variables.
[0130] wherein the Frank Copula function formula is:
[0131]
[0132]
[0133] wherein θ ∈ (-∞, +∞), and τ is the Kendall correlation coefficient.
[0134] 205. Selecting a target connection function model from the plurality of initial connection function models according to a goodness-of-fit test method.
[0135] In the embodiments of the present application, in order to select the optimal connection function model from the above-mentioned plurality of initial connection function models, a goodness-of-fit test method can be used. The specific description of the goodness-of-fit test method can refer to the description of step 203, and will not be repeated.
[0136] In some embodiments, the target connection function model is selected from the plurality of initial connection function models according to a goodness-of-fit test method, which can specifically include: calculating a test value corresponding to each initial connection function model according to the goodness-of-fit test method; and determining the initial connection function model with the smallest test value as the target connection function model.
[0137] It should be noted that in order to accurately select the target connection function model, a specific test value can be used to describe the characteristics of the connection function model, so that the selection is directly made by numerical comparison.
[0138] In the embodiments of the present application, the AIC criterion and the BIC criterion can be used for selection. The specific description of the AIC criterion and the BIC criterion can refer to the related description in step 203, and will not be repeated.
[0139] 206. Fusing the target marginal distribution model and the target connection function model to obtain a target joint distribution function.
[0140] In the embodiments of the present application, for the description of step 206, please refer to the detailed description of step 103 in the above embodiments, and the embodiments of the present application will not be repeated here.
[0141] 207. Simulate the measured wind power data by the target connection function model to obtain initial simulation sample data.
[0142] In the embodiments of the present application, according to the Copula function theory, the initial simulation sample data of the simulation data subject to the optimal target connection function model corresponding to the measured wind power data is simulated, and a corresponding scatter plot U can also be drawn according to the initial simulation sample data, as shown in Figure 5 .
[0143] 208. Convert the initial model sample data by the inverse function of the target marginal distribution model to obtain target simulation sample data.
[0144] In the embodiments of the present application, the initial model sample data is converted by the inverse function of the optimal target marginal distribution model to obtain equivalent samples subject to the optimal marginal distribution function and the optimal linking function, i.e. target simulation sample data, which can be as shown in Figure 6 , and it can be seen that the gap between the simulation data and the measured data is not large, which meets the requirements.
[0145] In some embodiments, after the target simulation sample data is calculated, the target simulation sample data can be detected. Specifically, the mean, standard deviation and other parameters of the target simulation sample data can be compared with the mean, standard deviation and other parameters of the measured data. When the difference is not more than a preset range, it means that the target simulation sample data is accurate.
[0146] The data simulation method for a wind farm provided in the embodiments of the present application can effectively represent the randomness characteristics of active power and wind speed from the perspective of probability theory, and the connection function can also effectively represent the correlation characteristics between wind speed and active power because the relationship between wind speed and active power is not a simple linear relationship. Therefore, the combination of the marginal distribution function and the connection function can more accurately and effectively realize the simulation and prediction of the missing data by constructing a joint distribution model.
[0147] As shown in Figure 7 , the present application provides a data simulation device for a wind farm, which can include:
[0148] The acquisition module 701 is configured to acquire measured wind power data of a wind turbine in a wind farm, and the measured wind power data includes wind speed and active power.
[0149] The processing module 702 is configured to construct a target edge distribution model and a target connection function model according to the measured wind power data, the target edge distribution model being used to represent randomness between the measured wind power data, and the target connection function model being used to represent correlation between the measured wind power data.
[0150] The processing module 702 is further configured to fuse the target edge distribution model and the target connection function model to obtain a target joint distribution function.
[0151] The processing module 702 is further configured to simulate the measured wind power data by using the target joint distribution function to obtain target simulation sample data.
[0152] In some embodiments, the processing module 702 is specifically configured to construct a plurality of initial edge distribution models according to the measured wind power data, the initial edge distribution models including at least two of a truncated normal distribution model, a lognormal distribution model, a truncated extreme value distribution model, and a Weibull distribution model.
[0153] The processing module 702 is specifically configured to select the target edge distribution model from the plurality of initial edge distribution models according to a goodness-of-fit test method.
[0154] In some embodiments, the processing module 702 is specifically configured to calculate a chi-square test value corresponding to each initial edge distribution model according to the goodness-of-fit test method.
[0155] The processing module 702 is specifically configured to determine the initial edge distribution model with the minimum chi-square test value as the target edge distribution model.
[0156] In some embodiments, the processing module 702 is specifically configured to construct a plurality of initial connection function models according to the measured wind power data, the initial connection function models including at least two of a Gaussian connection function model, a Plackett connection function model, and an Archimedean connection function model.
[0157] The processing module 702 is specifically configured to select the target connection function model from the plurality of initial connection function models according to the goodness-of-fit test method.
[0158] In some embodiments, the processing module 702 is specifically configured to calculate a chi-square test value corresponding to each initial connection function model according to the goodness-of-fit test method.
[0159] The processing module 702 is specifically configured to determine the initial connection function model with the minimum chi-square test value as the target connection function model.
[0160] In some embodiments, the processing module 702 is specifically configured to simulate the measured wind power data by using the target connection function model to obtain initial simulation sample data.
[0161] The processing module 702 is specifically configured to convert the initial simulation sample data by using an inverse function of the target edge distribution model to obtain target simulation sample data.
[0162] In the embodiments of the present application, each module can implement the wind farm data simulation method provided by the above method embodiments, and achieve the same technical effects. To avoid repetition, it will not be repeated here.
[0163] As shown in Figure 8 The embodiments of the present application also provide an electronic device, which can include:
[0164] The memory 801 stores executable program codes;
[0165] The processor 802 is coupled with the memory 801;
[0166] The processor 802 calls the executable program codes stored in the memory 801 to execute the wind farm data simulation method performed by the electronic device in each method embodiment.
[0167] The embodiments of the present application provide a computer readable storage medium, which stores a computer program. When the computer program is executed by a processor, each process of the wind farm data simulation method in the above method embodiments is implemented, and the same technical effects can be achieved. To avoid repetition, it will not be repeated here.
[0168] The embodiments of the present application also provide a computer program product, which stores a computer program. When the computer program is executed by a processor, each process of the wind farm data simulation method in the above method embodiments is implemented, and the same technical effects can be achieved. To avoid repetition, it will not be repeated here.
[0169] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can be in the form of a computer program product implemented on one or more computer usable storage media including computer usable program codes.
[0170] It should be understood that all the devices and methods disclosed in the embodiments of the present application can be implemented by other ways. The device embodiments described above are only schematic, and for instance, the flowcharts and the block diagrams in the figures illustrate the possible implementation modes of the devices, methods and computer program products according to the embodiments of the present application. In this regard, each block in the flowcharts or the block diagrams can represent a module, a program segment or a part of code, which contains one or more executable instructions for implementing the specified logic function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in a different order from that noted in the figures. For example, two consecutive blocks can actually be executed in parallel or in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or the flowcharts, and the combination of blocks in the block diagrams and / or the flowcharts, can be implemented by a dedicated hardware-based system for implementing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0171] In the present application, the processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0172] In the present application, the memory can include a non-persistent memory in a computer readable medium, random access memory (RAM) and / or non-volatile memory, etc., such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of the computer readable medium.
[0173] In the present application, those of ordinary skill in the art can understand that all or part of the steps of various methods in the above embodiments can be completed by relevant hardware instructed by a program, and the program can be stored in a computer readable storage medium, including permanent and non-permanent, removable and non-removable storage medium. The storage medium can realize information storage by any method or technology, and the information can be computer readable instructions, data structure, program module or other data. Examples of computer storage medium include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), other types of random access memory (RAM), read-only memory (ROM), one-time programmable read-only memory (OTPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, disk storage or other magnetic storage device or any other non-transmission medium that can be used to store information accessible by a computing device. According to the definition herein, the computer readable medium does not include transitory computer readable media such as modulated data signals and carriers.
[0174] It is to be noted that, in the present document, relational terms such as "first" and "second", and the like, are used solely to distinguish one entity or action from another entity or action, without necessarily requiring or implying any actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", "includes", "including", or the like, are intended to encompass non-exclusive inclusions, such that a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to those elements, but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. In addition, the term "coupled" or "coupling" is intended to mean a direct connection between two elements, or an indirect connection through one or more intermediate elements.
[0175] It is to be understood that the terminology "one embodiment" or "an embodiment" used throughout this document means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. Therefore, the appearances of the phrases "in one embodiment" or "in an embodiment" appearing in various places throughout the specification are not necessarily all referring to the same embodiment. In addition, the various specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It is also to be understood that the embodiments described in the specification are optional embodiments, and the actions and modules involved are not necessarily essential to the application. The above embodiments are not necessarily independent of each other, and are combined only to highlight different technical features in different embodiments. Those skilled in the art should know that the above embodiments can be combined in any manner.
[0176] In various embodiments of the present application, it should be understood that the size of the sequence number of the above processes does not mean the inevitable sequence of execution, and the execution sequence of the processes should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0177] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e. they can be located in one place, or can be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments of the present application.
[0178] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0179] The integrated units described above, if implemented in the form of software function units and sold or used as independent products, can be stored in a computer accessible memory. Based on such understanding, the technical solutions of the present application essentially or the part of the prior art that contributes to the present application or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a memory and includes a plurality of parts or all steps of the above-mentioned method for enabling a computer device (which can be a personal computer, a server or a network device, etc., and specifically can be a processor in the computer device) to execute the embodiments of the present application.
[0180] The above is only a specific implementation of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data simulation method for wind farms, characterized in that, The method includes: Obtain measured wind power data of the wind turbines in the wind farm, the measured wind power data including wind speed and active power; Based on the measured wind power data, a target edge distribution model and a target connection function model are constructed. The target edge distribution model is used to characterize the randomness between the measured wind power data, and the target connection function model is used to characterize the correlation between the measured wind power data. The target edge distribution model and the target connectivity function model are fused to obtain the target joint distribution function; The measured wind power data are simulated using the target joint distribution function to obtain target simulation sample data.
2. The method according to claim 1, characterized in that, The step of constructing a target edge distribution model based on the measured wind power data includes: Based on the measured wind power data, multiple initial marginal distribution models are constructed, including at least two of the following: truncated normal distribution model, log-normal distribution model, truncated extreme value distribution model, and Weibull distribution model; The target marginal distribution model is selected from the plurality of initial marginal distribution models according to the goodness-of-fit test method.
3. The method according to claim 2, characterized in that, The step of selecting the target marginal distribution model from the plurality of initial marginal distribution models according to the goodness-of-fit test method includes: Based on the goodness-of-fit test method, calculate the quasi-test value corresponding to each initial marginal distribution model; The initial edge distribution model with the smallest quasi-test value is determined as the target edge distribution model.
4. The method according to claim 1, characterized in that, The step of constructing the target connection function model based on the measured wind power data includes: Based on the measured wind power data, multiple initial connection function models are constructed, including at least two of the following: Gaussian connection function model, Precutter connection function model, and Archimedes connection function model. The target connection function model is selected from the plurality of initial connection function models according to the goodness-of-fit test method.
5. The method according to claim 4, characterized in that, The step of selecting the target connection function model from the plurality of initial connection function models according to the goodness-of-fit test method includes: Based on the goodness-of-fit test method, calculate the quasi-test value corresponding to each initial connection function model; The initial connection function model with the smallest quasi-test value is determined as the target connection function model.
6. The method according to claim 1, characterized in that, The step of simulating the measured wind power data using the target joint distribution function to obtain target simulation sample data includes: The measured wind power data are simulated using the target connection function model to obtain initial simulation sample data; The initial simulated sample data is transformed by the inverse function of the target marginal distribution model to obtain the target simulated sample data.
7. A data simulation device for a wind farm, characterized in that, include: The acquisition module is used to acquire measured wind power data of the wind turbines in the wind farm, including wind speed and active power. The processing module is used to construct a target edge distribution model and a target connection function model based on the measured wind power data. The target edge distribution model is used to characterize the randomness between the measured wind power data, and the target connection function model is used to characterize the correlation between the measured wind power data. The processing module is further configured to fuse the target edge distribution model and the target connectivity function model to obtain the target joint distribution function; The processing module is also used to simulate the measured wind power data using the target joint distribution function to obtain target simulation sample data.
8. An electronic device, characterized in that, include: Memory containing executable program code; and the processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the wind farm data simulation method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, include: The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the data simulation method for a wind farm as described in any one of claims 1 to 6.