Offshore wind power output evaluation method and system based on hybrid clustering and adaptive kernel density estimation

By employing a hybrid clustering and adaptive kernel density estimation method, the problems of multi-peak and spatiotemporal differences in offshore wind power output assessment were solved, enabling more accurate output characteristic assessment and improving the safety, stability and dispatch efficiency of the power system.

CN121923262APending Publication Date: 2026-04-24GUANGXI POWER GRID CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGXI POWER GRID CORP
Filing Date
2025-12-11
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing methods for assessing offshore wind power output are ill-suited to its multi-peak, asymmetric, and spatiotemporal variations, resulting in insufficient fitting accuracy and failing to adequately consider the impact of seasonal and regional factors, thus affecting the safe and stable operation of the power system.

Method used

A hybrid clustering and adaptive kernel density estimation method is adopted. The optimal number of clusters is determined by the Bayesian information criterion. By combining the Gaussian mixture model and nonparametric kernel density estimation, the bandwidth range is dynamically adjusted and spatiotemporal labels are associated to construct a multi-dimensional power output probability density function to evaluate the power output status of offshore wind power.

Benefits of technology

It significantly improves the computational efficiency and robustness of the model, and can more accurately reflect the multi-peak, asymmetric, and spatiotemporal differences of offshore wind power output. It is suitable for evaluating the output characteristics of single and multi-regional power generation, and improves the efficiency and reliability of grid dispatch.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121923262A_ABST
    Figure CN121923262A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of offshore wind power output evaluation, in particular to an offshore wind power output evaluation method and system based on hybrid clustering and adaptive kernel density estimation. According to the method, through adaptive bandwidth selection and abnormal value elimination, the calculation efficiency and robustness of the model are remarkably improved. Compared with a traditional parameter model method and a fixed bandwidth kernel density estimation method, the method has obvious advantages in fitting precision, and can better reflect multimodality, asymmetry and space-time difference of offshore wind power output. The method is not only suitable for depicting and evaluating the offshore wind power output characteristics of a single area, but also can be expanded and applied to offshore wind power output analysis of multiple areas and multiple time scales, and has good adaptability and expansibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of offshore wind power output assessment technology, and in particular to an offshore wind power output assessment method and system based on hybrid clustering and adaptive kernel density estimation. Background Technology

[0002] As the global energy structure transitions towards a low-carbon model, offshore wind power, as a clean and renewable energy source, is gradually becoming an important component of energy supply. However, offshore wind power output is characterized by strong intermittency, multiple peaks, and significant spatiotemporal variations. Its randomness and volatility pose serious challenges to the safe and stable operation of the power system. Therefore, accurately assessing the output characteristics of offshore wind power is crucial for optimizing wind farm site selection, improving grid dispatch efficiency, and reducing wind curtailment rates.

[0003] Currently, methods for evaluating the output characteristics of offshore wind power are mainly divided into two categories: parametric model methods and non-parametric model methods. Parametric model methods (such as Weibull distribution, normal distribution, etc.) model the output distribution by pre-setting the distribution form, but they rely on prior assumptions and are difficult to adapt to the multi-peak and asymmetric nature of offshore wind power output, resulting in insufficient fitting accuracy. Non-parametric model methods (such as kernel density estimation) are data-driven and do not require pre-setting the distribution form, which can better reflect complex output characteristics, but their accuracy is greatly affected by the choice of bandwidth parameter. Traditional fixed bandwidth or empirical selection methods are difficult to balance smoothness and local peak characteristics, which easily leads to overfitting or underfitting.

[0004] Furthermore, existing methods often neglect the spatiotemporal variability of offshore wind power output and fail to fully consider the impact of seasonal and regional factors on output characteristics, resulting in insufficient model universality. To address these issues, there is an urgent need for a method that can accurately characterize the output characteristics of offshore wind power, combining spatiotemporal labels to classify and extract output curves more reasonably and meticulously, providing reliable data support for power system planning and dispatch.

[0005] Therefore, it is necessary to study a method and system for evaluating offshore wind power output based on hybrid clustering and adaptive kernel density estimation. Summary of the Invention

[0006] To address the problems in existing technologies, this invention provides a method and system for evaluating offshore wind power output based on hybrid clustering and adaptive kernel density estimation. The specific technical solution is as follows: A method for evaluating offshore wind power output based on hybrid clustering and adaptive kernel density estimation includes the following steps: Step S1: Collect historical power output data of offshore wind power in multiple regions and process the data; Step S2: The optimal number of clusters k is determined using the Bayesian information criterion, and Gaussian mixture clustering is performed on the offshore wind power output curves to generate curve clusters that reflect different output modes. Step S3: Apply nonparametric kernel density estimation independently to each type of curve cluster, adaptively calculate the optimal bandwidth based on the criterion of minimizing the asymptotic integral mean square error, and dynamically adjust the bandwidth range through multimodality test; Step S4: Associate the clustering results with spatiotemporal labels, including seasonal and regional information, to construct a multi-dimensional power output probability density function. Based on the power output probability density function, calculate the confidence interval and upper and lower limits of power output fluctuation for offshore wind power, and assess the state of offshore wind power output.

[0007] Preferably, step S1, which involves collecting historical offshore wind power output data from multiple regions and processing the data, specifically includes the following steps: Based on the regional offshore wind power distribution and 8760 hours of wind measurement data from offshore wind towers, combined with sea location, distance from shore, and wind speed, the entire region's offshore wind power is divided into... N One region; By combining the offshore wind power scale of each wind zone within the sub-region and the original hourly offshore wind power output curves of each wind zone, the original hourly offshore wind power output curve of the sub-region can be obtained by weighted summation, as shown in the following formula: ; In the formula: A it This represents the original hourly offshore wind power output curve for wind zone i. m This represents the total number of wind zones within the current region; P it The original curve of offshore wind power output in wind zone i is t=1,2,...,8760; α i It represents the proportion of offshore wind power capacity in wind zone i to the total offshore wind power capacity in the current region; The power output sequences of offshore wind farms in all zones within the entire region are superimposed to obtain the power output curve of offshore wind power across the entire region.

[0008] Preferably, in step S2, the optimal number of clusters k is determined using the Bayesian information criterion as follows: If the error or disturbance of the Bayesian information criterion model follows a normal distribution, then the Bayesian information criterion model can be expressed as: ; in, k The number of model parameters represents the number of Gaussian distributions and the total number of parameters in the Gaussian mixture model. n The sample size represents the total number of offshore wind power output curves. LThe maximum value of the model likelihood function represents the maximum likelihood estimate of the sample data given the model. kln(n) Penalty items, S RSS The sum of squared residuals represents the estimated model; By calculating different cluster numbers k Choose the BIC value that minimizes the BIC value. k As the optimal number of clusters.

[0009] Preferably, step S2, which involves performing Gaussian mixture clustering on the offshore wind power output curves to generate curve clusters reflecting different output modes, specifically includes: Assume the wind power output per hour of the day is , i =1,2,…,24, then the Gaussian mixture model is expressed as: ; In the formula: x It is a random variable; Let x be the probability of a random variable. It is a weighting coefficient, and satisfies ; Let be the distribution of the k-th Gaussian component in the Gaussian mixture model; The expectation-maximization algorithm is used to estimate the three parameters in the above Gaussian mixture model. The three parameters are the mean of the k-th Gaussian component. Weighting coefficient π k and variance ; Transform the above equation into: ; (1) First, specify 3 parameters. m、 π e The initial value; (2) Calculate the posterior probability The calculation method is as follows: ; (3) Based on the posterior probability Solve The maximum likelihood function is as follows: ; in, For the first n Observed values ​​of one sample of offshore wind power; N The total number of samples; N k For the first k The number of valid samples for each Gaussian distribution component; (4) Based on the posterior probability and beg The maximum likelihood value is as follows: ; (5) Solve The maximum likelihood function is as follows: ; (6) If satisfied , and ,in α 1. α 2 and α 3 is the convergence threshold, then m ,π e The values ​​are respectively taken m (t) , π (t) , e (t) Otherwise, Repeat steps (2) to (6) until the algorithm converges; superscript and These represent the EM algorithm at the 1st, 2nd, and 3rd respectively. t Second and third t The parameter estimates obtained from -1 iterations.

[0010] Preferably, step S2 further includes removing outliers using box plots and extracting representative feature curves for each type of curve cluster, wherein removing outliers using box plots is specifically as follows: (1) Plot a box plot for the output data of each type of curve cluster, and calculate the lower quartile Q1, upper quartile Q3, interquartile range IQR, upper edge and lower edge; where the interquartile range IQR = Q3 - Q1; upper edge = Q3 + 1.5 × IQR; lower edge = Q1 - 1.5 × IQR; (2) Identify data points located outside the upper and lower edges, mark them as outliers, remove outliers, and retain normal data points.

[0011] Preferably, step S3 independently applies nonparametric kernel density estimation to each type of curve cluster, adaptively calculates the optimal bandwidth based on the criterion of minimizing the asymptotic integral mean square error, and dynamically adjusts the bandwidth interval through a multimodality test, specifically including the following steps: (1) Assumption p 1, p 2,…, p n Data sample set for offshore wind power p If a Gaussian function is used as the probability density estimation model for offshore wind power output, then this sample set p The nonparametric KDE probability density function can be expressed as: ; In the formula, This represents the wind power probability density function based on the nonparametric KDE. The bandwidth parameter for kernel density estimation; This is the i-th sample value of the active power output of wind power. u i =( p - p i ) / l bw ; n The total number of samples for offshore wind power output data; Bias produced by fitting estimation B and variance V They are respectively: ; ; in: ; ; In the formula, For the true probability density function of offshore wind power output; It is the second derivative of the true probability density function; Represents a higher-order infinitesimal; For kernel functions; is the independent variable of the kernel function; (2) The optimal bandwidth is adaptively calculated based on the criterion of minimizing the asymptotic integral mean square error. The optimal bandwidth is obtained when the asymptotic integral mean square error is minimized. The calculation formula is: ; ; Optimal bandwidth is determined by applying the normal reference criterion. The calculation formula is further simplified to: ; in, The standard deviation of the sample data for offshore wind power output; Replace the above formula with the following formula. The details are as follows: ; in, The divergence measure is the half-range. Represents the standard normal cumulative distribution function; The formula for calculating the bandwidth of the optimal kernel function is: .

[0012] Preferably, step S4 involves associating the clustering results with spatiotemporal labels, including seasonal and regional information, to construct a multi-dimensional power output probability density function. Based on this function, the confidence interval and upper and lower limits of power output fluctuations for offshore wind power are calculated. The assessment of the state of offshore wind power output specifically includes the following steps: (1) Spatiotemporal tags include seasons S and region R The method for constructing a multi-dimensional output probability density function is as follows: calculate the probability density function of the label in a specific time and space. conditional probability density function : ; In the formula: K The total number of clusters; For the first k Curve clusters based on adaptive bandwidth The resulting kernel density estimate probability density function; For the season S and region R Under the conditions k The probability weight of the occurrence of the power output mode, and satisfying the following: ; (2) Based on the power output probability density function, calculate the confidence interval of offshore wind power output and calculate the effective power output rate to evaluate the system operating status. The method for calculating the confidence interval is as follows: By analyzing the multidimensional power output probability density function Integrating yields the cumulative distribution function. Given a confidence level Calculate the upper and lower limits of force fluctuation. : ; In the formula, It is the inverse function of the cumulative distribution function; if the actual offshore wind power output exceeds this range, it is judged as an abnormal fluctuation. Calculate the effective power output of offshore wind power The formula is: ; In the formula, P WT This represents the actual power output of offshore wind power. P max The maximum power output of offshore wind power; based on the effective power output rate of offshore wind power. To assess the actual production efficiency of a wind farm, if If the value is below the preset threshold, it is identified as an inefficient operating period, and the operation strategy of the offshore wind turbine needs to be optimized. (3) Based on the power output probability density function, generate a typical scenario set for offshore wind power output and calculate the wind curtailment rate of offshore wind power for power system dispatch and planning; the calculation of the wind curtailment rate of offshore wind power The formula is: ; In the formula, For the curtailed wind power, based on The extent of insufficient grid absorption capacity is quantified, and areas with high wind curtailment rates are identified as key planning targets.

[0013] A system for assessing offshore wind power output based on hybrid clustering and adaptive kernel density estimation, comprising the following methods: The data acquisition and processing module is used to collect historical power output data of offshore wind power in multiple regions and process the data. The clustering module is used to determine the optimal number of clusters k using the Bayesian information criterion, and to perform Gaussian mixture clustering on the offshore wind power output curves to generate curve clusters that reflect different output modes. The bandwidth optimization module is used to independently apply nonparametric kernel density estimation to each type of curve cluster, adaptively calculate the optimal bandwidth based on the criterion of minimizing the asymptotic integral mean square error, and dynamically adjust the bandwidth range through multimodality test; The evaluation module is used to associate clustering results with spatiotemporal labels, including seasonal and regional information, to construct a multi-dimensional power output probability density function. Based on the power output probability density function, the confidence interval and upper and lower limits of power output fluctuation of offshore wind power are calculated to evaluate the state of offshore wind power output.

[0014] A computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device where the computer-readable storage medium is located to perform the aforementioned method for evaluating offshore wind power output based on hybrid clustering and adaptive kernel density estimation.

[0015] A processor for running a program, wherein the program executes the offshore wind power output assessment method based on hybrid clustering and adaptive kernel density estimation.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: The method of this invention significantly improves the computational efficiency and robustness of the model through adaptive bandwidth selection and outlier removal. Compared with traditional parametric modeling and fixed-bandwidth kernel density estimation methods, the method of this invention has a significant advantage in fitting accuracy and can better reflect the multi-peak, asymmetric, and spatiotemporal differences in offshore wind power output. This method is not only applicable to characterizing and evaluating the offshore wind power output characteristics of a single region, but can also be extended to the analysis of offshore wind power output in multiple regions and at multiple time scales, demonstrating good adaptability and scalability. Attached Figure Description

[0017] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.

[0018] Figure 1 This is a flowchart of the method of the present invention.

[0019] Figure 2 This is a flowchart of the method of the present invention.

[0020] Figure 3 This is a schematic diagram of the Gaussian mixture clustering results of the present invention.

[0021] Figure 4 This is a comparison chart of the fitting effects of the adaptive bandwidth KDE of the present invention and the traditional method.

[0022] Figure 5 This is a system schematic diagram of the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0025] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0026] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0027] Example 1: like Figure 1 As shown, this embodiment provides a method for evaluating offshore wind power output based on hybrid clustering and adaptive kernel density estimation, including the following steps: Step S1 involves collecting historical offshore wind power output data from multiple regions and processing the data. This includes the following steps: Based on the regional offshore wind power distribution and 8760 hours of wind measurement data from offshore wind towers, combined with sea location, distance from shore, and wind speed, the entire region's offshore wind power is divided into... N The dataset is divided into regions, such as eastern Guangdong, western Guangdong, and the Pearl River Delta region, and time scales, including day, season, and year. The dataset is then divided according to geographical regions and time scales, and the data is normalized.

[0028] By combining the offshore wind power scale of each wind zone within the sub-region and the original hourly offshore wind power output curves of each wind zone, the original hourly offshore wind power output curve of the sub-region can be obtained by weighted summation, as shown in the following formula: ; In the formula: A it This represents the original hourly offshore wind power output curve for wind zone i. m This represents the total number of wind zones within the current region; P it The original curve of offshore wind power output in wind zone i is t=1,2,...,8760; α i It represents the proportion of offshore wind power capacity in wind zone i to the total offshore wind power capacity in the current region; The power output sequences of offshore wind farms in all zones within the entire region are superimposed to obtain the power output curve of offshore wind power across the entire region.

[0029] Step S2: The optimal number of clusters k is determined using the Bayesian information criterion, and Gaussian mixture clustering is performed on the offshore wind power output curves to generate curve clusters that reflect different output modes.

[0030] The optimal number of clusters k is determined using the Bayesian information criterion as follows: The Bayesian information criterion is a statistical criterion used for model selection. Its goal is to find a balance between model complexity and goodness of fit, avoiding overfitting or underfitting. The formula for the Bayesian information criterion is: ; In the formula: k The number of model parameters represents the number of Gaussian distributions and the total number of their parameters (mean, variance, weight coefficients) in the Gaussian mixture model. n The sample size represents the total number of offshore wind power output curves. L The maximum value of the model likelihood function represents the maximum likelihood estimate of the sample data under a given model. Kln(n) The penalty term can effectively prevent the curse of dimensionality when the dimensionality is too high and the training sample data is relatively small.

[0031] If the error or disturbance of the Bayesian information criterion model follows a normal distribution, then the Bayesian information criterion model can be expressed as: ; in, S RSS This represents the sum of squared residuals of the estimated model.

[0032] By calculating different cluster numbers k Choose the BIC value that minimizes the BIC value. k As the optimal number of clusters.

[0033] BIC is S RSS and k The increasing function of BIC means that the introduction of residuals and unknown parameters will increase BIC. Therefore, when determining the optimal number of clusters for offshore wind power output, models with low BIC values ​​are preferred.

[0034] The core idea of ​​BIC: BIC introduces penalty terms. Kln(n) This penalizes the model's complexity. As the number of clusters increases... k As the complexity of the model increases, the BIC value increases; conversely, as the complexity of the model decreases, the BIC value decreases. Meanwhile, BIC is affected by -2ln( L This measures the goodness of fit of the model. A higher goodness of fit indicates a better fit. L The larger it is, -2ln( L The smaller the value of BIC, the smaller the BIC value. Therefore, a smaller BIC value indicates that the model has achieved a better balance between complexity and goodness of fit.

[0035] In characterizing the power output characteristics of offshore wind power, Gaussian mixture clustering is used to divide the power output curve into several classes, each corresponding to a typical power output mode (such as anti-peak shaving, high fluctuation, and stable). Different cluster numbers are calculated... k Choose the BIC value that minimizes the BIC value. k As the optimal number of clusters. For example, when k=3, the BIC value is the smallest, indicating that dividing the offshore wind power output curve into 3 categories can better balance model complexity and goodness of fit.

[0036] Advantages of BIC: Avoids overfitting through a penalty term. Kln(n) BIC effectively prevents the selection of too many clusters, avoiding overly complex models. It adapts to sample size by considering the number of samples (n), making it suitable for datasets of varying sizes. Furthermore, it is theoretically supported by Bayesian theory, possessing a strong statistical foundation and making it suitable for model selection with complex data distributions.

[0037] Using the methods described above, the Bayesian Information Criterion (BIC) can effectively determine the optimal number of clusters for Gaussian mixture clustering, providing theoretical support for the accurate characterization of offshore wind power output.

[0038] GMM is a probabilistic clustering method that assumes that the input samples follow a certain order. k A Gaussian distribution with unknown parameters is used to cluster samples that follow the same distribution into one class. GMM uses the Expectation-Maximization (EM) algorithm to fit k mixture Gaussian distributions to obtain the mean and covariance of each distribution. Compared to widely used... k Methods such as k-means clustering and hierarchical agglomerative clustering are used. GMM clustering can achieve better fitting results for complex distributions and its clustering effect is better than k-means.

[0039] Therefore, Gaussian mixture clustering is performed on the offshore wind power output curves to generate curve clusters reflecting different output modes, specifically including: Assume the wind power output per hour of the day is , i =1,2,…,24, then the Gaussian mixture model is expressed as: ; In the formula: x It is a random variable; Let x be the probability of a random variable. It is a weighting coefficient, and satisfies ; Let be the distribution of the k-th Gaussian component in the Gaussian mixture model.

[0040] The Gaussian mixture model above has three parameters that need to be estimated: the mean, the mean, and the mean. Weighting coefficient π k and variance .

[0041] The expectation-maximization algorithm is used to estimate the three parameters in the above Gaussian mixture model. The three parameters are the mean of the k-th Gaussian component. Weighting coefficient π k and variance .

[0042] Transform the above equation into: ; (1) First, specify 3 parameters. m、 π e The initial value; (2) Calculate the posterior probability The calculation method is as follows: ; (3) Based on the posterior probability Solve The maximum likelihood function is as follows: ; in, For the first n Observed values ​​of one sample of offshore wind power; N The total number of samples; N k For the first k The number of valid samples for the nth Gaussian distribution component (i.e., all samples belonging to the nth Gaussian distribution component) k (the sum of the posterior probabilities of the classes); (4) Based on the posterior probability and beg The maximum likelihood value is as follows: ; (5) Solve The maximum likelihood function is as follows: ; (6) If satisfied , and ,in α 1. α 2 and α 3 is the convergence threshold, then m ,π e The values ​​are respectively taken m (t) , π (t) , e (t) Otherwise, Repeat steps (2) to (6) until the algorithm converges; superscript and These represent the EM algorithm at the 1st, 2nd, and 3rd respectively. t Second and third t The parameter estimates obtained from -1 iterations.

[0043] Step S2 also includes using box plots to remove outliers and extracting representative characteristic curves for each curve cluster. A box plot is a statistical tool used for data distribution visualization and outlier detection; its structure includes the following key elements: Median (Q2): The middle value of the dataset, indicating the center of the data; Upper quartile (Q3): 75% of the data points in the dataset are less than or equal to this value; Lower quartile (Q1): 25% of the data points in the dataset are less than or equal to this value; Interquartile Range (IQR): IQR = Q3 - Q1, used to measure the dispersion of data; Top edge: Q3 + 1.5 × IQR, representing the upper limit of the data; Lower edge: Q1 - 1.5 × IQR, representing the lower limit of the data; Outliers: Data points located outside the top and bottom edges are considered outliers.

[0044] In the process of clustering offshore wind power output curves, box plots are used to identify and remove outliers, ensuring the accuracy and robustness of the clustering results.

[0045] The specific steps for removing outliers using box plots are as follows: (1) Plot a box plot for the output data of each type of curve cluster, and calculate the lower quartile Q1, upper quartile Q3, interquartile range IQR, upper edge and lower edge; where the interquartile range IQR = Q3 - Q1; upper edge = Q3 + 1.5 × IQR; lower edge = Q1 - 1.5 × IQR; (2) Identify data points located outside the upper and lower edges, mark them as outliers, remove outliers, and retain normal data points.

[0046] After outlier removal, representative characteristic curves for each cluster are extracted for subsequent adaptive kernel density estimation and spatiotemporal characteristic analysis. Outlier detection using box plots effectively improves the stability and reliability of clustering results, providing data support for the accurate characterization of offshore wind power output characteristics.

[0047] Step S3: Apply nonparametric kernel density estimation independently to each type of curve cluster, adaptively calculate the optimal bandwidth based on the criterion of minimizing the mean square error of the asymptotic integral, and dynamically adjust the bandwidth range through multimodality test.

[0048] Nonparametric kernel density estimation is a data-driven probability density estimation method that does not require assumptions that the data follows a specific distribution. It is suitable for offshore wind power output data, which has multimodal, asymmetric, and complex distribution characteristics.

[0049] The nonparametric kernel density estimation is applied independently to each type of curve family. The optimal bandwidth is adaptively calculated based on the criterion of minimizing the asymptotic integral mean square error, and the bandwidth interval is dynamically adjusted through a multimodality test. The specific steps include: (1) Assumption p 1, p 2,…, p n Data sample set for offshore wind power p If a Gaussian function is used as the probability density estimation model for offshore wind power output, then this sample set p The nonparametric KDE probability density function can be expressed as: ; according to K The requirements of continuity, symmetry, and nonnegativity of (p) K (p) The following constraints must also be met: ; In the formula: c is a constant. For K (p) In terms of kernel functions, there are many suitable options, such as Gaussian functions and polynomial functions, meaning there is diversity in the choice of kernel function. However, related studies have shown that when different kernel functions are selected, the errors in the fitting results are all within a small fluctuation range, which suggests that the kernel function... K (p) does not affect the fitting of the nonparametric KDE. Therefore, the Gaussian function is used as the probability density estimation model for offshore wind power output. K (p) can be represented as: ; Generally, let ui =( p - pi ) / l bw, further rewriting the nonparametric KDE of the probability density function of active power output of offshore wind power as: ; In the formula, This represents the wind power probability density function based on the nonparametric KDE. The bandwidth parameter for kernel density estimation; This is the i-th sample value of the active power output of wind power. u i =( p - p i ) / l bw ; n The total number of samples for offshore wind power output data; Based on the Asymptotic integral mean square error minimization criterion (AMISE), the optimal bandwidth is adaptively calculated, and the bandwidth range is dynamically adjusted through multi-peak test to ensure that the fitted curve takes into account both smoothness and local peak characteristics. The following is a further description of the optimal bandwidth selection for the kernel function: For nonparametric KDE models, bandwidth l The choice of bw parameters is crucial. If bandwidth... l If the bw value is too large, KDE will exhibit an excessive pursuit of the probability density function. The smoothness of the data masks the structural characteristics of the data samples, making it impossible to reflect the multi-peak nature of offshore wind power output characteristics, resulting in a large fitting error; if the bandwidth l When the bw value is too small, although the KDE fit can reflect some peak characteristics of the data, it is prone to overfitting, meaning that the resulting probability density function will have excessive fluctuations and will not reflect reality. Therefore, the selection of bandwidth should be based on an analysis of the actual error deviation and variance generated during fitting, and the optimal bandwidth should be calculated.

[0050] For the data sample set of offshore wind power p In other words, as the number of samples approaches positive infinity, the bandwidth of the nonparametric KDE... l bw , nl bw They tend towards 0 and +∞ respectively. At this point, the bias produced by the fitted estimate... B and variance V They are respectively: ; ; From the above, we can see that l Decreasing bw will lead to deviation B Reduce variance V Increase; conversely, decrease, then deviate. B Increased variance V Therefore, the bandwidth cannot simultaneously satisfy both the reduction of bias and variance in KDE estimation; a trade-off must be made between the two to arrive at the optimal bandwidth selection.

[0051] in: ; ; In the formula, For the true probability density function of offshore wind power output; It is the second derivative of the true probability density function; Represents a higher-order infinitesimal; For kernel functions; is the independent variable of the kernel function; (2) To this end, an optimal bandwidth calculation method is proposed to minimize the estimated asymptotic mean square error (AMISE). AMISE can comprehensively weigh the bias and variance of KDE. Based on the criterion of minimizing the asymptotic mean square error, the optimal bandwidth is adaptively calculated, and the optimal bandwidth is obtained when the asymptotic mean square error is minimized. The calculation formula is: ; ; Optimal bandwidth is determined by applying the normal reference criterion. The calculation formula is further simplified to: ; in, The standard deviation of the sample data for offshore wind power output; Generally, it is recommended to consider the more robust divergence measure, half-range. I qr Replace the above formula with the following formula. The details are as follows: ; in, The divergence measure is the half-range. Represents the standard normal cumulative distribution function; In this embodiment, the coefficient is set to 1.06 to achieve accurate estimation of the multi-peak probability density curve. The formula for calculating the optimal kernel function bandwidth is then: .

[0052] Plot probability density curves to analyze the distribution characteristics of offshore wind power output (such as multiple peak locations and fluctuation ranges).

[0053] Establish an evaluation index system, including the following indicators: ; ; ; In the formula p g,i , p o,iThese represent the fitted distribution data of the nonparametric KDE and the probability density values ​​corresponding to the actual frequency distribution interval, respectively. , Representative sample set p g,i , p o,i The sample mean.

[0054] Among the three indicators mentioned above, the smaller the error value, the smaller the difference between the fitted model of offshore wind power output distribution in the corresponding period and the data observation distribution, and the more accurate the nonparametric KDE fitting effect.

[0055] Step S4: Associate the clustering results with spatiotemporal labels, including seasonal and regional information, to construct a multi-dimensional power output probability density function. Based on the power output probability density function, calculate the confidence interval and upper and lower limits of power output fluctuation for offshore wind power, and assess the state of offshore wind power output.

[0056] The clustering results are associated with spatiotemporal labels, including seasonal and regional information, to construct a multi-dimensional power output probability density function. Based on the power output probability density function, the confidence interval and upper and lower limits of power output fluctuations of offshore wind power are calculated. The assessment of the state of offshore wind power output specifically includes the following steps: (1) Spatiotemporal tags include seasons S and region R The method for constructing a multi-dimensional output probability density function is as follows: calculate the probability density function of the label in a specific time and space. conditional probability density function : ; In the formula: K The total number of clusters; For the first k Curve clusters based on adaptive bandwidth The resulting kernel density estimate probability density function; For the season S and region R Under the conditions k The probability weight of the occurrence of the power output mode, and satisfying the following: ; (2) Based on the power output probability density function, calculate the confidence interval of offshore wind power output and calculate the effective power output rate to evaluate the system operating status. The method for calculating the confidence interval is as follows: By analyzing the multidimensional power output probability density function Integrating yields the cumulative distribution function. Given a confidence level Calculate the upper and lower limits of force fluctuation. : ; In the formula, It is the inverse function of the cumulative distribution function; if the actual offshore wind power output exceeds this range, it is judged as an abnormal fluctuation. Calculate the effective power output of offshore wind power The formula is: ; In the formula, P WT This represents the actual power output of offshore wind power. P max The maximum power output of offshore wind power; based on the effective power output rate of offshore wind power. To assess the actual production efficiency of a wind farm, if If the value is below the preset threshold, it is identified as an inefficient operating period, and the operation strategy of the offshore wind turbine needs to be optimized. (3) Based on the power output probability density function, generate a typical scenario set for offshore wind power output and calculate the wind curtailment rate of offshore wind power for power system dispatch and planning; the calculation of the wind curtailment rate of offshore wind power The formula is: ; In the formula, For the curtailed wind power, based on The extent of insufficient grid absorption capacity is quantified, and areas with high wind curtailment rates are identified as key planning targets.

[0057] This embodiment uses historical power output data from an offshore wind farm in a certain region over one year for simulation analysis to verify the effectiveness of the Gaussian mixture clustering method proposed in this invention. Figure 3 As shown, the Gaussian mixture clustering results of offshore wind power output optimized based on the Bayesian Information Criterion (BIC) are presented. The figure illustrates three typical output mode curve clusters: Mode 1 (marked as the "Stable" curve) exhibits low-output stability, mainly corresponding to seasons or periods with low winds, and shows relatively small fluctuations; Mode 2 (marked as the "High-Volatility" curve) exhibits medium-to-high output with significant fluctuations, reflecting the significant upslope and fluctuation characteristics of offshore wind power under the influence of monsoons, and is a key focus of grid dispatch; Mode 3 (marked as the "Anti-Peak Shaving" curve) exhibits full-load or high-output states, corresponding to periods with abundant wind resources. The above analysis shows that by introducing the BIC criterion to determine the optimal number of clusters (k=3), the GMM model can effectively identify and separate the multimodal characteristics of offshore wind power output, not only characterizing the average output level but also accurately capturing the output fluctuation patterns under different meteorological conditions, providing a precise data foundation for subsequent scenario-specific modeling.

[0058] This embodiment analyzes the impact of adaptive bandwidth kernel density estimation (KDE) on the fitting accuracy of output characteristics. The method of this invention is compared with the traditional fixed bandwidth KDE method and the parameter method (Weibull distribution), and the fitting results are shown in the appendix. Figure 4 As shown. By Figure 3 As can be seen, at the peak of the probability density curve (i.e., the high-probability interval), the fitting curve of the traditional fixed-bandwidth method (shown by the red dashed line in the figure) is relatively flat, failing to fully fit the peak characteristics of the data (underfitting). In contrast, the fitting curve of this invention, based on the AMISE criterion and adaptively adjusting the bandwidth (shown by the blue solid line in the figure), automatically reduces the bandwidth at local data-dense areas, making the fitting curve closely follow the histogram peaks and accurately reproducing the multi-peak nature of the power output. At the tail end of the probability density curve (i.e., the low-probability interval), the method of this invention appropriately increases the bandwidth, ensuring the smoothness of the curve and avoiding spurious fluctuations. Further analysis of error indicators shows that the root mean square error (RMSE) of the method of this invention is 0.015, which is approximately 46.4% lower than the traditional fixed-bandwidth KDE method (RMSE=0.028) and approximately 64.3% lower than the parametric method (RMSE=0.042). Therefore, the method of this invention has significant advantages in balancing the smoothness of the fitting curve with local detailed features, and can more accurately characterize the probability distribution characteristics of offshore wind power output.

[0059] Example 2: like Figure 5 As shown, based on the same inventive concept as Embodiment 1, this embodiment provides an offshore wind power output assessment system based on hybrid clustering and adaptive kernel density estimation. The method described includes: The data acquisition and processing module is used to collect historical power output data of offshore wind power in multiple regions and process the data. The clustering module is used to determine the optimal number of clusters k using the Bayesian information criterion, and to perform Gaussian mixture clustering on the offshore wind power output curves to generate curve clusters that reflect different output modes. The bandwidth optimization module is used to independently apply nonparametric kernel density estimation to each type of curve cluster, adaptively calculate the optimal bandwidth based on the criterion of minimizing the asymptotic integral mean square error, and dynamically adjust the bandwidth range through multimodality test; The evaluation module is used to associate clustering results with spatiotemporal labels, including seasonal and regional information, to construct a multi-dimensional power output probability density function. Based on the power output probability density function, the confidence interval and upper and lower limits of power output fluctuation of offshore wind power are calculated to evaluate the state of offshore wind power output.

[0060] Example 3: Based on the same inventive concept as Embodiment 1, this embodiment provides a computer-readable storage medium, which includes a stored program, wherein, when the program is executed, it controls the device where the computer-readable storage medium is located to execute the aforementioned method for evaluating offshore wind power output based on hybrid clustering and adaptive kernel density estimation.

[0061] Example 4: Based on the same inventive concept as Embodiment 1, this embodiment provides a processor for running a program, wherein the program executes the offshore wind power output assessment method based on hybrid clustering and adaptive kernel density estimation.

[0062] Those skilled in the art will recognize that the modules of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components of the examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of the invention.

[0063] In the embodiments provided by this invention, it should be understood that the division of modules is only a logical functional division. In actual implementation, there may be other division methods, such as multiple modules can be combined into one module, one module can be split into multiple modules, or some features can be ignored.

[0064] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0065] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0066] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.

Claims

1. A method for evaluating offshore wind power output based on hybrid clustering and adaptive kernel density estimation, characterized in that, Includes the following steps: Step S1: Collect historical power output data of offshore wind power in multiple regions and process the data; Step S2: The optimal number of clusters k is determined using the Bayesian information criterion, and Gaussian mixture clustering is performed on the offshore wind power output curves to generate curve clusters that reflect different output modes. Step S3: Apply nonparametric kernel density estimation independently to each type of curve cluster, adaptively calculate the optimal bandwidth based on the criterion of minimizing the mean square error of the asymptotic integral, and dynamically adjust the bandwidth range through multimodality test. Step S4: Associate the clustering results with spatiotemporal labels, including seasonal and regional information, to construct a multi-dimensional power output probability density function. Based on the power output probability density function, calculate the confidence interval and upper and lower limits of power output fluctuation for offshore wind power, and assess the state of offshore wind power output.

2. The method for evaluating offshore wind power output based on hybrid clustering and adaptive kernel density estimation according to claim 1, characterized in that, Step S1 involves collecting historical offshore wind power output data from multiple regions and processing the data, specifically including the following steps: Based on the regional offshore wind power distribution and 8760 hours of wind measurement data from offshore wind towers, combined with sea location, distance from shore, and wind speed, the entire region's offshore wind power is divided into... N One region; By combining the offshore wind power scale of each wind zone within the sub-region and the original hourly offshore wind power output curves of each wind zone, the original hourly offshore wind power output curve of the sub-region can be obtained by weighted summation, as shown in the following formula: ; In the formula: A it This represents the original hourly offshore wind power output curve for wind zone i. m This represents the total number of wind zones within the current region; P it The original curve of offshore wind power output in wind zone i is t=1,2,...,8760; α i It represents the proportion of offshore wind power capacity in wind zone i to the total offshore wind power capacity in the current region; The power output sequences of offshore wind farms in all zones within the entire region are superimposed to obtain the power output curve of offshore wind power across the entire region.

3. The method for evaluating offshore wind power output based on hybrid clustering and adaptive kernel density estimation according to claim 1, characterized in that, In step S2, the optimal number of clusters k is determined using the Bayesian information criterion, as follows: If the error or disturbance of the Bayesian information criterion model follows a normal distribution, then the Bayesian information criterion model can be expressed as: ; in, k The number of model parameters represents the number of Gaussian distributions and the total number of parameters in the Gaussian mixture model. n The sample size represents the total number of offshore wind power output curves. L The maximum value of the model likelihood function represents the maximum likelihood estimate of the sample data given the model. kln(n) Penalty items, S RSS The sum of squared residuals represents the estimated model; By calculating different cluster numbers k Choose the BIC value that minimizes the BIC value. k As the optimal number of clusters.

4. The method for evaluating offshore wind power output based on hybrid clustering and adaptive kernel density estimation according to claim 3, characterized in that, Step S2 involves performing Gaussian mixture clustering on the offshore wind power output curves to generate curve clusters reflecting different output modes. Specifically, this includes: Assume the wind power output per hour of the day is , i =1,2,…,24, then the Gaussian mixture model is expressed as: ; In the formula: x It is a random variable; Let x be the probability of a random variable. It is a weighting coefficient, and satisfies ; Let be the distribution of the k-th Gaussian component in the Gaussian mixture model; The expectation-maximization algorithm is used to estimate the three parameters in the above Gaussian mixture model. The three parameters are the mean of the k-th Gaussian component. Weighting coefficient π k and variance ; Transform the above equation into: ; (1) First, specify 3 parameters. μ、 π ε The initial value; (2) Calculate the posterior probability The calculation method is as follows: ; (3) Based on the posterior probability Solve The maximum likelihood function is as follows: ; in, For the first n Observed values ​​of one sample of offshore wind power; N The total number of samples; N k For the first k The number of valid samples for each Gaussian distribution component; (4) Based on the posterior probability and beg The maximum likelihood value is as follows: ; (5) Solve The maximum likelihood function is as follows: ; (6) If satisfied , and ,in α 1. α 2 and α 3 is the convergence threshold, then μ ,π ε The values ​​are respectively taken μ (t) , π (t) , ε (t) Otherwise, Repeat steps (2) to (6) until the algorithm converges; superscript and These represent the EM algorithm at the 1st, 2nd, and 3rd respectively. t Second and third t The parameter estimates obtained from -1 iterations.

5. The method for evaluating offshore wind power output based on hybrid clustering and adaptive kernel density estimation according to claim 1, characterized in that, Step S2 further includes removing outliers using box plots and extracting representative feature curves for each type of curve cluster. The specific steps for removing outliers using box plots are as follows: (1) Plot a box plot for the output data of each type of curve cluster, and calculate the lower quartile Q1, upper quartile Q3, interquartile range IQR, upper edge and lower edge; where the interquartile range IQR = Q3 - Q1; upper edge = Q3 + 1.5 × IQR; lower edge = Q1 - 1.5 × IQR; (2) Identify data points located outside the upper and lower edges, mark them as outliers, remove outliers, and retain normal data points.

6. The method for evaluating offshore wind power output based on hybrid clustering and adaptive kernel density estimation according to claim 1, characterized in that, Step S3 involves independently applying nonparametric kernel density estimation to each type of curve cluster, adaptively calculating the optimal bandwidth based on the criterion of minimizing the asymptotic integral mean square error, and dynamically adjusting the bandwidth interval through a multimodality test. Specifically, this includes the following steps: (1) Assumption p 1, p 2,…, p n Data sample set for offshore wind power p If a Gaussian function is used as the probability density estimation model for offshore wind power output, then this sample set p The nonparametric KDE probability density function can be expressed as: ; In the formula, This represents the wind power probability density function based on the nonparametric KDE. The bandwidth parameter for kernel density estimation; This is the i-th sample value of the active power output of wind power. u i =( p - p i ) / l bw ; n The total number of samples for offshore wind power output data; Bias produced by fitting estimation B and variance V They are respectively: ; ; in: ; ; In the formula, For the true probability density function of offshore wind power output; It is the second derivative of the true probability density function; Represents a higher-order infinitesimal; For kernel functions; is the independent variable of the kernel function; (2) The optimal bandwidth is adaptively calculated based on the criterion of minimizing the asymptotic integral mean square error. The optimal bandwidth is obtained when the asymptotic integral mean square error is minimized. The calculation formula is: ; ; Optimal bandwidth is determined by applying the normal reference criterion. The calculation formula is further simplified to: ; in, The standard deviation of the sample data for offshore wind power output; Replace the above formula with the following formula. The details are as follows: ; in, The divergence measure is the half-range. Represents the standard normal cumulative distribution function; The formula for calculating the bandwidth of the optimal kernel function is: 。 7. The method for evaluating offshore wind power output based on hybrid clustering and adaptive kernel density estimation according to claim 1, characterized in that, Step S4 associates the clustering results with spatiotemporal labels, including seasonal and regional information, to construct a multi-dimensional power output probability density function. Based on the power output probability density function, it calculates the confidence interval and upper and lower limits of power output fluctuations for offshore wind power, and evaluates the state of offshore wind power output. Specifically, this includes the following steps: (1) Spatiotemporal tags include seasons S and region R The method for constructing a multi-dimensional output probability density function is as follows: calculate the probability density function of the label in a specific time and space. conditional probability density function : ; In the formula: K The total number of clusters; For the first k Curve clusters based on adaptive bandwidth The resulting kernel density estimate probability density function; For the season S and region R Under the conditions k The probability weight of the occurrence of the power output mode, and satisfying ; (2) Based on the power output probability density function, calculate the confidence interval of offshore wind power output and calculate the effective power output rate to evaluate the system operating status. The method for calculating the confidence interval is as follows: By analyzing the multidimensional power output probability density function Integrating yields the cumulative distribution function. Given a confidence level Calculate the upper and lower limits of force fluctuation. : ; In the formula, It is the inverse function of the cumulative distribution function; if the actual offshore wind power output exceeds this range, it is judged as an abnormal fluctuation. Calculate the effective power output of offshore wind power The formula is: ; In the formula, P WT This represents the actual power output of offshore wind power. P max The maximum power output of offshore wind power; based on the effective power output rate of offshore wind power. To assess the actual production efficiency of a wind farm, if If the value is below the preset threshold, it is identified as an inefficient operating period, and the operation strategy of the offshore wind turbine needs to be optimized. (3) Based on the power output probability density function, generate a typical scenario set for offshore wind power output and calculate the wind curtailment rate of offshore wind power for power system dispatch and planning; the calculation of the wind curtailment rate of offshore wind power The formula is: ; In the formula, For the curtailed wind power, based on The extent of insufficient grid absorption capacity is quantified, and areas with high wind curtailment rates are identified as key planning targets.

8. A system for evaluating the output of offshore wind power based on hybrid clustering and adaptive kernel density estimation, characterized in that, The method described by any one of claims 1 to 7 includes: The data acquisition and processing module is used to collect historical power output data of offshore wind power in multiple regions and process the data. The clustering module is used to determine the optimal number of clusters k using the Bayesian information criterion, and to perform Gaussian mixture clustering on the offshore wind power output curves to generate curve clusters that reflect different output modes. The bandwidth optimization module is used to independently apply nonparametric kernel density estimation to each type of curve cluster, adaptively calculate the optimal bandwidth based on the criterion of minimizing the asymptotic integral mean square error, and dynamically adjust the bandwidth range through multimodality test; The evaluation module is used to associate clustering results with spatiotemporal labels, including seasonal and regional information, to construct a multi-dimensional power output probability density function. Based on the power output probability density function, the confidence interval and upper and lower limits of power output fluctuation of offshore wind power are calculated to evaluate the state of offshore wind power output.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device containing the computer-readable storage medium to perform the offshore wind power output assessment method based on hybrid clustering and adaptive kernel density estimation as described in any one of claims 1 to 7.

10. A processor, characterized in that, The processor is used to run a program, wherein the program executes a method for evaluating offshore wind power output based on hybrid clustering and adaptive kernel density estimation as described in any one of claims 1 to 7.