Integrated modeling method and system considering multivariable wind power curve

The abnormal data is removed through iForest and k-medoids++ algorithms, combined with BP neural network and Adaboost integrated algorithm, and multivariate factors are considered to model wind power curves, solving the reliability and accuracy of the wind power power curve model, and achieving the accuracy of wind power optimization monitoring.

CN120297091APending Publication Date: 2025-07-11STATE GRID LIAONING ELECTRIC POWER CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311807624.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The intermittent and unpredictability of wind power generation pose challenges to the power grid, and the reliability and accuracy of existing wind power curve models are insufficient.

Method used

The iForest algorithm is used to remove abnormal data, the k-medoids++ algorithm is used for data clustering, and the wind power curve modeling is combined with the BP neural network and the Adaboost integrated algorithm. Taking into account multivariate factors such as wind speed, air pressure, temperature, humidity and wind direction, the model accuracy is improved through data filtering, clustering and integrated modeling.

Benefits of technology

A reliable and accurate wind power curve model was established, which overcomes the model's sensitivity to outliers and improves the accuracy and accuracy of wind power optimization monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297091A_ABST
    Figure CN120297091A_ABST
Patent Text Reader

Abstract

The invention discloses an integrated modeling method and system considering a multivariable wind power curve, belongs to the field of power systems, and considers the influence of multivariate meteorological data on different sections of the power curve. The method comprises the following steps: firstly, considering abnormal values of data acquired by an original wind driven generator, and filtering the abnormal values by adopting an isolation forest; and based on the difference of the wind data, a k-mediid + + algorithm is adopted to classify the wind data. Furthermore, a BP neural network is used as a weak learner for different features of data. Meanwhile, Adaboost boosting is used for carrying out fusion on the data to construct a more accurate model. According to the wind power curve modeling method, the influence of the multivariate meteorological data on different sections of the power curve is also considered besides the multivariate meteorological data, so that the modeling of the wind power curve is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of power systems, and particularly relates to a method for modeling wind power curves. Background Art

[0002] Wind energy has the characteristics of rich reserves and no pollution, and has been widely developed in response to global energy shortages and climate change. The Global Wind Energy Council estimates that 7100 GW of new wind turbine (WT) capacity will be installed annually in the future, and 3.55 GW will be installed by 2024. With the support of policies, the penetration rate of wind power is continuously increasing. However, the intermittency and unpredictability of wind power have brought unprecedented challenges to the power grid and attracted people's attention. Summary of the Invention

[0003] The present invention proposes an integrated modeling method and system considering multi-variable wind power curves to solve the reliability and accuracy problems of wind power curve models.

[0004] The present invention adopts the following technical solutions:

[0005] An integrated modeling method considering multi-variable wind power curves, comprising the following steps:

[0006] Step 1: Data filtering, calculating the average path length of wind speed and power sample points through the iForest algorithm to judge abnormal data, and removing the abnormal data;

[0007] Step 2: Clustering the data processed in Step 1, and using the k-medoids++ algorithm to perform a strong correlation division on the wind speed and power data;

[0008] Step 3: Curve modeling: Based on the data divided in Step 2, using a BP neural network as a base learner to model the power curve;

[0009] Step 4: Integration: Construct an Adaboost integration algorithm, dynamically weight the modeling results of the base learners, and further adjust the proportions of different base learners until the number of iterations or the loss function of the learner reaches a threshold.

[0010] Further, the calculation of the sample points is as follows:

[0011]

[0012] Where c(x) is the average value of n data samples h(x), and E(h(x)) is the average path length of all trees in the sample. If s≈1, then x is an outlier; if s << 0.5, then x is a normal value.

[0013] Further, the data is divided into clusters by the k - mediid++ algorithm. The method for dividing the number of clusters is as follows:

[0014]

[0015]

[0016] d(x a ,x b )=||x a -x b || 2

[0017] where k is the number of clusters; r represents the sample points in the cluster C p ; o j represents the center in the cluster; P r is the probability distribution; d(·) is the distance from r to o j .

[0018] Further, step 3 specifically includes:

[0019] Assume there are data points (x1, y1), (x2, y2),…,(x m ,y m ),…,(x N ,y N ). The BP neural network is modeled as:

[0020] H m =z(W h x m +B h )

[0021] O m =W o H m +B o

[0022]

[0023] where H m is the output of the hidden layer, O m is the output of the neural network, W is the weight, B is the bias, z(·) is the activation function, J is the loss function, and λ is the loss function coefficient.

[0024] Further, the steps of the Adaboost ensemble algorithm are as follows

[0025] (1) Initialize the weights of the dataset The weights are evenly distributed, and the calculation formula is:

[0026]

[0027] where N is the total number of samples in the dataset pth, and the number of iterations is j = 1, 2, …, J;

[0028] (2) Calculate the error weights

[0029]

[0030] where is the weight of the mth sample in the jth iteration, and O p,j is the output of the neural network;

[0031] (3) Adjust the weights according to the error:

[0032]

[0033]

[0034] W m,j+1 = W m,j exp[α p,j * I(y m,p ≠ O p,j (x m,p )]

[0035] where I(·) is the indicator function;

[0036] (4) Combine the prediction result and the weights as the output of the model.

[0037]

[0038] (5) Repeat steps (1)-(4), and the clustering data of different clusters are repeatedly calculated to obtain the final power output model.

[0039] An integrated modeling system considering multi-variable wind power curves, comprising:

[0040] A data filtering module that judges abnormal data by calculating the average path length of wind speed and power sample points through the iForest algorithm and removes the abnormal data;

[0041] A clustering module that clusters the data processed by the data filtering module and uses the k-medoids++ algorithm to perform a strong correlation division on the wind speed and power data;

[0042] A curve modeling module: Based on the data divided by the clustering module, a BP neural network is used as a base learner to model the power curve;

[0043] Integrated module: Construct the Adaboost integrated algorithm to dynamically weight the modeling results of the base learners, and further adjust the proportions of different base learners until the number of iterations or the loss function of the learners reaches the threshold.

[0044] Further, the data filtering module calculates the sample point as:

[0045]

[0046] Where c(x) is the average value of n data samples h(x), and E(h(x)) is the average path length of all trees in the sample. If s≈1, then x is an outlier; if s<<0.5, then x is a normal value.

[0047] Further, the clustering module divides the data into clusters through the k - mediid++ algorithm. The method for dividing the number of clusters is as follows:

[0048]

[0049]

[0050] d(x a ,x b )=||x a -x b || 2

[0051] Where k is the number of clusters; r represents the sample points in the cluster C p ; o j represents the center in the cluster; P r is the probability distribution; d(·) is the distance from r to o j .

[0052] Further, the curve modeling module is modeled as:

[0053] Assume there are data points (x1,y1),(x2,y2),…,(x m ,y m ),…,(x N ,y N ). The BP neural network is modeled as:

[0054] H m =z(W h x m +B h )

[0055] O m =W o H m +B o

[0056]

[0057] Among them, H m is the output of the hidden layer, O m is the output of the neural network, W is the weight, B is the bias, z(·) is the activation function, J is the loss function, and λ is the loss function coefficient.

[0058] Furthermore, the process of integrating all the integration modules through the Adaboost integration algorithm is as follows:

[0059] (1) Initialize the weights of the dataset The weights are evenly distributed, and the calculation formula is:

[0060]

[0061] where N is the total number of samples in the dataset pth, and the number of iterations is j = 1, 2,..., J;

[0062] (2) Calculate the error weights

[0063]

[0064] where is the weight of the mth sample in the jth iteration, and O p,j is the output of the neural network;

[0065] (3) Adjust the weights according to the error:

[0066]

[0067]

[0068] W m,j+1 = W m,j exp[α p,j *I(y m,p ≠O p,j (x m,p )]

[0069] where I(·) is the indicator function;

[0070] (4) Combine the prediction result with the weights as the output of the model.

[0071]

[0072] (5) Repeat steps (1)-(4), and the clustering data of different clusters are repeatedly calculated to obtain the final power output model.

[0073] As can be seen from the above technical solutions, the beneficial effects of the present invention are as follows: An integrated modeling method and system for considering multi-variable wind power curves provided by the present invention takes into account the differences in the output power of wind turbines under different influencing factors. The present invention patent establishes an integrated modeling technology to achieve power evaluation based on multiple influencing factors. In order to overcome the sensitivity of the model to outliers, a data filtering technology is added to improve the accuracy of the model. In addition, the k-medoids++ algorithm is used to segment the data to theoretically describe the data differences. Finally, the Adaboost integration algorithm based on the BP neural network weak learner is used to model the wind power curve for different clustering data. The established reliable and accurate wind power curve model is of great significance for wind power optimization monitoring.

[0074] In addition to considering multivariate meteorological data, the present invention also considers the influence of multivariate meteorological data on different sections of the power curve, making the modeling of the wind power curve more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] Figure 1 Schematic diagram of the BP neural network weak learner provided by an embodiment of the present invention;

[0076] Figure 2 Framework of the power curve modeling provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0077] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0078] See Figure 2 As shown, an integrated modeling method for considering multi-variable wind power curves includes the following steps:

[0079] Step 1: Data filtering, calculating the average path length of the wind speed and power sample points through the iForest algorithm to judge abnormal data and removing the abnormal data. The data includes wind speed, air pressure, temperature, humidity and wind direction data;

[0080] Step 2: Clustering the data processed in Step 1, using the k-medoids++ algorithm to perform a strong correlation division on the wind speed and power data to improve the similarity within the clusters and reduce the similarity between the clusters, providing a good data set for power curve modeling;

[0081] Step 3: Curve modeling: Based on the data divided in Step 2, using the BP neural network as the base learner to model the power curve;

[0082] Step 4: Integration: Construct the Adaboost integration algorithm to dynamically weight the modeling results of the base learners, and further adjust the proportions of different base learners until the number of iterations or the loss function of the learner reaches the threshold.

[0083] The present invention determines the influencing factors of the wind power curve model.

[0084] Without considering the power conversion loss, the theoretical power output curve can usually be expressed as:

[0085]

[0086] Where C is p The power coefficient. ρ represents the air density. A is the blade area. v is the wind speed.

[0087] Combined with the above formula, it can be seen that the wind speed and the output power have a cubic relationship, and the output power is the main dependent variable. The power coefficient is a function of the blade pitch angle, and the power output is adjusted by controlling the blade pitch angle through the turbine controller. The density of the air is affected by the comprehensive influence of pressure, temperature and humidity. Among them, the change range of humidity is small, within 2%. Due to topographical factors, the atmospheric pressure changes by 10%; however, the temperature change range is very large, up to 20%. Compared with air pressure and humidity, temperature has the greatest influence on air density. The blade area is mainly affected by the wind direction. Through the above analysis, air pressure, temperature, humidity and wind direction are the main factors affecting the power output. Therefore, different from other studies, the present invention comprehensively considers factors such as wind speed, air pressure, temperature, humidity and wind direction for power curve modeling.

[0088] The calculation of the sample points is as follows:

[0089]

[0090] In the formula, c(x) is the average value in the n data samples h(x), and E(h(x)) is the average path length of all the trees in the sample. If s≈1, then x is an outlier, and if s<<0.5, then x is a normal value.

[0091] Data processing is performed based on historical power data.

[0092] In power modeling, outliers have a great impact on the modeling accuracy. iForest is an efficient data cleaning method that uses a collection of trees to identify outliers. Its main idea is that the characteristics of outliers are different from those of normal data, and outliers are the easiest to be isolated. In iForest, normal data (clusters with higher density) require more splitting times, while outliers (clusters with lower density) can be separated with fewer splitting times.

[0093] The iForest can determine whether it is an anomaly by calculating the average path length of the sample points. The calculation of the sample points in the sample is as follows

[0094]

[0095] Where c(x) is the average value in the n data samples h(x), and E(h(x)) is the average path length of all the trees in the sample. If s≈1, then x is an outlier; if s<<0.5, then x is a normal value.

[0096] Use k-medoids++ to perform a strong correlation division on the wind speed and power data.

[0097] In the process of clustering the power curve data, use k-medoids++ to perform a high correlation division on the data. In order to improve the similarity within the clusters and reduce the similarity between the clusters, the k-mediid++ algorithm divides the data into clusters. Different from the selection of the centroid in the k-medoids algorithm, k-medoids++ calculates the probability of possible centroids as the representative objects to determine the dissimilarity of each clustering. In the k-medoids++ algorithm, the division process iterates until the absolute error function is minimized. This patent only divides the wind speed and output power and does not consider other variables. The method of dividing the number of clusters using k-medoids++ is as follows.

[0098]

[0099]

[0100] d(x a ,x b )=||x a -x b || 2

[0101] Where k is the number of clusters; r represents the sample points in the cluster C p ; o j represents the center in the cluster; P r is the probability distribution; d(·) is the distance from r to o j .

[0102] Based on the data divided in step 3, use the BP neural network as the base learner to model the power curve.

[0103] Assume that there are data points (x1, y1), (x2, y2),…, (x m ,y m ),…, (x N ,y N ). The BP neural network modeling is as follows:

[0104] H m = z(W h x m + B h )

[0105] O m = W o H m + B o

[0106]

[0107] where H m is the output of the hidden layer, O m is the output of the neural network, W is the weight, and B is the bias. z(·)

[0108] is the activation function, J is the loss function, and λ is the loss function coefficient. The neural network model is as Figure 1 shown.

[0109] Construct the Adaboost ensemble algorithm to dynamically weight the modeling results of the base learners. The results of the ensemble model are as Figure 2 shown.

[0110] The Adaboost ensemble algorithm combines multiple machine learning models to ensure the minimum prediction model error. During the boosting process, the prediction results are dynamically weighted to further adjust the proportion of different weak learners until the number of iterations or the loss function of the learner reaches the threshold.

[0111] Based on the k - mediid++ algorithm, partition the dataset. The dataset is (x p1 , y p1 ), (x p2 , y p2 ), …, (x pm , y pm ), …, (x pN , y pN ) and is used to train the BP weak learner. The steps of the Adaboost ensemble algorithm are as follows

[0112] (1) Initialize the weights of the dataset The weights are uniformly distributed, and the calculation formula is

[0113]

[0114] where N is the total number of samples in the dataset pth. The number of iterations is j = 1, 2, …, J.

[0115] (2) Calculate the error weights

[0116]

[0117] Among them is the weight of the mth sample in the jth iteration. O p,j is the output of the neural network.

[0118] (3) Adjust the weights according to the error

[0119]

[0120]

[0121] W m,j+1 = W m,j exp[α p,j * I(y m,p ≠ O p,j (x m,p )]

[0122] where I(·) is the indicator function.

[0123] (4) Combine the prediction result and the weight as the output of the model.

[0124]

[0125] (5) Repeat steps (1)-(4), and the clustering data of different clusters are repeatedly calculated to obtain the final power output model.

[0126] Step 6: Conduct model evaluation.

[0127] To evaluate the accuracy of the constructed model, three different statistics are used, namely the sum of squared errors (SSE), the root mean square error (RMSE), and the coefficient of determination (R2):

[0128]

[0129]

[0130]

[0131] where y m is the actual value of the data, O m is the output of the model, is the average value of all actual values.

[0132] An integrated modeling system considering multi-variable wind power curves, comprising:

[0133] A data filtering module that judges abnormal data by calculating the average path length of wind speed and power sample points through the iForest algorithm and removes the abnormal data;

[0134] The clustering module clusters the data processed by the data filtering module, and uses the k-medoids++ algorithm to perform a strong correlation division on the data of wind speed and power;

[0135] Curve modeling module: Based on the data divided by the clustering module, the BP neural network is used as the base learner to model the power curve;

[0136] Ensemble module: Construct the Adaboost ensemble algorithm, dynamically weight the modeling results of the base learners, and further adjust the proportion of different base learners until the number of iterations or the loss function of the learner reaches the threshold.

[0137] The data filtering module calculates the sample point as:

[0138]

[0139] Where c(x) is the average value of n data samples h(x), and E(h(x)) is the average path length of all trees in the sample. If s≈1, then x is an outlier; if s<<0.5, then x is a normal value.

[0140] Further, the clustering module divides the data into clusters through the k-mediid++ algorithm, and the method for dividing the number of clusters is as follows:

[0141]

[0142]

[0143] d(x a ,x b )=||x a -x b || 2

[0144] Where k is the number of clusters; r represents the sample points in the cluster C p ; o j represents the center in the cluster; P r is the probability distribution; d(·) is the distance from r to o j .

[0145] The curve modeling module is modeled as:

[0146] Assume that there are data points (x1,y1), (x2,y2),…,(x m ,y m ),…,(x N ,y N ), and the BP neural network is modeled as:

[0147] H m =z(W h xm +B h )

[0148] O m =W o H m +B o

[0149]

[0150] where H m is the output of the hidden layer, O m is the output of the neural network, W is the weight, B is the bias, z(·) is the activation function, J is the loss function, and λ is the loss function coefficient.

[0151] The process of integrating the integrated modules through the Adaboost integration algorithm is as follows:

[0152] (1) Initialize the weights of the dataset The weights are evenly distributed, and the calculation formula is:

[0153]

[0154] where N is the total number of samples in the dataset pth, and the number of iterations is j = 1, 2,..., J;

[0155] (2) Calculate the error weights

[0156]

[0157] where is the weight of the mth sample in the jth iteration, and O p,j is the output of the neural network;

[0158] (3) Adjust the weights according to the error:

[0159]

[0160]

[0161] W m,j+1 =W m,j exp[α p,j *I(y m,p ≠O p,j (x m,p )]

[0162] where I(·) is the indicator function;

[0163] (4) Combine the prediction results with the weights as the output of the model.

[0164]

[0165] (5) Repeat steps (1)-(4), and the clustering data of different clusters are repeatedly calculated to obtain the final power output model.

[0166] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An integrated modeling method considering multi-variable wind power curves, characterized in that, It includes the following steps: Step 1: Data filtering. Calculate the average path length of wind speed and power sample points through the iForest algorithm to judge abnormal data and remove the abnormal data; Step 2: Cluster the data processed in Step 1. Use the k-medoids++ algorithm to perform a strong correlation division on the wind speed and power data; Step 3: Curve modeling: Based on the data divided in Step 2, use the BP neural network as the base learner to model the power curve; Step 4: Integration: Construct the Adaboost integration algorithm, dynamically weight the modeling results of the base learners, and further adjust the proportions of different base learners until the number of iterations or the loss function of the learner reaches the threshold.

2. The integrated modeling method for considering multi-variable wind power curves according to claim 1, characterized in that The calculation of the sample point is as follows: Where c(x) is the average value among n data samples h(x), and E(h(x)) is the average path length of all trees in the sample. If s≈1, then x is an outlier; if s<<0.5, then x is a normal value.

3. The integrated modeling method considering the multi-variable wind power curve according to claim 1, characterized in that, Divide the data into clusters through the k-mediid++ algorithm. The method for dividing the number of clusters is as follows: d(x a ,x b )=||x a -x b || 2 where k is the number of clusters; r represents the sample points in the cluster C p ; o j represents the center in the cluster; P r is the probability distribution; d(·) is the distance from r to o j .

4. The integrated modeling method for considering multi-variable wind power curves according to claim 1, characterized in that Step 3 specifically includes: Assume there are data points (x1, y1), (x2, y2), …, (x m , y m ), …, (x N , y N ), and the BP neural network model is as follows: H m = z(W h x m + B h ) O m = W o H m + B o where H m is the output of the hidden layer, O m is the output of the neural network, W is the weight, B is the bias, z(·) is the activation function, J is the loss function, and λ is the loss function coefficient.

5. The integrated modeling method considering the multi-variable wind power curve according to claim 1, characterized in that, The steps of the Adaboost integration algorithm are as follows (1) Initialize the weights of the dataset The weights are evenly distributed, and the calculation formula is: Where N is the total number of samples in the dataset pth, and the number of iterations is j = 1, 2, …, J; (2) Calculate the error weights where is the weight of the mth sample in the jth iteration, and O p,j is the output of the neural network; (3) Adjust the weights according to the error: W m,j+1 = W m,j exp[α p,j * I(y m,p ≠ O p,j (x m,p )] Where I(·) is the indicator function; (4) Combine the prediction result and the weights as the output of the model. (5) Repeat steps (1)-(4), and the clustering data of different clusters is repeatedly calculated to obtain the final power output model.

6. An integrated modeling system considering multi-variable wind power curves, characterized in that, It includes: A data filtering module that calculates the average path length of wind speed and power sample points through the iForest algorithm to judge abnormal data and remove the abnormal data; A clustering module that clusters the data processed by the data filtering module and uses the k-medoids++ algorithm to perform a strong correlation division on the wind speed and power data; A curve modeling module: Based on the data divided by the clustering module, use the BP neural network as the base learner to model the power curve; An integration module: Construct the Adaboost integration algorithm, dynamically weight the modeling results of the base learners, and further adjust the proportions of different base learners until the number of iterations or the loss function of the learner reaches the threshold.

7. The integrated modeling system for considering a multi-variable wind power curve according to claim 1, characterized in that, The data filtering module calculates the sample point as: Where c(x) is the average value among n data samples h(x), and E(h(x)) is the average path length of all trees in the sample. If s≈1, then x is an outlier; if s<<0.5, then x is a normal value.

8. The integrated modeling system for considering multi-variable wind power curves according to claim 1, characterized in that, The clustering module divides the data into clusters through the k-mediid++ algorithm. The method for dividing the number of clusters is as follows: d(x a ,x b )=||x a -x b || 2 where k is the number of clusters; r represents a sample point in the cluster C p ; o j represents the center in the cluster; P r is the probability distribution; d(·) is the distance from r to o j .

9. The integrated modeling system for considering multi-variable wind power curves according to claim 1, characterized in that, The curve modeling module models as: Suppose there are data points (x1, y1), (x2, y2), …, (x m , y m ), …, (x N , y N ), and the BP neural network model is as follows: H m = z(W h x m + B h ) O m = W o H m + B o Among them, H m is the output of the hidden layer, O m is the output of the neural network, W is the weight, B is the bias, z(·) is the activation function, J is the loss function, and λ is the loss function coefficient.

10. The integrated modeling method considering a multi-variable wind power curve according to claim 1, characterized in that The process of the integration module integrating through the Adaboost integration algorithm is: (1) Initialize the weights of the dataset The weights are evenly distributed, and the calculation formula is: Where N is the total number of samples in the dataset pth, and the number of iterations is j = 1, 2, …, J; (2) Calculate the error weights where is the weight of the mth sample in the jth iteration, and O p,j is the output of the neural network; (3) Adjust the weights according to the error: W m,j+1 = W m,j exp[α p,j * I(y m,p ≠ O p,j (x m,p )] Where I(·) is the indicator function; (4) Combine the prediction result and the weights as the output of the model. (5) Repeat steps (1)-(4), and the clustering data of different clusters are repeatedly calculated to obtain the final power output model.