Wind turbine blade icing prediction method, device, platform and storage medium

By combining the parallel coordinate method with recurrent neural networks, the accuracy problem of the wind turbine blade icing prediction model was solved, achieving more accurate icing prediction and processing.

CN116928042BActive Publication Date: 2025-09-16广东省工业边缘智能创新中心有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310365369.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-03
Publication Date
2025-09-16
Estimated Expiration
2043-04-03

AI Technical Summary

Technical Problem

Existing technology makes it difficult to build a high-precision wind turbine blade icing prediction model, mainly because the characteristic data of sample wind turbines are rich in content, uneven in quality and complex in structure, which makes data processing and calculation difficult and makes it impossible to accurately find features related to icing.

Method used

The parallel coordinate method is used to convert the complex first feature data into visual second feature data. By determining the first cluster center and performing clustering, the first and second clustering results are obtained. Combined with recurrent neural network modeling, the third feature data related to icing is extracted.

Benefits of technology

The accuracy of the wind turbine blade icing prediction model is improved, making the prediction of the blade icing condition more accurate and better able to prevent or deal with the blade icing problem.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116928042B_ABST
    Figure CN116928042B_ABST
Patent Text Reader

Abstract

The embodiments of the present application relate to the technical field of wind turbine blade icing prediction, and specifically to a wind turbine blade icing prediction method, apparatus, modeling platform, and storage medium. The method comprises: obtaining first feature data of a sample wind turbine, processing the first feature data according to the parallel coordinate method to obtain second feature data, determining a first cluster center and clustering the second feature data to obtain a first cluster result, determining a second cluster center to re-cluster the second feature data to obtain a second cluster result, extracting third feature data from the second cluster result as a training sample to train a wind turbine blade icing prediction model, and predicting the wind turbine blade to be predicted based on the icing prediction model. Through the above steps, the third feature data used for training has a high correlation with whether the wind turbine blade is iced or not, so that the trained wind turbine blade icing prediction model makes predictions more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of wind turbine blade icing prediction, and specifically to a wind turbine blade icing prediction method, device, modeling platform, and storage medium. Background Art

[0002] During actual operation, wind turbine blades can frost and cause failure. To address blade icing, it's necessary to identify the characteristics associated with blade icing—those that cause it. Because there are multiple factors that can cause blade icing, such as ambient temperature and blade speed, it's necessary to assess these characteristics and identify the most significant ones. Subsequent actions, whether preventing blade icing or addressing already frozen blades, can be guided by these key characteristics.

[0003] Wind turbine blade icing prediction models can be constructed to predict blade icing, thereby identifying features related to blade icing and helping users resolve blade icing issues. However, the feature data of the sample wind turbines used to build the model is rich in content, has varying data quality, and is extremely complex. This makes processing and calculation of the sample wind turbine data difficult, making it difficult to construct a highly accurate wind turbine blade icing prediction model to predict blade icing. The features that are identified are also inaccurate. Therefore, the challenge of building a highly accurate wind turbine blade icing prediction model remains. Summary of the Invention

[0004] In view of the above problems, an embodiment of the present application provides a method for predicting icing of wind turbine blades, which is used to solve the problem of how to construct a wind turbine blade icing prediction model with high accuracy.

[0005] According to a first aspect of an embodiment of the present application, a method for predicting icing of wind turbine blades is provided, which is applied to a modeling platform. The method includes:

[0006] Acquire first characteristic data of a sample wind turbine, each of the first characteristic data including at least two characteristics and data corresponding to the characteristics;

[0007] Processing the first feature data according to a parallel coordinate method to obtain second feature data, each second feature data including only one feature of the first feature data and data corresponding to the feature;

[0008] Determining a first cluster center for each of the features based on the second feature data, wherein each feature includes at least one first cluster center, and the first cluster center is used to cluster the second feature data on the feature where the first cluster center is located;

[0009] Calculating a first Euclidean distance between each second feature data and each first cluster center on a feature included in the second feature data, and determining the first cluster center having the shortest first Euclidean distance to the second feature data;

[0010] Dividing the second feature data into the first cluster center having the shortest first Euclidean distance to the second feature data, to obtain a first clustering result;

[0011] Calculating, for each of the first cluster centers, a mean value of the second feature data of each of the first cluster centers according to the first clustering result as a second cluster center;

[0012] Calculating a second Euclidean distance from each second feature data to each second cluster center on a feature included in the second feature data, and determining the second cluster center having the shortest second Euclidean distance to the second feature data;

[0013] Dividing the second characteristic data into the second cluster center having the shortest second Euclidean distance to the second characteristic data, to obtain a second clustering result;

[0014] extracting third feature data according to the second clustering result;

[0015] Modeling the third characteristic data using a recurrent neural network to obtain a wind turbine blade icing prediction model;

[0016] The wind turbine blade icing condition to be predicted is predicted according to the wind turbine blade icing prediction model.

[0017] In an optional manner, after the first feature data is processed according to the parallel coordinate method to obtain the second feature data, the method further includes:

[0018] The second feature data is subjected to missing value filling, outlier removal and standardization according to the box plot method.

[0019] In an optional manner, the first characteristic data includes at least one characteristic of wind speed, wind direction, average wind direction, ambient temperature, time, generator speed, pitch motor temperature, main shaft speed and blade speed.

[0020] In an optional manner, calculating the first Euclidean distance between each second feature data and each first cluster center on the feature included in the second feature data, and determining the first cluster center having the shortest first Euclidean distance to the second feature data, includes:

[0021] According to the formula Calculate the shortest first Euclidean distance from the second feature data to each first cluster center on the feature corresponding to the second feature data, where dist(x m ,c n ) is the shortest first Euclidean distance from the second feature data to each first cluster center on the feature corresponding to the second feature data, X includes all the second feature data, x m is the mth second feature data, C is the set of the first cluster centers, c n is x m The first cluster center on the corresponding feature.

[0022] In an optional manner, calculating, for each first cluster center according to the first clustering result, a mean value of the second feature data of each first cluster center as a second cluster center includes:

[0023] According to the formula The second cluster center is calculated, where c i2 is the i-th second cluster center, N is the number of the second feature data divided into the i-th first cluster center, x m is the mth second feature data, c i1 is the i-th first cluster center.

[0024] In an optional manner, after dividing the second feature data into the second cluster centers having the shortest second Euclidean distance to the second feature data to obtain a second clustering result, the method further includes:

[0025] According to the formula Calculate the error sum of squares of the second clustering result, where E is the error sum of squares, k is the number of the second cluster centers, and c i2 is the second cluster center of the i-th cluster, dist(x m , c i2 ) is x m To the second cluster center c where it is located i2 The Euclidean distance, x m is the mth second feature data;

[0026] Determining whether the sum of squared errors of the second clustering result is less than a first threshold;

[0027] If yes, output the second clustering result;

[0028] If not, the second clustering result is used as the first clustering result, and the process proceeds to the step of calculating the mean of the second feature data of each first cluster center as the second cluster center for each first cluster center according to the first clustering result.

[0029] In an optional manner, the second feature data has an icing label or a non-icing label, and extracting the third feature data according to the second clustering result includes:

[0030] For each of the second cluster centers, determining a first number of the second feature data having an icing label and a second number of the second feature data having a non-icing label;

[0031] Calculating a first ratio of the first number to the total number of the second characteristic data of the second cluster center, and a second ratio of the second number to the total number of the second characteristic data of the second cluster center;

[0032] Determining whether a ratio between the first ratio and the second ratio is greater than or equal to a second threshold;

[0033] If so, extract the second feature data corresponding to the ratio in the second cluster center as the third feature data.

[0034] According to a second aspect of an embodiment of the present application, a wind turbine blade icing prediction device is provided, which is applied to a modeling platform. The device includes:

[0035] A first acquisition module is used to acquire first characteristic data of a sample wind turbine, each of which includes at least two characteristics and data corresponding to the characteristics;

[0036] A first processing module is configured to process the first feature data according to a parallel coordinate method to obtain second feature data, wherein each second feature data includes only one feature of the first feature data and data corresponding to the feature;

[0037] A first determining module is configured to determine a first cluster center of each feature according to the second feature data, wherein each feature includes at least one first cluster center, and the first cluster center is used to cluster the second feature data on the feature where the first cluster center is located;

[0038] A first calculation module is used to calculate a first Euclidean distance between each second feature data and each first cluster center on the feature included in the second feature data, and determine the first cluster center having the shortest first Euclidean distance to the second feature data;

[0039] A second processing module is configured to divide the second feature data into the first cluster center having the shortest first Euclidean distance to the second feature data, to obtain a first clustering result;

[0040] A second determining module is configured to calculate, for each of the first cluster centers, a mean value of the second feature data of each of the first cluster centers according to the first clustering result as a second cluster center;

[0041] A second calculation module is configured to calculate a second Euclidean distance between each second feature data and each second cluster center on a feature included in the second feature data, and determine the second cluster center having the shortest second Euclidean distance to the second feature data;

[0042] A data partitioning module: configured to partition the second characteristic data into the second cluster center having the shortest second Euclidean distance to the second characteristic data, to obtain a second clustering result;

[0043] Data extraction module: used for extracting third feature data according to the second clustering result;

[0044] A data modeling module is configured to use a recurrent neural network to model the third characteristic data to obtain an icing prediction model for wind turbine blades;

[0045] Prediction module: used for predicting the icing condition of the wind turbine blade to be predicted according to the wind turbine blade icing prediction model.

[0046] According to a third aspect of an embodiment of the present application, a modeling platform is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus;

[0047] The memory is used to store at least one executable instruction, and the executable instruction enables the processor to execute the operation of the wind turbine blade icing prediction method as described in any of the above embodiments.

[0048] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, wherein the storage medium stores at least one executable instruction, and the executable instruction, when run, executes the operation of the wind turbine blade icing prediction method as described in any of the above embodiments.

[0049] In an embodiment of the present application, the complex and diverse first feature data is converted into visual second feature data by the parallel coordinate method, thereby determining the first cluster center for subsequent clustering based on the first cluster center to obtain the first cluster result. The second cluster center is determined based on the first cluster result, and the second cluster result is obtained by clustering the second feature data multiple times, so that the correlation between the second feature data of the second cluster centers is higher. By judging the degree of correlation between the second feature data of each second cluster center and icing or non-icing from the second cluster result, the second feature data with an icing label or a non-icing label of the second cluster center with a high degree of correlation with icing or non-icing is used as the third feature data for subsequent modeling. Modeling is performed based on the third feature data with a greater correlation with icing or non-icing, and the resulting wind turbine blade icing prediction model is more accurate, so that the icing condition of the wind turbine blade to be predicted can be predicted more accurately in the end.

[0050] The above description is only an overview of the technical solutions of the embodiments of the present application. In order to more clearly understand the technical means of the embodiments of the present application, they can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the embodiments of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The accompanying drawings are only used to illustrate the embodiments and are not to be considered as limiting the present application. In addition, the same reference symbols are used to represent the same components throughout the drawings. In the drawings:

[0052] Figure 1 A schematic diagram of a process for predicting icing of wind turbine blades provided in an embodiment of the present application is shown;

[0053] Figure 2 A schematic diagram showing second characteristic data obtained by the parallel coordinate method according to an embodiment of the present application is shown;

[0054] Figure 3 A schematic diagram of a first cluster center obtained according to second characteristic data provided in an embodiment of the present application is shown;

[0055] Figure 4 A schematic diagram of a first cluster center obtained after dividing the second feature data provided in an embodiment of the present application is shown;

[0056] Figure 5 A schematic diagram showing a second clustering result provided in an embodiment of the present application is shown;

[0057] Figure 6 A schematic flow chart of a method for predicting icing of wind turbine blades provided in another embodiment of the present application is shown;

[0058] Figure 7 A schematic diagram of the structure of a wind turbine blade icing prediction device provided in an embodiment of the present application is shown;

[0059] Figure 8 A structural diagram of the modeling platform provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0060] The exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited to the embodiments set forth herein.

[0061] The inventors have noted that during actual operation, wind turbine blades can malfunction due to blade icing. Icing on wind turbine blades can change the blade shape and disrupt the aerodynamic properties of the blades, leading to reduced wind turbine efficiency and unstable operation, which in turn affects the stable operation of the power grid. To address the issue of blade icing, it is necessary to identify the characteristics associated with blade icing, that is, the characteristics that cause blade icing. Because there are multiple characteristics that can cause blade icing in wind turbines, such as ambient temperature and blade speed, it is necessary to assess the various characteristics of the wind turbine and identify those associated with blade icing.

[0062] Typically, wind turbine blade icing prediction models are constructed to predict blade icing and identify features related to blade icing. However, the feature data from the sample wind turbines used to build the model is rich in content, has varying quality, and has an extremely complex structure. This makes processing and calculation of the sample wind turbine data difficult, making it difficult to construct a highly accurate wind turbine blade icing prediction model to predict icing on wind turbine blades. Consequently, the features identified are inaccurate. Therefore, developing a highly accurate wind turbine blade icing prediction model has become a challenge.

[0063] After in-depth research, the inventors designed a method for predicting wind turbine blade icing. The method processes sample data used to construct a wind turbine blade icing prediction model, and extracts data related to whether the sample wind turbine blades are frozen or not. Furthermore, based on the above data, a more accurate wind turbine blade icing prediction model can be obtained by modeling through a recurrent neural network, which makes the blade icing prediction more accurate.

[0064] To extract more accurate data related to whether the sample wind turbine blades are iced or not from the sample wind turbine data, the k-means algorithm can be used to mine the sample data. However, the k-means algorithm requires the input of the number of expected cluster centers during its first operation, and the final clustering result of the k-means algorithm is closely related to the input number of expected cluster centers. Therefore, it is necessary to determine an appropriate number of expected cluster centers; otherwise, the output clustering results may not meet the requirements. However, faced with sample data with rich content, uneven data quality, and extremely complex data structure, it is impossible to directly obtain the number of expected cluster centers. Therefore, it is necessary to reduce the dimensionality of the sample data. The multidimensional and complex sample data is converted to a two-dimensional plane using the parallel coordinate method, making the sample data visual. This allows the required number of expected cluster centers to be intuitively obtained through the parallel coordinate plot.

[0065] Figure 1 A flow chart of the wind turbine blade icing prediction method provided by an embodiment of the present application is shown. The method is executed by a modeling platform, which may be a modeling platform including one or more processors, which may be a central processing unit (CPU), or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application, which are not limited here. The one or more processors included in the modeling platform may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs, which are not limited here. The modeling platform may be a server or a server cluster. Figure 1 As shown, according to the first aspect of the embodiment of the present application, the method includes the following steps:

[0066] Step 101: Acquire first characteristic data of a sample wind turbine, where each first characteristic data includes at least two characteristics and data corresponding to the characteristics.

[0067] The sample wind turbines are the source of sample data for training the wind turbine blade icing prediction model in the embodiment of the present application. As the sample wind turbines operate, some of the sample wind turbine blades are frozen, while some of the sample wind turbine blades are not. Based on whether the sample wind turbine blades are frozen, each first feature data is labeled. The first feature data of the sample wind turbine with frozen blades is labeled as frozen, and the first feature data of the sample wind turbine with unfrozen blades is labeled as unfrozen.

[0068] Sensors are installed on the sample wind turbines to obtain first characteristic data of the sample wind turbines. Multiple sensors are provided, each configured to collect different characteristics of the sample wind turbines. The sensors collect wind turbine data at a set frequency (e.g., once per second). Each sensor collects data once to generate one piece of first characteristic data. All sensors on the same sample wind turbine collect data simultaneously to generate the first characteristic data.

[0069] The first characteristic data includes at least two characteristics and data corresponding to the characteristics. In some embodiments, the first characteristic data includes at least one characteristic selected from the group consisting of wind speed, wind direction, average wind direction, ambient temperature, time, generator speed, pitch motor temperature, main shaft speed, and blade speed. Among them, wind speed, wind direction, average wind direction, ambient temperature, and time are environmental characteristics, generator speed, pitch motor temperature, and main shaft speed are generator characteristics, and blade speed is a blade characteristic. For example, at a certain moment in time, the sample wind turbine has a wind speed of 10 m / s, a pitch motor temperature of 15°C, and an ambient temperature of 10°C. At this moment, the first characteristic data acquired by the sensor is expressed as [10 m / s, 15°C, 10°C].

[0070] Step 102: Process the first feature data according to the parallel coordinate method to obtain second feature data, where each second feature data only includes one feature of the first feature data and data corresponding to the feature.

[0071] Step 103: Determine the first cluster center of each feature based on the second feature data. Each feature includes at least one first cluster center. The first cluster center is used to cluster the second feature data on the feature where it is located.

[0072] In steps 102 to 103, since the first characteristic data of the wind turbine blades has a large amount of data and a complex type, it is impossible to directly observe the patterns of the data (i.e., directly determine the important features related to icing of the wind turbine blades, hereinafter referred to as important features) when facing the first characteristic data, and the first characteristic data needs to be processed to obtain the important features.

[0073] In order to extract important features from the first feature data, the embodiment of the present application adopts the k-means algorithm to perform data mining on the first feature data, and clusters the data multiple times through the k-means algorithm to obtain important features. However, when the k-means algorithm performs the first clustering, it is necessary to pre-set the number of cluster centers. If the number of cluster centers is not set properly, the final clustering result will be greatly affected by the number of cluster centers, which will lead to a lower accuracy of the final clustering result. However, since the number of first feature data is large, and each first feature data has multiple features, each feature corresponds to corresponding data, it is impossible to directly determine the number of cluster centers based on the first feature data. The first feature data can be processed by the parallel coordinate method to obtain visualized second feature data. The second feature data can be displayed on a parallel coordinate graph for people or computers to more intuitively observe the distribution of the second feature data, so that subsequent people or computers can set a more reasonable number of cluster centers.

[0074] After the parallel coordinate method is used, the first feature data is processed according to the feature number to obtain the second feature data corresponding to the feature number. The second feature data will continue to use the label of the first feature data, and each second feature data has an ice label or a non-ice label. For example, referring to Figure 2 The first characteristic data is [10m / s, 15℃, 10℃]. The first characteristic data is processed to obtain three second characteristic data, namely the second characteristic data of wind speed of 10m / s, the second characteristic data of pitch motor temperature of 15℃ and the second characteristic data of ambient temperature of 10℃. The second characteristic data are connected by lines to indicate that they belong to the same first characteristic data.

[0075] On each feature, the first cluster center is determined based on the distribution of the second feature data. The number of the first cluster centers can be the number of clusters on each feature. For example, referring to Figure 3 , there are 5 first feature data in total, and each first feature data is processed to obtain 6 second feature data. On feature B, the second feature data is clustered into two groups, as shown by circle a and circle b. Therefore, two first cluster centers are set on this feature, and one first cluster center is set in each of the two groups of second feature data, that is, one first cluster center is set in each of circle a and circle b. If there is no clustering on a certain feature, such as Figure 3 As shown in the data of feature D in , the second feature data is discretely distributed, so only one first cluster center will be set on feature D. In addition, since discrete data has no reference value for predicting any results, the only first cluster center can be set at any value of feature D.

[0076] After determining the first cluster centers of all features in the second feature data, clustering of the second feature data is started based on the first cluster centers.

[0077] Step 104: Calculate the first Euclidean distance between each second feature data and each first cluster center on the feature included in the second feature data, and determine the first cluster center having the shortest first Euclidean distance to the second feature data.

[0078] The shorter the first Euclidean distance, the higher the correlation between the second feature data and the first cluster center. Therefore, by determining the first cluster center with the shortest first Euclidean distance to the second feature data, the second feature data with higher correlation can be clustered according to the first cluster center.

[0079] In some embodiments, according to the formula Calculate the shortest first Euclidean distance from the second feature data to each first cluster center on the feature corresponding to the second feature data, where dist(x m ,c n ) is the shortest first Euclidean distance from the second feature data to each first cluster center on the feature corresponding to the second feature data, X includes all second feature data, x m is the mth second feature data, C is the set of the first cluster centers, c n is x m The first cluster center on the corresponding feature.

[0080] Compare the calculated first Euclidean distances between each second feature data and each first cluster center on the features included therein, and select the first cluster center with the shortest first Euclidean distance as the cluster center of the second feature data.

[0081] Step 105: Divide the second feature data into a first cluster center having the shortest first Euclidean distance to the second feature data, to obtain a first clustering result.

[0082] like Figure 4 As shown, there are two first cluster centers c and d on feature G. Comparing the first Euclidean distance from the second feature data e to the first cluster center c and the first Euclidean distance from the second feature data e to the first cluster center d, the first Euclidean distance from the second feature data e to the first cluster center c is shorter, and accordingly the second feature data e is divided into the first cluster center c.

[0083] Step 106: Calculate the mean of the second characteristic data of each first cluster center according to the first clustering result as the second cluster center.

[0084] In some embodiments, according to the formula Calculate the second cluster center, where c i2is the i-th second cluster center, N is the number of second feature data divided into the i-th first cluster center, x m is the mth second feature data, c i1 is the first cluster center of the i-th cluster.

[0085] Step 107: Calculate the second Euclidean distance between each second feature data and each second cluster center on the features included in the second feature data, and determine the second cluster center having the shortest second Euclidean distance to the second feature data.

[0086] The specific calculation formula of step 107 is basically the same as the calculation formula of the aforementioned step 104. It is only necessary to replace the first cluster center with the second cluster center and the first Euclidean distance with the second Euclidean distance. Please refer to the previous description and will not repeat it here.

[0087] The calculated second Euclidean distances between each second feature data and each second cluster center on the features included therein are compared, and the second cluster center with the shortest second Euclidean distance is selected as the cluster center of the second feature data.

[0088] Step 108: Divide the second characteristic data into a second cluster center having the shortest second Euclidean distance to the second characteristic data, to obtain a second clustering result.

[0089] Step 109: Extract third feature data based on the second clustering results. This step extracts third feature data that is highly correlated with whether wind turbine blades are iced or not from the second clustering results. Because the second clustering results have been clustered multiple times, the second feature data of each second cluster center are closely related. Therefore, third feature data can be extracted from the second clustering results as training samples.

[0090] From the second clustering results, determine whether most of the second feature data of each second cluster center have the same label. For example, if the second feature data with the icing label in a second cluster center accounts for a large proportion of the total number of second feature data in the second cluster center, it means that the second feature data with the icing label in the second cluster center is related to icing on the wind turbine blades. The second feature data with the icing label in the second cluster center can be extracted as the third feature data of the training sample. A specific method for extracting the third feature data is provided below. In some embodiments, step 109 includes:

[0091] Step a01: for each second cluster center, determine a first number of second feature data with an icing label and a second number of second feature data with a non-icing label.

[0092] Step a02: Calculate a first ratio of the first number to the total number of second feature data of the second cluster center, and a second ratio of the second number to the total number of second feature data of the second cluster center.

[0093] Step a03: Determine whether the first ratio and the second ratio have a ratio greater than or equal to a second threshold. If so, proceed to step a04; otherwise, terminate the process, indicating that the second cluster center does not contain second feature data related to whether the wind turbine blade is iced or not.

[0094] Step a04: extracting the second feature data corresponding to the ratio greater than or equal to the second threshold in the second cluster center as the third feature data.

[0095] The total number of second feature data in each second cluster center is the first number plus the second number. The first ratio is used to represent the proportion of second feature data with an icing label in the second cluster center, and the second ratio is used to represent the proportion of second feature data with a non-icing label in the second cluster center.

[0096] In order to determine the degree of correlation between the second characteristic data of each second cluster center and whether the blade is frozen or not, it is necessary to set a judgment standard (i.e., a second threshold value). If the first ratio or the second ratio is greater than the second threshold value, the correlation is high; if it is less than the second threshold value, the correlation is low. If the first ratio or the second ratio is greater than the second threshold value, it is determined that the second characteristic data with an icing label or the second characteristic data with an unicing label accounts for a larger proportion of the total number of the second characteristic data of the second cluster center to which it corresponds, thereby determining that the correlation between icing or not and the second characteristic data of the second cluster center is large, and determining that the second characteristic data with an icing label of the second cluster center is an important feature related to icing. In some embodiments, the second threshold value is 90%. For example, referring to Figure 5 When the first ratio value of the second cluster center h on feature B is greater than or equal to 90%, it indicates that the second feature data with the icing label accounts for a large proportion of the second cluster center, and the second feature data with the icing label of the second cluster center has a high correlation with icing of the wind turbine blades. Accordingly, the second feature data with the icing label in the second cluster center is extracted as the important data corresponding to the important feature (feature B) (i.e., the third feature data). Similarly, when the second ratio value of the second cluster center g on feature B is greater than or equal to 90%, it indicates that the second feature data with the non-icing label accounts for a large proportion of the second cluster center, and the second feature data with the non-icing label of the second cluster center has a high correlation with icing of the wind turbine blades. Accordingly, the second feature data with the non-icing label in the second cluster center is extracted as the important data corresponding to the important feature (feature B) (i.e., the third feature data).

[0097] Step 110: Modeling the third characteristic data using a recurrent neural network to obtain a wind turbine blade icing prediction model.

[0098] Since the characteristic data of icing on wind turbine blades are generally time series data, that is, each first characteristic data includes at least time, the extracted third characteristic data also includes at least time, and there is a certain correlation between the third characteristic data at the previous moment and the third characteristic data at the subsequent moment in the time series data, the third characteristic data at the previous moment and the third characteristic data at the subsequent moment of the same sample wind turbine are used as a group of training samples, and all training samples are modeled through a recurrent neural network.

[0099] In some embodiments, part of the third characteristic data is saved for subsequent model evaluation of the wind turbine blade icing model, and the part of the third characteristic data does not participate in the model training process.

[0100] The model evaluation indicators used are mean square error (MSE), root mean square error (RMSE) and mean absolute percentage error (MAPE). MSE is used as the loss function of the recurrent neural network to estimate the error of the current state of the wind turbine blade icing prediction model.

[0101] The mean square error MSE is calculated according to the formula Calculated, where z i is the true value of the i-th sample point, is the predicted value for the i-th sample point, and N is the number of sample points. When the predicted value and the true value completely match, the generated wind turbine blade icing prediction model is considered perfect. A larger mean square error (MSE) indicates a greater error in the generated wind turbine blade icing prediction model. Taking the square root of the MSE yields the root mean square error (RMSE). The RMSE more intuitively reflects the magnitude of the deviation between the predicted value and the true value.

[0102] The mean absolute percentage error MAPE is calculated according to the formula It is calculated that the mean absolute percentage error (MAPE) is equivalent to normalizing the error of each sample point, reducing the influence of outliers. The smaller the value of the mean absolute percentage error (MAPE), the more accurate the generated wind turbine blade icing prediction model is.

[0103] Step 111: predicting the icing condition of the wind turbine blade to be predicted according to the wind turbine blade icing prediction model.

[0104] The complex and diverse first feature data is converted into visual second feature data through the parallel coordinate method, thereby determining the first cluster center for subsequent clustering based on the first cluster center to obtain the first cluster result. The second cluster center is determined based on the first cluster result, and the second cluster result is obtained by clustering the second feature data multiple times, so that the correlation between the second feature data of the second cluster centers is higher. By judging the degree of correlation between the second feature data of each second cluster center and icing or non-icing from the second cluster result, the second feature data with icing labels or non-icing labels of the second cluster center with a high degree of correlation with icing or non-icing is used as the third feature data for subsequent modeling. Modeling is performed based on the third feature data with a high correlation with icing or non-icing, and the resulting wind turbine blade icing prediction model is more accurate, so that the icing condition of the wind turbine blade to be predicted can be predicted more accurately.

[0105] After obtaining the prediction results, the user can process the wind turbine accordingly according to the prediction results to prevent the wind turbine blades from freezing or to process the wind turbine blades that have already frozen. Specifically, the user inputs the same feature data of the wind turbine blades to be predicted as the features used to train the wind turbine blade icing prediction model. For example, the wind turbine blade icing prediction model is trained using the features of wind speed, ambient temperature, and pitch motor temperature. Then, the user will input the current state values ​​of the wind speed, ambient temperature, and pitch motor temperature of the wind turbine blades to be predicted. The wind turbine blade icing prediction model then makes a prediction based on the input values ​​to obtain a prediction result of whether the wind turbine blades to be predicted have frozen, are about to freeze, or will not freeze. Based on the prediction results, the user can make adjustments to the state of the wind turbine to prevent the freezing of wind turbine blades that are about to freeze or to solve the problem of wind turbine blades that have already frozen.

[0106] Since the first feature data obtained generally has problems such as incomplete data, lack of unified standards between different feature data, outliers and redundant data, even if the first feature data is processed by the parallel coordinate method to obtain the second feature data, it is impossible to directly mine the second feature data. Even if it can be mined, the mining results are unsatisfactory. Therefore, in order to improve the quality of data mining, it is necessary to preprocess the second feature data to improve the accuracy of subsequent extraction of the third feature data. In some embodiments, after the first feature data is processed according to the parallel coordinate method to obtain the second feature data, the method further includes the following steps:

[0107] Step b01: Fill missing values ​​in the second feature data, remove outliers according to the boxplot method, and perform standardization.

[0108] The missing value filling, outlier removal using the boxplot method, and standardization can be performed on the second feature data in any order. For example, the missing value filling can be performed on the second feature data first, then outlier removal can be performed, and finally, the second feature data can be standardized. Alternatively, the outlier removal can be performed on the second feature data first, then missing value filling can be performed, and finally, the standardization can be performed. This is not limited here and can be set according to actual needs.

[0109] In some embodiments, the second feature data is first processed for missing values, and the missing values ​​of the second feature data are filled according to the mean method. N is a set of N values ​​near the missing value, then the mean of the set is Where y is the missing value that needs to be filled (that is, the mean of the set of N values ​​near the missing value).

[0110] In some embodiments, since most of the second feature data are non-normally distributed data, it is necessary to find outliers from the second feature data according to the box plot method, and then remove outliers from the second feature data. First, draw a number axis, the unit size of the number axis is consistent with the unit of the second feature data, and the length of the number axis is slightly longer than the full range of the second feature data. Secondly, draw a rectangular box, the positions of the two end sides correspond to the upper quartile and lower quartile of the second feature data respectively, and draw a median line at the position of the median inside the rectangular box. Then draw a line segment at a position 1.5 times the interquartile range above the upper quartile, and draw a line segment at a position 1.5 times the interquartile range below the lower quartile. These two line segments are called inner limits. Then draw a line segment at a position 3 times the interquartile range above the upper quartile, and draw a line segment at a position 3 times the interquartile range below the lower quartile. These two line segments are called outer limits. The interquartile range is the difference between the upper quartile and the lower quartile. The second characteristic data within the inner limit is a normal value, the second characteristic data between the inner limit and the outer limit is a mild abnormal value, and the second characteristic data outside the outer limit is an extreme abnormal value. Finally, all abnormal values ​​of the second characteristic data are deleted.

[0111] In some embodiments, since the second feature data of different features have different calculation units, it is necessary to remove the dimension effect of the second feature data. The second feature data is normalized, where x * is the standardized second feature data, x is the second feature data, μ is the mean of the second feature data, and σ is the variance of the second feature data. By standardizing the second feature data, the dimension effect is removed, the numerical differences between the second feature data of different characteristics are reduced, and subsequent processing of the second feature data is more accurate.

[0112] In some embodiments, reference Figure 6 After step 108, the method further includes:

[0113] Step c01: According to the formula Calculate the error sum of squares of the second clustering result, where E is the error sum of squares, k is the number of the second cluster centers, and c i2 is the second cluster center of the i-th cluster, dist(x m , c i2 ) is x m To the second cluster center c i2 The Euclidean distance, x m is the mth second feature data.

[0114] Step c02: Determine whether the sum of squared errors of the second clustering result is less than a first threshold. If so, proceed to step c03; otherwise, proceed to step c04. Step c03: Output the second clustering result.

[0115] Step c04: Use the second clustering result as the first clustering result and go to step 106.

[0116] Through step c01 to step c02, the second feature data is iteratively clustered, and the second cluster centers are continuously optimized until the sum of square errors of the second clustering results is less than the first threshold value, so that the clustering effect of the second clustering results finally output is better, so that important features related to whether the wind turbine blades are icing or not can be accurately extracted subsequently.

[0117] Figure 7 The following is a schematic diagram of a wind turbine blade icing prediction device 200 provided in an embodiment of the present application, which is applied to a modeling platform. As shown in the figure, the device includes:

[0118] The first acquisition module 201 is used to acquire first characteristic data of a sample wind turbine, each first characteristic data including at least two characteristics and data corresponding to the characteristics;

[0119] The first processing module 202 is configured to process the first feature data according to the parallel coordinate method to obtain second feature data, where each second feature data only includes one feature of the first feature data and data corresponding to the feature;

[0120] A first determining module 203 is configured to determine a first cluster center of each feature based on the second feature data. Each feature includes at least one first cluster center, and the first cluster center is used to cluster the second feature data of the feature where the first cluster center is located.

[0121] The first calculation module 204 is configured to calculate a first Euclidean distance between each second feature data and each first cluster center on the feature included in the second feature data, and determine a first cluster center having the shortest first Euclidean distance to the second feature data;

[0122] The second processing module 205 is configured to divide the second feature data into a first cluster center having the shortest first Euclidean distance to the second feature data, to obtain a first clustering result;

[0123] The second determining module 206 is configured to calculate, for each first cluster center according to the first clustering result, a mean value of the second characteristic data of each first cluster center as a second cluster center;

[0124] The second calculation module 207 is used to calculate the second Euclidean distance between each second feature data and each second cluster center on the feature included in the second feature data, and determine the second cluster center with the shortest second Euclidean distance to the second feature data;

[0125] Data partitioning module 208: configured to partition the second feature data into a second cluster center having the shortest second Euclidean distance to the second feature data, to obtain a second clustering result;

[0126] Data extraction module 209: used to extract third feature data according to the second clustering result;

[0127] Data modeling module 210: configured to use a recurrent neural network to model the third characteristic data to obtain a wind turbine blade icing prediction model;

[0128] Prediction module 211: used to predict the icing condition of the wind turbine blades to be predicted according to the wind turbine blade icing prediction model.

[0129] Figure 8 A structural diagram of the modeling platform provided in an embodiment of the present application is shown. The specific embodiment of the present application does not limit the specific implementation of the modeling platform.

[0130] like Figure 8 As shown, the modeling platform may include: a processor 302 , a communications interface 304 , a memory 306 , and a communication bus 308 .

[0131] Processor 302, communication interface 304, and memory 306 communicate with each other via communication bus 308. Communication interface 304 is used to communicate with other devices, such as client devices or other server network elements. Processor 302 is used to execute program 310, which may specifically perform the steps described in the aforementioned embodiment of the method for predicting icing on wind turbine blades.

[0132] Specifically, the program 310 may include program code including computer-executable instructions.

[0133] Processor 302 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the modeling platform may be processors of the same type, such as one or more CPUs, or may be processors of different types, such as one or more CPUs and one or more ASICs.

[0134] The memory 306 is used to store the program 310. The memory 306 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0135] An embodiment of the present application provides a computer-readable storage medium storing executable instructions. When the executable instructions are executed on a modeling platform, the modeling platform executes the wind turbine blade icing prediction method in any of the above method embodiments.

Claims

1. A wind turbine blade icing prediction method, applied to a modeling platform, characterized in that: The method comprises: Acquire first characteristic data of a sample wind turbine, each of the first characteristic data including at least two characteristics and data corresponding to the characteristics; Processing the first feature data according to a parallel coordinate method to obtain second feature data, each second feature data including only one feature of the first feature data and data corresponding to the feature, the second feature data including an ice label or a non-ice label; Determining a first cluster center for each of the features based on the second feature data, wherein each feature includes at least one first cluster center, and the first cluster center is used to cluster the second feature data on the feature where the first cluster center is located; Calculating a first Euclidean distance between each second feature data and each first cluster center on a feature included in the second feature data, and determining the first cluster center having the shortest first Euclidean distance to the second feature data; Dividing the second feature data into the first cluster center having the shortest first Euclidean distance to the second feature data, to obtain a first clustering result; Calculating, for each of the first cluster centers, a mean value of the second feature data of each of the first cluster centers according to the first clustering result as a second cluster center; Calculating a second Euclidean distance from each second feature data to each second cluster center on a feature included in the second feature data, and determining the second cluster center having the shortest second Euclidean distance to the second feature data; Dividing the second characteristic data into the second cluster center having the shortest second Euclidean distance to the second characteristic data, to obtain a second clustering result; For each of the second cluster centers, determining a first number of the second feature data having an icing label and a second number of the second feature data having a non-icing label; Calculating a first ratio of the first number to the total number of the second characteristic data of the second cluster center, and a second ratio of the second number to the total number of the second characteristic data of the second cluster center; Determining whether a ratio between the first ratio and the second ratio is greater than or equal to a second threshold; If so, extracting the second feature data corresponding to the ratio in the second cluster center as the third feature data; Modeling the third characteristic data using a recurrent neural network to obtain a wind turbine blade icing prediction model; The wind turbine blade icing condition to be predicted is predicted according to the wind turbine blade icing prediction model.

2. The wind turbine blade icing prediction method according to claim 1, wherein: After the first feature data is processed according to the parallel coordinate method to obtain the second feature data, the method further includes: The second feature data is subjected to missing value filling, outlier removal and standardization according to the box plot method.

3. The wind turbine blade icing prediction method according to claim 1, wherein: The first characteristic data includes at least one characteristic of wind speed, wind direction, average wind direction, ambient temperature, time, generator speed, pitch motor temperature, main shaft speed and blade speed.

4. The wind turbine blade icing prediction method according to claim 1, wherein: The calculating a first Euclidean distance between each second feature data and each first cluster center on a feature included in the second feature data, and determining the first cluster center having the shortest first Euclidean distance to the second feature data, comprises: According to the formula The shortest first Euclidean distance from the second feature data to each first cluster center on the feature corresponding to the second feature data is calculated, where is the shortest first Euclidean distance from the second feature data to each first cluster center on the feature corresponding to the second feature data, including all the second characteristic data, For the m The second characteristic data, C is the set of the first cluster centers, for The first cluster center on the corresponding feature.

5. The wind turbine blade icing prediction method according to claim 1, wherein: The calculating, for each of the first cluster centers according to the first clustering result, a mean value of the second feature data of each of the first cluster centers as a second cluster center, includes: According to the formula The second cluster center is calculated, where For the i The second cluster center, N To be divided into the i The number of the second feature data of the first cluster center, For the m The second characteristic data, For the i The first cluster centers.

6. The wind turbine blade icing prediction method according to claim 1, wherein: After dividing the second feature data into the second cluster centers having the shortest second Euclidean distance to the second feature data to obtain a second clustering result, the method further includes: According to the formula Calculate the error sum of squares of the second clustering result, where E is the error sum of squares, k is the number of the second cluster centers, c i2 is the second cluster center of the i-th cluster, dist (x m ,c i2 ) for x m To the second cluster center where it is located c i2 The Euclidean distance, x m For the m the second characteristic data; Determining whether the sum of squared errors of the second clustering result is less than a first threshold; If yes, output the second clustering result; If not, the second clustering result is used as the first clustering result, and the process proceeds to the step of calculating the mean of the second feature data of each first cluster center as the second cluster center for each first cluster center according to the first clustering result.

7. A wind turbine blade icing prediction device, applied to a modeling platform, characterized in that: The device comprises: A first acquisition module is used to acquire first characteristic data of a sample wind turbine, each of which includes at least two characteristics and data corresponding to the characteristics; A first processing module is configured to process the first feature data according to a parallel coordinate method to obtain second feature data, wherein each second feature data includes only one feature of the first feature data and data corresponding to the feature, and the second feature data includes an ice label or a non-ice label; A first determining module is configured to determine a first cluster center of each feature according to the second feature data, wherein each feature includes at least one first cluster center, and the first cluster center is used to cluster the second feature data on the feature where the first cluster center is located; A first calculation module is used to calculate a first Euclidean distance between each second feature data and each first cluster center on the feature included in the second feature data, and determine the first cluster center having the shortest first Euclidean distance to the second feature data; A second processing module is configured to divide the second feature data into the first cluster center having the shortest first Euclidean distance to the second feature data, to obtain a first clustering result; A second determining module is configured to calculate, for each of the first cluster centers, a mean value of the second feature data of each of the first cluster centers according to the first clustering result as a second cluster center; A second calculation module is configured to calculate a second Euclidean distance between each second feature data and each second cluster center on a feature included in the second feature data, and determine the second cluster center having the shortest second Euclidean distance to the second feature data; A data partitioning module: configured to partition the second characteristic data into the second cluster center having the shortest second Euclidean distance to the second characteristic data, to obtain a second clustering result; A data extraction module is configured to determine, for each second cluster center, a first number of the second feature data having an ice label and a second number of the second feature data having an uniced label; calculate a first ratio of the first number to the total number of the second feature data of the second cluster center, and a second ratio of the second number to the total number of the second feature data of the second cluster center; determine whether the first ratio and the second ratio have a ratio greater than or equal to a second threshold; and if so, extract the second feature data corresponding to the ratio in the second cluster center as the third feature data; A data modeling module is configured to use a recurrent neural network to model the third characteristic data to obtain an icing prediction model for wind turbine blades; Prediction module: used for predicting the icing condition of the wind turbine blade to be predicted according to the wind turbine blade icing prediction model.

8. A modeling platform, characterized in that: include: A processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform the operation of the wind turbine blade icing prediction method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The storage medium stores at least one executable instruction, and the executable instruction, when running, executes the operation of the wind turbine blade icing prediction method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Wind power prediction method based on wind speed information and wind direction information

    CN105654207A

  • Wind power blade icing speculation method based on data modeling

    CN109209790A