A wind turbine abnormal data combination cleaning method

CN117454099BActive Publication Date: 2026-10-09CHINA THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311166188.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-11
Publication Date
2026-10-09
Estimated Expiration
2043-09-11

AI Technical Summary

Technical Problem

[0003]现有的风电异常数据清洗方法多为利用单一方法进行,对特定类别下的异常数据清洗效果不明显,且清洗过程中超参数的选取人为主观性比较强,最优参数的获取较为繁琐,耗费时间较长

Benefits of technology

[0025] 1. The combined data cleaning method provided by this invention improves the speed of filtering out abnormal data compared to the single adaptive DBSCAN, further enhancing the efficiency of data cleaning. Specifically, it utilizes a parameter optimization strategy to optimize the key parameters of the DBSCAN algorithm. It generates a list of candidate ε and MinPts parameters based on the distribution characteristics of the dataset itself. Based on the generated parameter information, it generates the cluster count result. When the cluster count result shows a stable trend, the corresponding ε and MinPts parameters are taken as the optimal parameters. Since the adaptive algorithm first needs to calculate the distance distribution matrix of the dataset, it can easily lead to memory overflow when the number of clusters is large. Secondly, as ε increases, the clustering result may only have one cluster, losing its classification significance and further increasing the algorithm's unnecessary time, making it unnecessary to continue obtaining the parameter list. To solve these two problems, the adaptive DBSCAN algorithm in this invention incorporates a decision loop when obtaining the parameter list. The operation ends when the cluster count in the output clustering result is 1, and the optimal parameters are selected from the candidate list generated at this moment to reduce the algorithm's running time. Meanwhile, considering that the clustering time will increase with the increase of the number of data in the dataset, the original dataset is randomly sliced ​​into multiple sub-datasets for computation. Finally, the clustering results are weighted and averaged to obtain the final result, thereby reducing the overall computation time of combined data cleaning and improving efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117454099B_ABST
    Figure CN117454099B_ABST
Patent Text Reader

Abstract

A wind turbine abnormal data combination cleaning method, comprising the following steps: S101, pre-processing wind turbine data to obtain pre-processed data; S102, combining data cleaning algorithm to clean the pre-processed data; S103, the abnormal data after cleaning is filled by using random forest algorithm. The combination data cleaning method provided by the application has improved the speed of screening abnormal data compared with single adaptive DBSCAN, and further improves the efficiency of data cleaning. Specifically, the key parameters of the DBSCAN algorithm are optimized by using the parameter optimization strategy, the to-be-selected epsilon and MinPts parameter list are generated by using the distribution characteristics of the data set itself, the cluster number result is generated according to the generated parameter information, and when the cluster number result changes stably, the epsilon and MinPts parameters corresponding to this time are used as the optimal parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wind power big data processing technology, and in particular to a method for combined cleaning of abnormal wind turbine data. Background Technology

[0002] In recent years, wind power, as one of the most mature new energy development and utilization technologies, has achieved significant progress, with installed capacity increasing year by year. However, wind turbines have complex structures and closely interconnected components, generating massive amounts of operational data during long-term operation. This data characterizes various parameters and operational status of the wind turbine. In-depth research on this data helps to understand the operating status of the wind turbine, the working conditions of its components, fault identification, power prediction, and other aspects, accelerating technological progress in the wind power industry. However, due to factors such as wind turbine operating conditions, sensor equipment limitations, and weather conditions, the collected wind power data contains a large amount of anomalies. Directly using this data for research can negatively impact the results. Therefore, it is necessary to consider the distribution characteristics of wind power data to identify and clean anomalies. In cases where the proportion of outliers is large, further data imputation is required to ensure data integrity and improve data quality.

[0003] Existing methods for cleaning abnormal wind power data mostly rely on a single approach, which is not very effective for cleaning abnormal data of specific categories. Furthermore, the selection of hyperparameters during the cleaning process is highly subjective, and obtaining the optimal parameters is cumbersome and time-consuming. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method for combined cleaning of abnormal data of wind turbines, which improves the efficiency and effect of data cleaning and reduces time consumption.

[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a method for combined cleaning of abnormal data of wind turbines, comprising the following steps:

[0006] S101, preprocess the wind turbine data to obtain preprocessed data;

[0007] S102, The combined data cleaning algorithm cleans the preprocessed data;

[0008] S103, the cleaned outlier data is filled using the random forest algorithm.

[0009] Preferably, S101 includes: marking abnormal data points where the wind speed is greater than zero and the power is zero when the wind turbine stops operating due to a fault or maintenance; abnormal data points where the power is greater than zero and the wind speed is zero are generated when the sensor fails or data is transmitted; and reducing the density of data corresponding to the state of not being able to operate normally.

[0010] Preferably, step S102 includes the following steps:

[0011] S201, Use the adaptive DBSCAN algorithm to determine the optimal parameters for the preprocessed data;

[0012] S202, perform cleaning using the DBSCAN algorithm;

[0013] S203, the elbow rule determines the optimal k for the k-means algorithm, and the k-means algorithm is then cleaned.

[0014] Preferably, step S201 first calculates the Euclidean distance between each data point in the dataset and its k-th nearest neighbor, and calculates the average of all distances. The current k value is used as the neighborhood radius value, and a neighborhood parameter list is obtained by calculating all k values. Second, based on the obtained neighborhood radius list, the expected value of the number of data points in the corresponding neighborhood is calculated, and this is used as the minimum number of neighborhood samples. Finally, the cluster density parameter is calculated using the obtained two parameter list. When the density parameter remains unchanged, it is considered that the expected clustering result has been achieved, thus obtaining the optimal neighborhood radius and the minimum number of neighborhood samples, and completing the determination of the optimal parameters.

[0015] Preferably, in step S202, the DBSCAN parameters are set to clean the wind turbine preprocessing data to obtain a normal dataset P1 and an abnormal dataset V1.

[0016] Preferably, in step 203, the elbow rule is used to determine the k value of the P1 dataset obtained in step S202. After determining the k value, the P1 dataset is cleaned using the k-means algorithm. At this time, the k value in the k-means algorithm is the value determined by the elbow rule, resulting in the normal dataset P2 and the abnormal dataset V2.

[0017] Preferably, in step S103, the normal dataset P2 after combined data cleaning is used as training data and fed into the built random forest model for training. After the trained model is obtained, the abnormal datasets V1 and V2 are used as data to be filled and input into the model to obtain the filled dataset, thus completing the data cleaning.

[0018] Preferably, the specific steps of the k-means algorithm are as follows:

[0019] Determine a value for k, then select k data points from the dataset as centroids, calculate the distance from each data point to the centroid, and assign the data point to the set to which the data point is closer to the centroid.

[0020] Let x i (a i ,b i Let be any point in the dataset, and let its distance to a certain centroid k be... i (a k ,bk The distance is

[0021] After all the data in the dataset is grouped into sets, there are k sets in total. The centroids of the k sets are recalculated. If the distance between the recalculated centroid and the original centroid is less than the set threshold, the clustering is considered to have achieved the expected result and the clustering ends.

[0022] To select the optimal value of k, the elbow rule is used to determine the value of k, as described in detail below:

[0023] Set distortion degree: the sum of squared distances between the centroid of the set and the positions of its internal members. Assuming n samples are divided into K sets, C k Let u represent the k-th set (k = 1, 2, ..., K), with u as its centroid. k Then the degree of distortion of the k-th set is: Define the total distortion of all sets. The elbow rule uses a line graph to plot the k-value versus the total distortion level. When the distortion level is significantly improved, the current k-value can be selected as the number of clusters.

[0024] This invention provides a method for cleaning abnormal data from wind turbines, which has the following beneficial effects:

[0025] 1. The combined data cleaning method provided by this invention improves the speed of filtering out abnormal data compared to the single adaptive DBSCAN, further enhancing the efficiency of data cleaning. Specifically, it utilizes a parameter optimization strategy to optimize the key parameters of the DBSCAN algorithm. It generates a list of candidate ε and MinPts parameters based on the distribution characteristics of the dataset itself. Based on the generated parameter information, it generates the cluster count result. When the cluster count result shows a stable trend, the corresponding ε and MinPts parameters are taken as the optimal parameters. Since the adaptive algorithm first needs to calculate the distance distribution matrix of the dataset, it can easily lead to memory overflow when the number of clusters is large. Secondly, as ε increases, the clustering result may only have one cluster, losing its classification significance and further increasing the algorithm's unnecessary time, making it unnecessary to continue obtaining the parameter list. To solve these two problems, the adaptive DBSCAN algorithm in this invention incorporates a decision loop when obtaining the parameter list. The operation ends when the cluster count in the output clustering result is 1, and the optimal parameters are selected from the candidate list generated at this moment to reduce the algorithm's running time. Meanwhile, considering that the clustering time will increase with the increase of the number of data in the dataset, the original dataset is randomly sliced ​​into multiple sub-datasets for computation. Finally, the clustering results are weighted and averaged to obtain the final result, thereby reducing the overall computation time of combined data cleaning and improving efficiency.

[0026] 2. The combined data cleaning method provided by this invention can improve data quality while ensuring data integrity. Specifically, the wind turbine data after combined data cleaning meets the data distribution under normal wind turbine operation conditions, but the dataset is missing filtered outlier data, which cannot guarantee the integrity of the dataset. This invention uses meteorological data other than wind speed and power, combined with the random forest algorithm to fill in the outlier data, ensuring data integrity and improving data quality. It can guarantee high-quality data information in wind turbine performance evaluation, condition monitoring, power prediction and other work, thereby improving the accuracy of prediction models. Attached Figure Description

[0027] The present invention will be further described below with reference to the accompanying drawings and embodiments:

[0028] Figure 1 This is a flowchart of the method of the present invention;

[0029] Figure 2 This is a flowchart of the combined data cleaning algorithm of the present invention;

[0030] Figure 3 This is a schematic diagram of the data filling process of the present invention;

[0031] Figure 4 This is a scatter plot of the original wind speed-power of a certain model of fan according to the present invention;

[0032] Figure 5 This is a scatter plot of wind speed and power after the original pretreatment steps of a certain type of fan according to the present invention;

[0033] Figure 6 This is a wind speed-power scatter plot provided by the present invention after combined data cleaning;

[0034] Figure 7 This is a flowchart of the random forest model algorithm provided by the present invention;

[0035] Figure 8 This is the wind speed-power scatter plot provided by the present invention after data filling. Detailed Implementation

[0036] like Figure 1 As shown, a method for cleaning abnormal data from wind turbines includes the following steps:

[0037] S101, preprocess the wind turbine data to obtain preprocessed data;

[0038] In this embodiment, abnormal wind turbine data is divided into four categories: The first category is data with wind speed greater than zero and power of zero. This type of data is mainly generated when the wind turbine is shut down due to fault or maintenance. The second category is data with wind speed of zero and power greater than zero. This type of data is mainly generated due to anemometer malfunction. The third category is data with central accumulation. This type of data is mainly caused by wind turbine curtailment. As the wind speed increases, in order to maintain the stable operation of the power grid, the wind turbine maintains a state below full power and does not follow the wind speed change. The fourth category is discrete abnormal data. This type of data is irregularly distributed mainly due to weather conditions, sensor malfunction, and electromagnetic interference during data transmission. It is some distance away from the main data area and presents a scattered state.

[0039] In this embodiment, the wind turbine data is first preprocessed to initially clean up the first and second types of abnormal data, preparing for further cleaning of the wind power data.

[0040] S102, The combined data cleaning algorithm cleans the preprocessed data;

[0041] S103, the cleaned outlier data is filled using the random forest algorithm.

[0042] In this implementation, a data imputation model is established using the random forest algorithm. First, the normal data P2 after data cleaning is used as training data to train the model. Then, the abnormal data V1 and V2 are imputed using the trained model to ensure data integrity, further improve data reliability, and the whole process has a low degree of human intervention.

[0043] Preferably, S101 includes: marking abnormal data points where the wind speed is greater than zero and the power is zero when the wind turbine stops operating due to a fault or maintenance; abnormal data points where the power is greater than zero and the wind speed is zero are generated when the sensor fails or data is transmitted; and reducing the density of data corresponding to the state of not being able to operate normally.

[0044] Preferred, such as Figure 2 As shown, step S102 includes the following steps:

[0045] S201, Use the adaptive DBSCAN algorithm to determine the optimal parameters for the preprocessed data;

[0046] S202, perform cleaning using the DBSCAN algorithm;

[0047] S203, the elbow rule determines the optimal k for the k-means algorithm, and the k-means algorithm is then cleaned.

[0048] The combined data cleaning method of the adaptive DBSCAN algorithm and k-means includes: obtaining one year of historical wind power data for a wind turbine; performing per-unit processing on the obtained wind power data and constructing the wind speed-power curve of the wind turbine; the selection of hyperparameters for the adaptive DBSCAN algorithm includes: adaptively selecting the minimum value of the radius parameter and the neighborhood density; the selection of hyperparameters for the k-means algorithm includes: determining the optimal value of the number of clusters through the elbow rule, so that the power curve generated by the data after the combined data cleaning method has the highest similarity with the standardized power curve.

[0049] The adaptive DBSCAN algorithm selects the optimal parameters based on a parameter optimization strategy. It uses the data distribution characteristics of the dataset itself to generate candidate ε and MinPts parameters, generates the number of clusters based on the candidate parameters, and selects the ε and MinPts parameters corresponding to the minimum density threshold when the number of clusters changes stably.

[0050] The adaptive algorithm first needs to calculate the distance distribution matrix of the dataset. When the amount of data to be clustered is large, it can easily lead to computer memory overflow. Secondly, as ε increases, the clustering will result in only one cluster, which is meaningless and only increases the algorithm's unnecessary computation time, making it unnecessary to continue obtaining the parameter list. To solve these two problems, this embodiment of the invention improves the adaptive DBSCAN algorithm when obtaining the parameter list. By judging the current clustering situation, when the aggregated cluster value is less than 1, the algorithm exits and generates a candidate list of parameters. The optimal parameters are selected from the current list, reducing the algorithm's running time.

[0051] When the dataset is large, the DBSCAN algorithm may cause memory overflow during operation, and the computation time will increase significantly with the increase in the amount of dataset. To improve efficiency, the original dataset is randomly sliced ​​into multiple sub-datasets for clustering, and the results are weighted and averaged to obtain the final result, thereby reducing the algorithm's running time and improving efficiency.

[0052] Preferably, step S201 first calculates the Euclidean distance between each data point in the dataset and its k-th nearest neighbor, and calculates the average of all distances. The current k value is used as the neighborhood radius value, and a neighborhood parameter list is obtained by calculating all k values. Second, based on the obtained neighborhood radius list, the expected value of the number of data points in the corresponding neighborhood is calculated, and this is used as the minimum number of neighborhood samples. Finally, the cluster density parameter is calculated using the obtained two parameter list. When the density parameter remains unchanged, it is considered that the expected clustering result has been achieved, thus obtaining the optimal neighborhood radius and the minimum number of neighborhood samples, and completing the determination of the optimal parameters.

[0053] Suppose the sample dataset D = {x1, x2, ..., x} m}, where m is the number of samples, the DBSCAN algorithm is described in detail below:

[0054] ε-neighborhood: for x j ∈D, its neighborhood contains samples in the sample set D that are the same as x. j The subset of samples whose distance is no greater than ε, i.e., N = {x} i ∈D|dis(x i ,x j )≤ε},N∈(x j The number of this subset is denoted as |N∈(x) j )|. Where dis(x) i ,x j ) is x i To x j The distance.

[0055] Core object: For any sample x j ∈D, if its ε-neighborhood corresponds to |N∈(x j It contains at least MinPts samples, that is, if |N∈(x) j If x ≥ MinPts, then x j It is the core object. MinPts represents the minimum number of samples in the neighborhood of a sample with a distance of ε.

[0056] Density directly: If x i Located at x j In the ε-neighborhood of x, and j If it is a core object, then it is called x. i By x j Density reaches directly.

[0057] Density achievable: for x in dataset D i and x j If a sample sequence p1, p2, ..., p exists, t t is the number of sample sequences satisfying p1 = x i p T =x j , and p t+1 By p t Density accessibility. In other words, density accessibility is transitive. Transitive samples p1, p2, ..., p in the sequence. T-1 All are core objects; only core objects can enable the density of other samples to be directly accessed.

[0058] Density connected: for x i and x j If a core object sample x exists in D k , making x i and xj All are composed of x k If the density is achievable, then x is called x. t and x j Density connected.

[0059] Clusters and noise: Randomly select an object x from dataset D p From x p We begin by searching the dataset D for all points that satisfy the ε and MinPts conditions and are density-reachable, forming a cluster. Objects that do not belong to any cluster are labeled as noise points.

[0060] Preferably, in step S202, the DBSCAN parameters are set to clean the wind turbine preprocessing data to obtain a normal dataset P1 and an abnormal dataset V1.

[0061] Preferably, in step 203, the elbow rule is used to determine the k value of the P1 dataset obtained in step S202. After determining the k value, the P1 dataset is cleaned using the k-means algorithm. At this time, the k value in the k-means algorithm is the value determined by the elbow rule, resulting in the normal dataset P2 and the abnormal dataset V2.

[0062] Preferred, such as Figure 3 As shown, in step S103, the normal dataset P2 after combined data cleaning is used as training data and fed into the built random forest model for training. After the trained model is obtained, the abnormal datasets V1 and V2 are used as data to be filled and input into the model to obtain the filled dataset, thus completing the data cleaning.

[0063] Preferably, the specific steps of the k-means algorithm are as follows:

[0064] Determine a value for k, then select k data points from the dataset as centroids, calculate the distance from each data point to the centroid, and assign the data point to the set to which the data point is closer to the centroid.

[0065] Let x i (a i ,b i Let be any point in the dataset, and let its distance to a certain centroid k be... i (a k ,b k The distance is

[0066] After all the data in the dataset is grouped into sets, there are k sets in total. The centroids of the k sets are recalculated. If the distance between the recalculated centroid and the original centroid is less than the set threshold, the clustering is considered to have achieved the expected result and the clustering ends.

[0067] To select the optimal value of k, the elbow rule is used to determine the value of k, as described in detail below:

[0068] Set distortion degree: the sum of squared distances between the centroid of the set and the positions of its internal members. Assuming n samples are divided into K sets, C k Let u represent the k-th set (k = 1, 2, ..., K), with u as its centroid. k Then the degree of distortion of the k-th set is: Define the total distortion of all sets. The elbow rule uses a line graph to plot the k-value versus the total distortion level. When the distortion level is significantly improved, the current k-value can be selected as the number of clusters. Where J... k x represents the distortion level of the k-th cluster. i Let u be the position of point i in the cluster. k C represents the location of the cluster's center point. k Let K be the k-th cluster, where K is the number of clusters.

[0069] The present application is verified through a specific embodiment below, selecting a wind turbine in a wind farm, such as... Figure 4 The image shows the original wind speed-power scatter plot of the wind turbine. The original data was first preprocessed to reduce the density of outliers, resulting in the following image. Figure 5 The wind speed-power scatter plot shown below Figure 5 Only the preprocessed data is retained.

[0070] according to Figure 5 The data is processed using the adaptive DBSCAN algorithm to determine parameter values. When the cluster density parameter remains stable, the neighborhood radius and the minimum number of neighborhood samples are considered to be the optimal values. At this time, the two parameters are 0.321 and 10, respectively. Then, the wind turbine data is cleaned using the DBSCAN clustering algorithm based on the two parameters to obtain the normal dataset P1 and the abnormal dataset V1.

[0071] The optimal number of clusters k for the P1 normal dataset was determined to be 11 using the elbow rule. Then, the k-means algorithm was used to clean the P1 dataset, resulting in the normal dataset P2 and the abnormal dataset V2. The final cleaned result is as follows. Figure 6 As shown in the scatter plots obtained by comparing the results before and after cleaning, it can be seen that... Figure 6 Most of the abnormal data has been cleared, and the obtained data is now quite accurate.

[0072] Using the cleaned and normal data P2 from the combined data as training data, a random forest imputation model is built. The flowchart of the random forest algorithm is as follows. Figure 7As shown, firstly, N sample datasets are obtained by randomly sampling with replacement in the given P2 dataset; secondly, each sample dataset is input into each decision tree, and a portion of the feature sample dataset is randomly selected from the total features as the feature dataset, and a decision tree is generated in the feature dataset in the optimal splitting manner; finally, N decision trees are generated to form a random forest.

[0073] The random forest model's decision tree parameters are set to 101. Outlier datasets V1 and V2 are used as imputation data, and the random forest model is used to fill in the gaps, resulting in the final overall wind turbine wind speed-power scatter plot, as shown below. Figure 8 As shown in the figure, the abnormal data has been filled in, and the obtained data is relatively complete.

[0074] The above embodiments are merely preferred technical solutions of the present invention and should not be considered as limitations on the present invention. The scope of protection of the present invention should be limited to the technical solutions described in the claims, including equivalent substitutions of the technical features described in the claims. That is, equivalent substitutions and improvements within this scope are also within the scope of protection of the present invention.

Claims

1. A method for cleaning abnormal data from wind turbines, characterized in that, Includes the following steps: S101, preprocess the wind turbine data to obtain preprocessed data; S102, The combined data cleaning algorithm cleans the preprocessed data; S103, the cleaned outlier data is filled using the random forest algorithm; Step S102 includes the following steps: S201, Use the adaptive DBSCAN algorithm to determine the optimal parameters for the preprocessed data; S202, perform cleaning using the DBSCAN algorithm; S203, the elbow rule determines the optimal k for the k-means algorithm, and the k-means algorithm is cleaned. Step S201 first calculates the Euclidean distance between each data point in the dataset and its k-th nearest neighbor, and calculates the average of all distances. The current k value is used as the neighborhood radius value, and a neighborhood parameter list is obtained by calculating all k values. Second, based on the obtained neighborhood radius list, the expected value of the number of data points in the corresponding neighborhood is calculated, and this is used as the minimum number of neighborhood samples. Finally, the cluster density parameter is calculated using the obtained two parameter lists. When the density parameter remains unchanged, it is considered that the expected clustering result has been achieved, thus obtaining the optimal neighborhood radius and the minimum number of neighborhood samples, and completing the determination of the optimal parameters. By randomly slicing the original dataset into multiple sub-datasets for clustering, the results are weighted and averaged to obtain the final result. Sample dataset , Given the number of samples, the DBSCAN algorithm is described in detail below: Neighborhood: for Its neighborhood contains the sample set Zhongyu The distance is no greater than The subset of samples, that is, , The number of this subset is denoted as ;in for arrive The distance; Core object: For any sample If its Neighborhood corresponding At least include A sample, that is, if ,but It is the core object; among which The sample distance is The minimum number of neighborhood samples; Density reaches directly: If lie in of In the neighborhood, and If it is a core object, then it is called Depend on Density reaches directly; Density achievable: for datasets middle and If a sample sequence exists , The number of sample sequences satisfies , ,and Depend on Density reaches directly; Transmission samples in the sequence All are core objects, and the density of other samples is directly accessible to the core objects. Density connected: For and ,if There are core object samples in it. ,make and All by If the density can be reached, then it is called... and Density connected; Clusters and noise: from the dataset Choose any object ,from Start in the dataset Search in the middle satisfies and All points that are conditionally and density-reachable form a cluster, while objects that do not belong to any cluster are marked as noise points.

2. The method for combined cleaning of abnormal wind turbine data according to claim 1, characterized in that, S101 includes: marking abnormal data points where the wind speed is greater than zero and the power is zero when the wind turbine stops operating due to fault or maintenance; abnormal data points where the power is greater than zero and the wind speed is zero are generated when the sensor fails or data is transmitted; and reducing the density of data corresponding to the state of not being able to operate normally.

3. The method for combined cleaning of abnormal wind turbine data according to claim 1, characterized in that, In step S202, the DBSCAN parameters are set to clean the wind turbine preprocessing data, resulting in a normal dataset P1 and an abnormal dataset V1.

4. The method for combined cleaning of abnormal wind turbine data according to claim 3, characterized in that, In step 203, the elbow rule is used to determine the k value of the P1 dataset obtained in step S202. After determining the k value, the P1 dataset is cleaned using the k-means algorithm. At this time, the k value in the k-means algorithm is the value determined by the elbow rule, resulting in the normal dataset P2 and the abnormal dataset V2.

5. The method for combined cleaning of abnormal wind turbine data according to claim 4, characterized in that, In step S103, the normal dataset P2 after combined data cleaning is used as training data and fed into the built random forest model for training. After the trained model is obtained, the abnormal datasets V1 and V2 are used as data to be filled and input into the model to obtain the filled dataset, thus completing the data cleaning.

6. The method for combined cleaning of abnormal wind turbine data according to claim 1, characterized in that, The specific steps of the k-means algorithm are as follows: Determine a value for k, then select k data points from the dataset as centroids, calculate the distance from each data point to the centroid, and assign the data point to the set to which the data point is closer to the centroid. set up For any point in the dataset, how far does it travel to a certain centroid? The distance is ; After all the data in the dataset is grouped into sets, there are k sets in total. The centroids of the k sets are recalculated. If the distance between the recalculated centroid and the original centroid is less than the set threshold, the clustering is considered to have achieved the expected result and the clustering ends. To select the optimal value of k, the elbow rule is used to determine the value of k, as described in detail below: Set distortion degree: the sum of squared distances between the centroid of the set and the positions of its internal members. Assuming n samples are divided into K sets, Describes the k-th set The centroid of this set is Then the degree of distortion of the k-th set is: Define the total distortion of all sets. The elbow rule uses a line graph to plot the k-value versus the total distortion level. When the distortion level is significantly improved, the current k-value can be selected as the number of clusters.