Wind power blade early damage identification method and system

By employing a clustering method that involves two cleaning processes on the wind speed-power data of the SCADA system and combining it with a convolutional neural network model and acoustic emission signals, the problem of identifying early damage to wind turbine blades was solved, enabling efficient and real-time damage identification and maintenance optimization.

CN120819476APending Publication Date: 2025-10-21SHENYANG UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511207946.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively identify early damage to wind turbine blades. Traditional detection methods are inefficient and costly, and SCADA data cleaning quality is poor, making it impossible to accurately identify blade damage in real time. This affects the safe and reliable operation and maintenance costs of wind turbine units.

Method used

By combining real-time wind speed-power data from the SCADA system with two data cleaning processes, an improved density-based noisy spatial clustering method and an improved K-means algorithm are used to construct a convolutional neural network model, and acoustic emission signals are used to identify early blade damage.

Benefits of technology

It improves the reliability and accuracy of SCADA data, enables efficient and real-time identification of early damage to wind turbine blades, reduces maintenance costs, and provides important reference for repair solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120819476A_ABST
    Figure CN120819476A_ABST
Patent Text Reader

Abstract

The invention provides a wind turbine blade early damage identification method and system, and belongs to the technical field of wind turbine generator set state monitoring, and the method comprises the steps: carrying out the two-time data cleaning of the wind speed-power real-time data of an SCADA system; comparing the wind speed-power real-time data after the last cleaning with a standard wind speed-power curve to obtain an initial wind power blade damage state; and inputting the real-time acoustic emission signals into the wind power blade early damage prediction model to obtain the type and degree of early damage identification of the wind power blade, the wind power blade early damage prediction model, classifying the acoustic emission signals collected by a fatigue loading test through an improved K-means algorithm, and constructing a convolutional neural network. According to the method, the problem that an unsupervised learning algorithm excessively depends on experience to set a threshold value is avoided, the defects of poor data cleaning quality and low algorithm generalization are overcome, and the reliability, accuracy and usability of SCADA data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wind turbine generator set status monitoring, and particularly relates to a method and system for identifying early damage to wind turbine blades. Background Art

[0002] As wind turbines become larger, their components are subject to complex alternating loads, leading to frequent failures. Blades are key components for wind turbines to capture wind energy, accounting for 15%-20% of the total cost. Due to manufacturing defects such as delamination, wrinkling, and poor curing, as well as the combined effects of outdoor environments like wind, snow, and frost, early damage such as fiber breakage, matrix cracks, base fiber debonding, and interlaminar cracking typically begins after 2-5 years of operation. During this time, the blades may be within their warranty period, making damage easily overlooked and undetected. Main beam damage is a major cause of fracture accidents. Routine maintenance such as replacing, repairing, painting, and removing coverings for wind turbine blades, due to their large size, inconvenient transportation and disassembly, and the need for high-altitude work, leads to soaring maintenance costs, significantly reducing wind turbine power generation time and output. At the same time, if the degree and type of early blade damage cannot be identified in a timely and accurate manner, it will be difficult to formulate a blade repair plan based on the severity and type of damage, and it will be impossible to formulate a reasonable blade preventive maintenance strategy during the maintenance of the wind turbine. It is difficult to effectively maintain the important load-bearing structure of the blade. Therefore, the early damage identification of the blade main beam is of great significance to the safe and reliable operation of the wind turbine, extending the service life, improving the power generation efficiency, and reducing the operation and maintenance costs.

[0003] Most blade main beams are unidirectional laminate structures composed of a matrix of carbon fiber, glass fiber, carbon / glass hybrid fiber and epoxy resin, which bear 80% of the load of the blade structure. Early damage on the main beam surface can be restored to the original design strength after repair at a low cost. However, once the golden period for repair is missed, the probability of fracture failure will increase greatly. At present, the traditional inspection methods for wind turbine blades mainly include the percussion sound identification method, telescope observation method, and close visual inspection by skilled workers. These inspection methods are inefficient and mainly rely on manual experience, resulting in a large number of blades not being repaired in a timely manner. Over-repair and under-repair often occur. However, the current online identification methods for early blade damage still have the following difficulties:

[0004] (1) Installing vibration, strain, fiber optic and other sensors for real-time status monitoring. On the one hand, the collected signals cannot identify early damage to the blade main beam laminate structure, and can usually only detect serious blade failures (such as imbalance, ice coating, cracks, fractures, etc.). On the other hand, wind turbine blades are giant hollow rods, and online monitoring requires a large number of sensors to identify local damage. The equipment cost, installation and maintenance, and data storage are expensive and difficult to implement. In addition, the use of non-destructive testing technologies such as infrared and ultrasonic testing is extremely costly and requires downtime for testing, which cannot identify the service status of the blade in real time.

[0005] (2) Most wind farms are equipped with supervisory control and data acquisition (SCADA) systems. SCADA systems typically record data every 10 minutes to reduce data storage. Wind turbine condition monitoring based on SCADA data mining has become one of the most researched methods. However, wind turbines operate in harsh environments. Under the influence of multiple complex loads, communication failures, extreme weather, instrument damage, wind farm curtailment, power consumption restrictions, and other factors inevitably generate a large amount of abnormal data. The reliability, accuracy, and availability of SCADA-collected data are directly related to the effectiveness of early damage monitoring.

[0006] (3) Wind power data is one of the key data in the SCADA system, which is used to ensure that the wind turbine operates in the optimal performance range in real time. However, in actual wind farms, the wind power data set contains a lot of noise, and abnormal noise data needs to be cleaned in real time, online, and efficiently. At present, the commonly used unsupervised learning algorithms rely too much on experience to set thresholds, and have the disadvantages of poor cleaning quality and weak generalization. The discrete interval box plot method based on statistical principles divides the data interval according to the wind speed interval and uses segmented filtering to effectively process abnormal data points. However, discrete abnormal data points close to the normal power data band are often prone to misjudgment, and these wind power data scatter points just contain information about early damage to the blades. Summary of the Invention

[0007] In response to the deficiencies in the prior art, this application proposes a method and system for identifying early damage to wind turbine blades.

[0008] In a first aspect, the present application proposes a method for identifying early damage to wind turbine blades, comprising:

[0009] For the wind turbine blades to be identified, obtain the real-time wind speed and power data from the SCADA system;

[0010] Performing two data cleanings on the wind speed-power real-time data of the SCADA system;

[0011] Compare the real-time wind speed-power data after the last cleaning with the standard wind speed-power curve to obtain the preliminary damage status of the wind turbine blades;

[0012] If the initial wind turbine blade damage status is damaged, real-time acoustic emission signals are collected;

[0013] The real-time acoustic emission signal is input into a pre-established wind turbine blade early damage prediction model to obtain the type and degree of wind turbine blade early damage identification. Wind turbine blade early damage refers to damage identified during the operational warranty period of the wind turbine blade. The pre-established wind turbine blade early damage prediction model is established by collecting acoustic emission signals through fatigue loading tests on blade specimens made of the same material as the wind turbine blade to be identified. Peak frequency and amplitude are used as feature values. The acoustic emission signals collected from the fatigue loading test are classified using an improved K-means algorithm, and a convolutional neural network is constructed to obtain the pre-established wind turbine blade early damage prediction model.

[0014] The wind speed-power real-time data of the SCADA system is cleaned twice, including:

[0015] The first data cleaning of the wind speed-power real-time data of the SCADA system was performed, including: using the threshold setting method to remove the abnormal data accumulated at the bottom of the wind speed-power real-time data of the SCADA system; using the density-based spatial clustering method with noise to remove the abnormal data sparse in the middle of the wind speed-power real-time data of the SCADA system;

[0016] If after the first data cleaning, the first remaining data volume, the first data rejection rate, the first cleaning time, and the Spearman coefficients of the wind speed-power real-time data before and after the first data cleaning all reach the first data cleaning threshold, then a second data cleaning is performed on the wind speed-power real-time data of the SCADA system, including: using an improved density-based spatial clustering method with noise to perform a second data cleaning on the central accumulated abnormal data in the wind speed-power real-time data of the SCADA system.

[0017] The abnormal data accumulated at the bottom is defined as: the difference between the wind power value corresponding to the wind speed in the wind speed-power real-time data of the SCADA system and zero is less than or equal to the wind power threshold, the number of wind speed-power real-time data located at the bottom of the wind speed-power curve is greater than the number threshold, and the density of the wind speed-power real-time data is greater than the density threshold and is distributed in a long strip shape;

[0018] The middle sparse abnormal data is defined as: the wind speed-power real-time data is sparsely distributed in the middle area of ​​the wind speed-power curve and below the standard wind speed-power curve, and the density of the wind speed-power real-time data is less than or equal to the density threshold;

[0019] The middle accumulation abnormal data is defined as: the wind speed-power real-time data of the SCADA system is located in the middle area of ​​the wind speed-power curve, and the density of the wind speed-power real-time data is greater than a density threshold.

[0020] The method of using density-based spatial clustering method with noise to remove sparse abnormal data in the middle of the wind speed-power real-time data of the SCADA system includes:

[0021] Step S2.2.1: Initialize the wind speed-power real-time data of the SCADA system, use the wind speed-power real-time data of the SCADA system as the first data set, and mark all data in the first data set as unprocessed;

[0022] Step S2.2.2: Traverse the first data set and, for each unprocessed data point, perform the following operations: mark the current data point as processed data, check all data points within the neighborhood of the current data point, where the neighborhood is the circle centered on the current data point and the neighborhood radius is the radius of the circle. The resulting circular area is called the neighborhood. If the number of all data points within the neighborhood point is less than the minimum number of points, mark the current data point as a noise point.

[0023] Step S2.2.3: If the number of all data in the neighborhood is greater than or equal to the minimum number of points, mark the current data as a core point and create a new cluster with the new cluster as the current cluster;

[0024] Step S2.2.4: For each point in the neighborhood of the core point, perform the following operations: If the current point has not been processed, mark the current point as processed and check the points in the neighborhood of the current point; if the number of points in the neighborhood of the current point is greater than or equal to the minimum number of points, and the current point has not been assigned to any cluster, add the current point to the current cluster and repeat step S2.2.4 until there are no new points that can be added to the current cluster;

[0025] Step S2.2.5: Repeat steps S2.2.2 to S2.2.4 until all data in the first data set are marked as processed data;

[0026] Step S2.2.6: Output the clustering results, including: noise clusters composed of all noise points and data clusters composed of each core point. Remove all noise clusters and use the data clusters composed of each core point as the real-time wind speed-power data after the first data cleaning.

[0027] The improved density-based spatial clustering method with noise is used to perform a second data cleaning on the central accumulation abnormal data in the wind speed-power real-time data of the SCADA system, including:

[0028] The wind speed-power real-time data of the SCADA system after the first cleaning is used as the second data set;

[0029] For each point in the second data set, calculate the distance from any point M to all other points, output the distance graph from any point M to all other points, and find the k nearest neighbor points from any point M. Record the distance from each point in the second data set to the corresponding k-th nearest neighbor point, which is called the k-nearest neighbor distance;

[0030] Sort the k-nearest-neighbor distances of all points in the second dataset in ascending order, and number all points in the second dataset according to the size of the k-nearest-neighbor distances. Draw a relationship graph between the sorted k-nearest-neighbor distances and the point numbers, which is called a k-nearest-neighbor distance graph.

[0031] The k-nearest neighbor distance corresponding to the inflection point in the k-nearest neighbor distance graph is used as the neighborhood radius of the density-based spatial clustering method with noise. The data points before the inflection point belong to the data cluster, and the data points after the inflection point belong to the noise cluster. The data points before the inflection point are used as the real-time wind speed-power data of the SCADA system after the second data cleaning.

[0032] Acoustic emission signals are collected through fatigue loading tests on blade specimens. The material of the blade specimens is consistent with the wind turbine blade to be identified. Peak frequency and amplitude are used as feature values. The acoustic emission signals collected from the fatigue loading tests are classified using an improved K-means algorithm. A convolutional neural network is constructed to obtain a pre-established wind turbine blade early damage prediction model, including:

[0033] Acquiring acoustic emission signals through a fatigue loading test of a blade specimen, the material of which is consistent with the wind turbine blade to be identified;

[0034] Preprocessing the collected acoustic emission signals;

[0035] Principal component analysis is used to reduce the dimension of the pre-processed acoustic emission signals;

[0036] The peak frequency and amplitude are used as feature values, and the improved K-means algorithm is used to classify the acoustic emission signals after dimension reduction.

[0037] The classified acoustic emission signals are used to construct a convolutional neural network to obtain a pre-established wind turbine blade early damage prediction model.

[0038] The improved K-means algorithm includes:

[0039] Step S5.4.1: using the acoustic emission signal after dimension reduction as the third data set;

[0040] Step S5.4.2: Randomly select a data point from the third data set as the first cluster center;

[0041] Step S5.4.3: Calculate the shortest distance D(x) from each point in the third data set to the first cluster center;

[0042] Step S5.4.4: Select a new point from the third data set as the second cluster center, and the probability of selection is the same as D(x) 2 proportional to;

[0043] Step S5.4.5: Repeat steps S5.4.2 to S5.4.4 until K cluster centers are selected;

[0044] Step S5.4.6: Using the selected K cluster centers as initial values, calculate the distance from each data point in the third data set to the K cluster centers, assign each data point to the cluster with the nearest cluster center, update the cluster center, and repeat step S5.4.6 until the clustering results converge.

[0045] The activation function of the convolutional neural network is the ReLU function. When the input value of the ReLU function is less than zero, the output of the ReLU function is 0. When the input value of the ReLU function is greater than or equal to zero, the output value of the ReLU function is equal to the input.

[0046] In a second aspect, the present application proposes a wind turbine blade early damage identification system, comprising:

[0047] The data acquisition module is used to obtain the real-time wind speed and power data of the SCADA system for the wind turbine blade to be identified;

[0048] A data cleaning module is used to perform two data cleanings on the wind speed-power real-time data of the SCADA system;

[0049] The initial damage identification module is used to compare the real-time wind speed-power data after the last cleaning with the standard wind speed-power curve to obtain the initial damage status of the wind turbine blade;

[0050] An acoustic emission signal acquisition module is used to collect real-time acoustic emission signals if the preliminary wind turbine blade damage status is damaged;

[0051] The damage type judgment module is used to input the real-time acoustic emission signal into the pre-established wind turbine blade early damage prediction model to obtain the type and degree of wind turbine blade early damage identification. Wind turbine blade early damage refers to damage identified during the wind turbine blade's operational warranty period. The pre-established wind turbine blade early damage prediction model is established through the following process: collecting acoustic emission signals through fatigue loading tests on blade specimens. The material of the blade specimens is consistent with the wind turbine blade to be identified. Peak frequency and amplitude are used as feature values. The acoustic emission signals collected in the fatigue loading test are classified using an improved K-means algorithm, and a convolutional neural network is constructed to obtain the pre-established wind turbine blade early damage prediction model.

[0052] In a third aspect, the present application proposes a computer program product, comprising a computer program or instructions, which, when executed by a processor, implements the aforementioned method for identifying early damage to wind turbine blades.

[0053] Beneficial effects:

[0054] This application proposes a method and system for identifying early damage to wind turbine blades. The method of this application combines SCADA historical data, real-time data and acoustic emission monitoring signals to solve the problems of different temporal and spatial resolutions of multi-sensor signals, high algorithm complexity and unreliable real-time monitoring. It can not only efficiently use the historical data of the SCADA system to monitor the service status of blades, but also adopts an offline method to establish a blade early damage prediction model. It can identify the type and degree of early blade damage in real time based on the characteristics of the acoustic emission monitoring signal without shutting down the machine.

[0055] The method of this application uses blade plywood material specimens for fatigue testing, collects acoustic emission monitoring signals and uses them as training sets and test sets after preprocessing, and constructs a convolutional neural network model for wind turbine blade damage identification. Thanks to the excellent recognition and classification capabilities of convolutional neural networks, it integrates online monitoring and offline damage assessment methods without the need for relevant physical knowledge modeling, solves the problem of online identification of early damage to wind turbine blades, and provides an important reference for formulating blade repair plans and strategies.

[0056] The method of the present application targets the causes and distribution characteristics of abnormal wind power data, adopts a two-step cleaning method for bottom accumulation, middle sparseness, and middle accumulation abnormal data, and adaptively calculates the neighborhood radius value through the slope mutation method, avoiding the problem of unsupervised learning algorithms relying too much on experience to set thresholds, solving the drawbacks of poor data cleaning quality and weak algorithm generalization, and greatly improving the reliability, accuracy and availability of SCADA collected data. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 A flow chart of a method for identifying early damage to a wind turbine blade according to an embodiment of the present application;

[0058] Figure 2 A schematic diagram of a normalized standard wind power curve according to an embodiment of the present application;

[0059] Figure 3 Standardized preprocessing of collected wind power data in the embodiment of the present application;

[0060] Figure 4 The clustering results of wind speed-power real-time data collected from SCADA in the embodiment of the present application;

[0061] Figure 5 Schematic diagram of k-nearest neighbor distances in an embodiment of the present application;

[0062] Figure 6 Schematic diagram of wind speed-power real-time data before cleaning according to an embodiment of the present application;

[0063] Figure 7 Real-time data of wind speed and power after cleaning in the embodiment of the present application;

[0064] Figure 8 Schematic diagram of preparing a composite laminate specimen of a blade main beam according to an embodiment of the present application;

[0065] Figure 9 Schematic diagram of the fatigue loading test equipment according to an embodiment of the present application;

[0066] Figure 10 Schematic diagram of cluster centers in an embodiment of the present application;

[0067] Figure 11 Schematic diagram of clustering results in the embodiment of the present application;

[0068] Figure 12 Flowchart of the convolutional neural network training for blade main beam loss according to an embodiment of the present application;

[0069] Figure 13 A schematic diagram of the training accuracy of an embodiment of the present application;

[0070] Figure 14 Schematic diagram of training loss value in an embodiment of the present application;

[0071] Figure 15 Schematic diagram of the confusion matrix of the embodiment of the present application;

[0072] Figure 16 Schematic diagram of the learning curve of the model in the embodiment of the present application. DETAILED DESCRIPTION

[0073] The specific implementation of the present application is further described in detail below with reference to the accompanying drawings and examples.

[0074] Example 1:

[0075] This embodiment takes a 2.5MW wind turbine generator set in service at a certain wind farm as an example. The model of the wind turbine set is UP86-1500, and the blade model is GP42-1.5MW-IECⅢ / Ⅳ. The main beam of the blade is used as the data source. In specific implementation, other parts of the blade can be used as the data source.

[0076] This embodiment proposes a method for identifying early damage to wind turbine blades. Figure 1 Shown, including:

[0077] Step S1: For the wind turbine blade to be identified, obtain the real-time wind speed-power data of the SCADA system;

[0078] In this embodiment, a standard wind speed-power curve of a wind turbine is drawn based on blade design parameters, and real-time wind speed-power data from the SCADA system of in-service wind turbines is collected at the wind farm. The actual wind power curve is subjected to standardization preprocessing, including:

[0079] The standard wind speed-power curve of blade design is one of the important performance indicators. The theoretical formula of power P captured by wind turbine blades is shown in formula (1):

[0080] (1);

[0081] Where: ρ is the air density, kg / m 3 ; A is the blade swept area, m 2 ; C p is the wind energy utilization coefficient, v is the wind speed, m / s. According to the design swept area and wind energy utilization coefficient of the blade model, the normalized standard wind speed-power curve of the blade is drawn, such as Figure 2 The standard wind speed-power curve can be divided into three parts: the first part is before the wind speed is cut in, when the power is zero. At this time, the wind speed is too low to cause the blades to rotate. The second part is when the wind speed does not reach the rated wind speed, and the power increases with the cube of the wind speed, satisfying the above equation. At this time, the blades can achieve energy conversion and generate electricity, but the operating power does not reach the rated power. The third part is when the wind speed reaches the rated wind speed, and the unit operates at the fixed rated power.

[0082] The standard wind speed-power curve serves as an important reference for this embodiment. If the blade is undamaged, the collected wind power data will closely follow the standard wind speed-power curve after data cleaning. Conversely, if the wind power data significantly deviates from the standard power curve, the service condition of the wind turbine blade can be preliminarily determined based on the data deviation distribution.

[0083] Since the characteristic scales of wind speed and power data are inconsistent, in order to reduce the influence of the numerical size of the monitoring data on the data clustering effect, so that the distance metric is dominated by the features with larger scales, resulting in poor cleaning effect, the data should be standardized and pre-processed first. To this end, the real-time 10-minute wind speed-power data of the SCADA system is collected on-site at the wind farm. The data is scaled to a given interval through data standardization so that all features have the same weight in the distance calculation, reducing the influence of the numerical size on the clustering result. It can be understood that: SCADA historical data, namely wind power data, is used to monitor the service status of blades. This historical data is the wind power standardization curve obtained by collecting and processing the 10-minute real-time data. Therefore, after the real-time data is collected and processed, it becomes historical data. This embodiment uses SCADA historical data to identify early damage. This embodiment uses the min-max standardization method to scale the wind power data to the [0-1] interval. The data in the data set is represented by x. The standardization pre-processing is shown in formula (2):

[0084] (2);

[0085] Where: x max is the maximum data value in the wind power dataset; x min is the minimum data value in the wind power data set; x scaled is the normalized data value.

[0086] The real-time 10-minute wind speed-power data collected from the SCADA system is pre-processed and standardized. Figure 3 shown.

[0087] Step S2: performing two data cleanings on the wind speed-power real-time data of the SCADA system;

[0088] In this embodiment, according to the noise data distribution characteristics of the wind speed-power curve, a reasonable method is selected to perform preliminary data cleaning. In this embodiment, the wind speed-power real-time data of the SCADA system is divided into three types of abnormal data: bottom accumulation, middle sparseness, and middle accumulation. The three types of abnormal data are obtained through historical data analysis of the SCADA system; the wind power data collected in the actual wind farm is a large amount of noise data generated by the SCADA system due to harsh environment, communication failure, extreme weather, instrument damage, power consumption and other reasons. According to the causes and distribution characteristics of SCADA abnormal data, the wind power data of wind turbine blades is divided into three types of abnormal data: bottom accumulation, middle sparseness, and middle accumulation. According to the noise data distribution characteristics of the wind power curve, a reasonable method is selected to perform data cleaning.

[0089] The wind speed-power real-time data of the SCADA system is cleaned twice, including:

[0090] Step S2.1: Performing the first data cleaning on the wind speed-power real-time data of the SCADA system, including: using a threshold setting method to remove abnormal data accumulated at the bottom of the wind speed-power real-time data of the SCADA system; using a density-based spatial clustering method with noise to remove abnormal data sparse in the middle of the wind speed-power real-time data of the SCADA system;

[0091] In this embodiment, the bottom-accumulated real-time data, particularly the data points, are clustered at the bottom, with noise points scattered outside this bottom. Therefore, the threshold setting is very obvious and can be observed by the human eye. However, this varies for different wind turbine models and wind farm resources, and is generally a special case that is prone to occur during installation and commissioning. Therefore, the threshold setting is a manually entered threshold. Abnormal data caused by wind turbine failure, communication failure, unplanned power outages, and other reasons is characterized by a wind power value corresponding to the wind speed close to zero, a large amount of data accumulated at the bottom of the wind speed-power curve, high density, and a long strip-like distribution. This bottom-accumulated data can be directly cleaned by setting a threshold. Therefore, the bottom-accumulated abnormal data is defined as: the difference between the wind power value corresponding to the wind speed in the SCADA system's real-time wind speed-power data and zero is less than or equal to the wind power threshold, the number of wind speed-power real-time data located at the bottom of the wind speed-power curve is greater than the quantity threshold, and the density of the wind speed-power real-time data is greater than the density threshold and is distributed in a long strip-like distribution. Bottom-accumulated data: a large amount of data accumulated at the bottom of the wind power data graph, high density, and a long strip-like distribution. This type of noise data is generated by wind turbine failure, communication equipment failure, unplanned power outages for maintenance, etc. The characteristic of this noise data is that the wind speed exists, but the power is close to zero.

[0092] Central sparse anomaly data, generated by limiting wind turbine rated power during wind farm scheduling, is characterized by sparse and scattered distribution in the middle region of the graph and below the standard wind power curve, with low data density. This data can be cleaned using the density-based spatial clustering with noise (DBSCAN) algorithm. Central sparse anomaly data is defined as: real-time wind speed-power data sparsely distributed in the middle region of the wind speed-power curve and below the standard wind speed-power curve, with a density less than or equal to the density threshold. Central sparse data is relatively sparse and randomly distributed in the middle region of the graph, lacking a clear distribution pattern but exhibiting low data density. This type of data is typically generated when a wind farm needs to control some or all turbines to a specified power level according to the tracking scheduling plan, but encounters anomalies due to communication failures, extreme weather, or other factors. This data is characterized by random and scattered distribution and low data density.

[0093] like Figure 3As shown in the figure, the noise data distribution in the wind speed-power curve is characterized by horizontal concentration. The proportion of sparse data in the middle is the largest. Its data density is significantly different from other available data. In addition, most noise data is discretely distributed around the available data. It is suitable for an unsupervised learning clustering algorithm based on the density principle. The detailed processing steps are as follows:

[0094] Step S2.2.1: Initialize the wind speed-power real-time data of the SCADA system, use the wind speed-power real-time data of the SCADA system as the first data set, and mark all data in the first data set as unprocessed;

[0095] Step S2.2.2: Traverse the first data set and for each unprocessed data point, perform the following operations: mark the current data point as processed data, check all data points in the neighborhood of the current data point (the neighborhood is determined according to the given radius Eps), and if the number of all data points in the neighborhood point is less than the minimum number of points MinPts, mark the current data point as a noise point;

[0096] Step S2.2.3: Determine the core point. If the number of all data in the neighborhood is greater than or equal to the minimum number of points MinPts, then mark the current data as a core point and create a new cluster with the new cluster as the current cluster.

[0097] Step S2.2.4: Expand the cluster. For each point in the neighborhood of the core point, perform the following operations: If the current point has not been processed, mark the current point as processed and check the points in the current point's neighborhood. If the number of points in the current point's neighborhood is greater than or equal to the minimum number of points, MinPts, and the current point has not been assigned to any cluster, add the current point to the current cluster. Repeat step S2.2.4 until there are no new points that can be added to the current cluster.

[0098] Step S2.2.5: Repeat steps S2.2.2 to S2.2.4 until all data in the first dataset are marked as processed data;

[0099] Step S2.2.6 Output the clustering results, including the noise clusters composed of all noise points and the data clusters composed of each core point.

[0100] The SCADA data set collected at the wind farm is clustered into several data clusters using a density-based spatial clustering method with noise, such as Figure 4 shown.

[0101] Step S2.2: If after the first data cleaning, the first remaining data volume, the first data rejection rate, the first cleaning time, and the Spearman coefficient of the wind speed-power real-time data before and after the first data cleaning all reach the first data cleaning threshold, then the wind speed-power real-time data of the SCADA system is subjected to a second data cleaning, including: using an improved density-based noisy spatial clustering method to perform a second data cleaning on the central accumulated abnormal data in the wind speed-power real-time data of the SCADA system.

[0102] In this embodiment, the first data elimination rate is greater than or equal to 25%, and the first cleaning time is less than or equal to 40 seconds, which represent the cleaning efficiency; the data elimination rate = (the amount of original data - the amount of remaining data) / the amount of original data, that is, the effect expressed by the first remaining data amount and the data elimination rate are consistent, representing the size of the data; the Spearman coefficient (Spearman rank correlation coefficient) represents the cleaning quality and is generally greater than 0.9.

[0103] According to step S2, perform a first cleansing of the wind farm's wind power data. Evaluate the efficiency and quality indicators before and after cleaning, including the amount of remaining data, data rejection rate, cleaning time, and Spearman coefficient. If the cleaning indicators are poor, re-set the parameters for a second cleansing. Compare the wind power data from the second cleansing process to the standard wind speed-power curve to determine the damage status of the wind turbine blades.

[0104] In this embodiment, according to step S2, the comparison Figure 3 and Figure 4 ,The amount of remaining data, data elimination rate, and ,cleaning time were used as indicators for evaluating the ,cleaning efficiency, which were 13173, 23.07%, and 59s respectively. The cleaning time was long, and ,secondary cleaning should be performed.

[0105] The cleaning quality is assessed by comparing the deviations between the wind power curves before and after cleaning and the standard wind turbine power curve. The correlation coefficient can quantitatively assess the degree of correlation between wind speed and power, with a high correlation coefficient indicating good cleaning quality. The Spearman correlation coefficient is highly versatile and applicable to both linear and nonlinear relationships. Therefore, this embodiment uses the Spearman correlation coefficient as an evaluation indicator for cleaning quality. The formula for solving the Spearman correlation coefficient is shown in Equation (3):

[0106] (3);

[0107] Among them, x i and y i are the sample values ​​of the two variables; is the mean of the variable x, is the mean of variable y; r ranges from -1 to 1. r > 0: Positive correlation, meaning that when one variable increases, the other tends to increase as well. r < 0: Negative correlation, meaning that when one variable increases, the other tends to decrease. r = 0: No correlation, meaning there is no linear relationship between the two variables.

[0108] The Spearman coefficient of wind speed-power data before and after cleaning was calculated respectively. The closer the correlation coefficient is to 1, the better the positive correlation between the data. Figure 3 The Spearman coefficient is 0.705, Figure 4 The correlation coefficient is 0.740, indicating that the cleaning quality is not ideal, and the cleaning quality of the abnormal data accumulated in the middle is poor.

[0109] This type of data is caused by random factors such as fault propagation and extreme weather. This type of abnormal data accumulates in the middle of the power curve. Unlike the sparse data in the middle, this data has a higher density, which easily leads to the mixing of abnormal and normal data, making misjudgments very likely. This results in poor cleaning efficiency and quality, and requires a secondary cleaning using an improved density-based spatial clustering method with noise. This type of data differs from the sparse data in its density. This type of data is caused by relatively random factors such as signal propagation noise and extreme weather. These noise points are densely packed in the coordinate system and accumulate in the middle of the wind power data graph.

[0110] The middle accumulation abnormal data is defined as: the wind speed-power real-time data of the SCADA system is located in the middle area of ​​the wind speed-power curve, and the density of the wind speed-power real-time data is greater than a density threshold.

[0111] Data points are classified into core points, boundary points, and noise points based on the number of other points in their neighborhood. Core points are points with a sufficient number of points in the neighborhood, boundary points are points that are not core points but are located in the neighborhood of a core point, and noise points are points that do not belong to any cluster. MinPts is the minimum number of points in the neighborhood, and Eps is the neighborhood radius. Determining Eps and MinPts directly affects the clustering effect of density-based spatial clustering with noise.

[0112] A smaller MinPts value makes the algorithm more sensitive to noise but may generate too many data clusters. A larger value makes the algorithm more robust to noise but may ignore smaller data clusters. Generally speaking, the MinPts value is greater than or equal to the data dimension plus one. For a two-dimensional wind power dataset, this threshold can be set to 3. If the Eps value is too large, too many data points will be assigned to the same cluster. If the Eps value is too small, all data points will be marked as noise, with no core points. Currently, determining the Eps value is not only error-prone and inefficient, but also lacks correlation between data groups.

[0113] To this end, this embodiment improves the DBSCAN algorithm, visualizes the wind power data, calculates the k-nearest neighbor distance, draws a k-distance graph to identify the density distribution of the data set, and adaptively calculates the Eps value using the slope mutation method. The calculation steps are as follows:

[0114] Step S2.2.1: using the wind speed-power real-time data of the SCADA system after the first cleaning as the second data set;

[0115] Step S2.2.2: For each point in the second dataset, calculate the distance from any point M to all other points, output a distance graph from any point M to all other points, and find the k nearest neighbors from any point M. Record the distance from each point in the second dataset to its corresponding k-th nearest neighbor, called the k-nearest neighbor distance;

[0116] In this embodiment, for each point in the data set, the distance from it to all other points is calculated, and a distance graph from a certain point to all other points is output. Its k nearest neighbor points are found, and the distance from each point to its kth nearest neighbor point is recorded, which is called the k nearest neighbor distance.

[0117] Step S2.2.3: Sort the k-nearest-neighbor distances of all points in the second dataset in ascending order, number all points in the second dataset according to the size of the k-nearest-neighbor distances, and draw a graph of the relationship between the sorted k-nearest-neighbor distances and the point numbers, which is called the k-nearest-neighbor distance graph;

[0118] In this embodiment, the k-nearest neighbor distances of all points are sorted in ascending order, and then a relationship diagram between the sorted k-nearest neighbor distances and the index of the point is drawn, as shown in FIG. Figure 5 As shown, Figure 5 The x-coordinate refers to the real-time wind power data points of the SCADA system, that is, the number of data points in the wind power-wind speed curve that has not been standardized.

[0119] Step S2.2.4: The k-nearest neighbor distance corresponding to the inflection point in the k-nearest neighbor distance graph is used as the neighborhood radius of the density-based spatial clustering method with noise. The data points before the inflection point belong to the data cluster, and the data points after the inflection point belong to the noise cluster. The data points before the inflection point are used as the real-time wind speed-power data of the SCADA system after the second data cleaning.

[0120] In this embodiment, an "elbow point" is often observed in the k-nearest neighbor distance graph, i.e., a point where the slope of the curve changes significantly. The k-nearest neighbor distance corresponding to this point can be used as an estimate of the EPS threshold. This point appears as an inflection point in the k-nearest neighbor distance graph, where the slope changes suddenly. Data points before the inflection point belong to the dense area, while data points after the inflection point belong to the sparse area or noise points.

[0121] Comparing the Spearman coefficients of the original data, the Spearman coefficients of the data after the first cleaning, and the Spearman coefficients of the data after the second cleaning, they are 0.705, 0.740, and 0.939 respectively. The wind power data before and after the algorithm improvement are shown in the figure below. Figure 6 、 Figure 7 As shown in the figure, the DBSCAN algorithm cannot effectively divide the data clusters for the bottom accumulation data and the middle accumulation data, resulting in poor cleaning quality. The cleaning quality of the wind power data is improved by improving the algorithm. The middle accumulation data has been effectively cleaned and is significantly more concentrated around the standard power curve, indicating that the secondary cleaning quality is better.

[0122] Step S3: Compare the real-time wind speed-power data after the last cleaning with the standard wind speed-power curve to obtain a preliminary damage status of the wind turbine blade;

[0123] In this example, comparing the cleaned wind power data with the standard wind power curve reveals that the wind speed at which the power data in the dataset reaches its maximum value is significantly lower than the rated wind speed. This means that the power limit is reached at a lower wind speed, and the cut-in wind speed is significantly higher than that of the standard power curve. This indicates that a higher wind speed is required to achieve the same output power. This suggests that the wind turbine blades' wind energy capture efficiency has decreased, and internal blade damage may have occurred. The densely packed central scatter points in the wind power data indicate blade damage and extreme weather conditions, requiring further identification of the specific type and process of blade damage.

[0124] Step S4: If the preliminary wind turbine blade damage status is damaged, real-time acoustic emission signals are collected;

[0125] Step S5: Input the real-time acoustic emission signal into a pre-established wind turbine blade early damage prediction model to obtain the type and degree of wind turbine blade early damage identification. Wind turbine blade early damage refers to damage identified during the wind turbine blade's operational warranty period. The pre-established wind turbine blade early damage prediction model is established by collecting acoustic emission signals through a fatigue loading test of a blade specimen, where the material of the blade specimen is consistent with the wind turbine blade to be identified. Peak frequency and amplitude are used as feature values. The acoustic emission signals collected through the fatigue loading test are classified using an improved K-means algorithm, and a convolutional neural network is constructed to obtain the pre-established wind turbine blade early damage prediction model.

[0126] Acoustic emission signals are collected through fatigue loading tests on blade specimens. The material of the blade specimens is consistent with the wind turbine blade to be identified. Peak frequency and amplitude are used as feature values. The acoustic emission signals collected from the fatigue loading tests are classified using an improved K-means algorithm. A convolutional neural network is constructed to obtain a pre-established wind turbine blade early damage prediction model, including:

[0127] Step S5.1: Acquiring acoustic emission signals through a fatigue loading test of a blade specimen, where the material of the blade specimen is consistent with the wind turbine blade to be identified;

[0128] In this embodiment, a blade specimen is prepared using a vacuum assisted resin infusion (VARI) process to prepare a composite laminate specimen made of the same material as the blade, and an acoustic emission signal is collected during a fatigue loading test.

[0129] The specific method of preparing the test piece is as follows: lay the pre-prepared fiber layer on the bottom mold, cover the template after laying, and ensure that the lower mold and the upper mold can be completely closed; use a vacuum bag to wrap and seal, and then use a vacuum pump to evacuate to a negative pressure state; inject the resin liquid into the entire mold through the grease injection port to allow the resin to fully penetrate the layered structure; heat and cure at 70 to 80 degrees Celsius and then seal for a period of time. The composite laminate test piece used in the experiment is as follows: Figure 8 As shown in the figure, the specimen's length × width × height is 250 mm × 25 mm × 3.8 mm. To prevent the specimen from slipping relative to the fatigue testing machine fixture or from being directly crushed at both ends during the fatigue test, reinforcement treatment was implemented at 50 mm on both sides of the specimen.

[0130] The specific materials used in this example are: unidirectional glass fiber cloth (ECW600-1270, 600 g / m²) was selected as the ply material. The epoxy resin used in the preparation of the glass fiber reinforced plastic (GFRP) composite material is Araldite LY 1564 Sp, with a density of 1.1-1.2 g / cm³. The curing agent is Aradur 3486, with a density of 0.94-0.95 g / cm³. The mass ratio of epoxy resin to curing agent in the resin matrix is ​​100:34. Since epoxy resin is a thermosetting resin, the resin solution and curing agent should be thoroughly stirred and mixed before heating during the preparation process.

[0131] Fatigue loading test and acoustic emission signal collection: The fatigue testing machine starts loading and the acoustic emission detection system is started. Real-time monitoring and data collection can be found. The channel window connected to the sensor begins to show acoustic emission signals (voltage fluctuations). Fatigue load is continuously applied to the specimen until the composite specimen breaks, which is considered the end of the experiment. An overview of the equipment used in the experiment is shown below: Figure 9 shown.

[0132] The sensor used in the experiment was an RS-2A sensor with a diameter of 18.8 mm, a height of 15 mm, an M5-KY connector, and a ceramic sensing surface. Vaseline was used as a coupling agent to connect the sensor to the sensor base. A mixture of AB glue and glue was applied to the sensor base in equal proportions. The two sensors were 100 mm apart and 50 mm from the centerline of the specimen. To protect the acoustic emission sensor, it was placed on the sensor base before installation. Vaseline was applied between the sensor and the base to prevent the sensor base from affecting the accuracy of the acoustic emission data acquisition. To protect the sensor from potential damage during specimen failure and to minimize signal attenuation in the sensor base, the preamplifier used in the acoustic emission experiment had a signal gain of 40 dB (100x), which reduced attenuation and noise during signal transmission and provided initial amplification of weak acoustic emission signals. The AE experiments used a DS2-8A full-information AE analyzer, supporting up to eight channels of simultaneous data acquisition, with a continuous data throughput rate of 131MB / s and a waveform data throughput rate of 96MB / s. This instrument was used to monitor AE activity during the loading process in real time and store the signals. The AE detection system, a SAEU2S multi-channel AE monitoring system, offers functions such as AE signal parameters, waveform display, rough AE source location, and basic spectrum analysis.

[0133] At the same time, to minimize the impact of fatigue testing machine vibration on acoustic emission signal acquisition, direct sampling is required prior to acquisition to analyze inherent noise and set an appropriate threshold value. This ensures that the sensor can stably and accurately acquire acoustic emission signals. The composite material specimens used in this experiment typically take more than four hours to break under fatigue load. If continuous sampling is used, the amount of data collected would be too large for the equipment to handle and would be inconvenient for subsequent data processing and analysis. Furthermore, the collected acoustic emission signals may contain excessive noise interference, making it difficult to effectively identify damage signals. Therefore, based on the actual experimental conditions, the sensor acquisition trigger mode for this experiment was set to threshold trigger.

[0134] Before the experiment, a pass-through method (without setting a threshold) was used to acquire the signal. The impact of noise on the acquired AE signal was analyzed. It was found that setting the threshold to 50mV and the preamplifier gain to 40dB effectively filtered out most ambient noise. Three key parameters for AE acquisition signals are Peak Definition Time (PDT), Hit Definition Time (HDT), and Hit Lockout Time (HLT). For an AE system, the time required to acquire an impact signal is: T = Signal Duration (Duration) + Hit Definition Time (HDT) + Hit Lockout Time (HLT). Before performing pass-through sampling on the specimen, three sets of values ​​were tested for each interval. The results showed that setting the PDT, HDT, and HLT values ​​to 50us, 200us, and 300us, respectively, yielded the best sampling results. The AE signal sampling parameter settings are summarized in Table 1.

[0135] Step S5.2: pre-processing the collected acoustic emission signals;

[0136] In this embodiment, by analyzing the collected acoustic emission signals, a total of 12 parameters of usable signal segments are extracted. Each signal segment includes amplitude, duration, rise time, ring count, rise count, energy, RMS, ASL, number of impacts, impact rate, centroid frequency, and peak frequency, i.e., 12-dimensional data. However, the data set cannot be used directly and needs to be cleaned first to deal with missing values, duplicate values, and erroneous values ​​(logical errors, labeling errors, and other errors) in the data.

[0137] Step S5.3: using principal component analysis to perform dimensionality reduction on the pre-processed acoustic emission signal;

[0138] Table 1 Acoustic emission signal acquisition parameters;

[0139] ;

[0140] In this embodiment, principal component analysis (PCA) is a commonly used method for data dimensionality reduction and is a multivariate statistical method.

[0141] Each variable is first dimensionless (i.e., standardized). After the variables are standardized to ensure the availability of the data set, the data in the acoustic emission signal parameter table is reduced in dimension according to the following steps:

[0142] (1) Select the covariance matrix or correlation matrix and calculate the characteristic roots and corresponding eigenvectors of the acoustic emission signal parameter data set;

[0143] (2) Calculate the variance contribution rate of each acoustic emission parameter dimension data, and select the number of principal components that can represent information greater than the threshold value according to the threshold of its variance contribution rate;

[0144] (3) Reduce the data dimension and save it as a new data file.

[0145] The characteristic parameters of AE signals after PCA data cleaning are summarized in Table 2. After data dimensionality reduction and data elimination, the total number of data is 91,621.

[0146] Table 2 AE signal characteristic parameters;

[0147] ;

[0148] Step S5.4: Using the peak frequency and amplitude as feature values, the improved K-means algorithm is used to classify the AE signal after dimension reduction.

[0149] In this example, the collected raw acoustic emission (AE) signals were preprocessed using PCA (Principal Component Analysis) to reduce dimensionality and eliminate preprocessing errors. A total of 91,621 data points were obtained, as shown in Table 2. Table 2 contains numerous characteristic parameters for the AE signals. In this example, only two parameters, peak frequency and amplitude, were used for clustering, while the other parameters were not used. First, theoretical analysis shows that the stress distribution characteristics of early blade damage are closely related to these two parameters. Second, extensive experiments have shown that clustering using peak frequency and amplitude provides the best performance and time efficiency compared to other parameters. Therefore, in this example, only peak frequency and amplitude were used as characteristic values. Due to the large volume of AE signal parameter data, it is not possible to label the data row by row and directly train the model using supervised learning. Therefore, an unsupervised learning approach using the K-means algorithm was used for data classification. For datasets containing several unlabeled data, the datasets were divided into multiple clusters based on their inherent similarity, ensuring high similarity within the same cluster and low similarity between clusters. The elbow method is usually used to select the optimal number of clusters k. The core indicator of the elbow method is SSE (sum of the squared errors), as shown in formula (4):

[0150] (4);

[0151] Among them, Ci represents the I-th cluster, p represents the sample point in Ci, m i represents the centroid of Ci (the mean of all samples in Ci), SSE represents the sum of squared errors of all samples, and represents the quality of the clustering result. The optimal number of clusters k is the number corresponding to the "elbow part" of the image.

[0152] K-means can only be used when the average value of the cluster can be defined. It is not suitable for experimental data sets and is sensitive to initial values. Different initial values ​​may lead to different results. In addition, this algorithm is also sensitive to noise and isolated data. To effectively solve the problem of inconsistent cluster centers, an improved K-means algorithm is used, including:

[0153] Step S5.4.1: using the acoustic emission signal after dimension reduction as the third data set;

[0154] Step S5.4.2: Randomly select a data point from the third data set as the first cluster center;

[0155] In this embodiment, the first center point is randomly selected. The first cluster center is randomly selected from the data set as the initial seed point.

[0156] Step S5.4.3: Calculate the shortest distance D(x) from each point in the third data set to the first cluster center;

[0157] Step S5.4.4: Select a new point from the third data set as the second cluster center, and the probability of selection is the same as D(x) 2 proportional to;

[0158] In this embodiment, distances are calculated and the next center point is determined. For each point in the dataset, the shortest distance between it and the selected center point is calculated. Points with greater distances have a higher probability of being selected as the next center point. That is, points farther from the existing cluster center are more likely to be selected as the new cluster center. This probability-weighted mechanism ensures a more reasonable distribution of initial center points.

[0159] Step S5.4.5: Repeat steps S5.4.2 to S5.4.4 until K cluster centers are selected;

[0160] In this embodiment, the selection process is repeated until initialization is complete. Following the above rules, K initial centers (i.e., K cluster centers) are selected in sequence. This step effectively avoids the "center point clustering" problem that can occur with traditional K-means by maximizing the dispersion between the initial centers.

[0161] Step S5.4.6: Using the selected K cluster centers as initial values, calculate the distance from each data point in the third data set to the K cluster centers, assign each data point to the cluster with the nearest cluster center, update the cluster center, and repeat step S5.4.6 until the clustering results converge.

[0162] In this embodiment, standard K-means clustering is performed. After obtaining the optimized initial center point, the conventional K-means algorithm is iterated: 1) each data point is assigned to the cluster to which the nearest center point belongs; 2) the center point (mean) of each cluster is recalculated; 3) the above process is repeated until the clustering result is stable. Using the selected K cluster centers as the initial value, the distance from each data point to the K cluster centers is calculated, and the data point is assigned to the cluster with the nearest cluster center. The cluster centers are then updated, and this process is repeated until the clustering result converges. The specific updating of the cluster centers is as follows: first, assume how many classes there are, then loop through each class and cluster once. Finally, according to the quality SSE index of the clustering result, the assumed class is determined to be the best, and convergence is completed.

[0163] The improved K-means algorithm (abbreviated as K++ algorithm) has the following advantages: (1) K-means randomly selects the initial cluster center, which may cause the result to fall into the local optimum. The K++ algorithm selects new cluster centers based on the distance between the data point and the selected cluster center, making the initial cluster center distribution more uniform and more representative of the data distribution, thereby improving the quality and stability of the clustering results. (2) Due to the optimization of the initial cluster center, the clustering results obtained by the K++ algorithm are usually closer to the actual distribution of the data, can better identify different categories in the data, and reduce clustering errors. (3) The initial cluster center selected by K++ is more reasonable, so that the algorithm can converge to a better result faster during the iteration process, reduce the number of iterations, and improve computational efficiency. This embodiment uses the K++ algorithm to obtain the type of damage mode.

[0164] By calculating SSE, the result is as follows Figure 10 As shown, the elbow method can be used to determine the position of the "elbow" at 3, i.e., the optimal value of K for this data set is 3, which divides the data into three clusters. This embodiment uses the elbow method to obtain the proportion of different types of injury patterns, which helps to assess the extent of early-stage injury.

[0165] After dimensionality reduction preprocessing, the acoustic emission data was imported into the unsupervised learning program based on the K-means++ algorithm. The data set was clustered and labeled (Cluster-1, Cluster-2, Cluster-3). The data volume of each label is shown in Table 3. The peak frequency and amplitude are selected to more intuitively display the clustering results. Figure 11 The clustering of AE signals is shown in the figure. It can be seen from the figure that the channel signals are well divided into 3 categories.

[0166] Table 3 Data categories after clustering;

[0167] ;

[0168] By analyzing the peak frequencies of the three cluster signals obtained after clustering the data set, in the three frequency bands of (100-300kHz), (300-420kHz), and (420-580kHz), respectively, it was identified that the damage mode of the damage signal in Cluster-1 with a frequency band of (100-300kHz) was matrix cracking, and it was inferred that the damage mode of the damage signal in Cluster-3 with a frequency band of (420-580kHz) was fiber breakage. In addition to matrix cracking and fiber breakage, a more common damage mode was delamination, so it was inferred that the damage mode of the damage signal in Cluster-2 with a frequency band of (300-420kHz) was delamination.

[0169] According to the corresponding intervals of the peak frequencies of several common damage types: matrix cracking (100-300kHz), delamination (300-420kHz) and fiber breakage (420-580kHz), that is, by extracting the peak frequency of a certain acoustic emission signal of the wind turbine blade, the damage mode can be identified through the interval where the peak frequency is located.

[0170] Step S5.5: Construct a convolutional neural network using the classified acoustic emission signals to obtain a pre-established wind turbine blade early damage prediction model.

[0171] In this example, the raw data set of AE damage signals for composite materials under tension-tension fatigue loading totaled 101,919 sets. After data preprocessing and removal of unusable data, 91,621 sets of usable labeled data remained. The data set was clustered into three categories based on peak frequency and amplitude: 16,231 sets of acoustic emission data for the first category, 36,657 sets for the second category, and 38,733 sets for the third category. The data volumes for each category are shown in Table 3. Considering the input characteristics of the convolutional neural network, peak frequency and amplitude were selected as feature values ​​to establish a sample database for training the convolutional neural network. Considering the total data volume of the sample database, the training set and test set data ratio was set to 8:2. This means that 80% of the labeled data in the sample database was used as the training set for training the neural network, and 20% of the labeled data in the sample database was used as the test set for evaluating the accuracy of the neural network.

[0172] Learning rate parameters for convolutional neural networks:

[0173] The learning rate determines the step size for parameter updates during neural network training. A larger learning rate allows the model to converge quickly, but it can also result in missed optimal solutions or even in non-convergence, reducing model stability. A smaller learning rate allows the model to find the optimal solution more accurately, but it also increases training time, consumes more computing resources, and slows convergence. The initial learning rate for neural networks is typically set to 0.01. A small-scale experiment using a sample database revealed that when the initial learning rate was set to 0.01, the model's loss curve exhibited persistent fluctuations, lacking a clear, sustained downward trend. Therefore, to ensure computational efficiency and stability of the neural network, a lower learning rate was considered. In a small-scale experiment, the initial learning rate was set to 0.001, resulting in good convergence and a clear downward trend in the loss curve. The batch size was set to 64, the maximum number of training rounds was set to 20, the maximum number of iterations was set to 22,900, and the number of iterations per round was set to 1,145.

[0174] The loss function of convolutional neural network is:

[0175] The loss function is generally expressed as L(y, f(x)), which is an indicator that measures the difference between the model's predicted value f(x) and the true value y. The smaller the value, the better the model's prediction effect, and vice versa. The commonly used loss functions are mainly the following. Mean square error (MSE), mean absolute error (MAE), and cross entropy loss function. Among them, the cross entropy loss function can well reflect the gap between the predicted result and the true label in the classification problem, and the gradient calculation is relatively stable, and the problem of gradient disappearance will not occur. Therefore, this study uses the cross entropy loss function to calculate the error between the predicted result and the true label, and uses this as a basis to update the parameters of the convolutional neural network to optimize the performance of the model. The cross entropy loss function loss is expressed as:

[0176] (5);

[0177] Among them, n is the number of samples, y i and p i are the true value of the i-th sample and the predicted value of the convolutional neural network, respectively.

[0178] Activation function of convolutional neural network:

[0179] Convolution operation is a linear operation. If only linear convolution layers and fully connected layers are used to form a network, the output data of the last layer is essentially a linear combination of the input layers. The superposition of a large number of linear operations is still linear. In order to solve the problem of nonlinear data association in acoustic emission signals, a nonlinear activation function is introduced between the convolution layer and the pooling layer of the convolutional neural network to perform nonlinear fitting on the input data, so that different layers of the convolutional neural network can have nonlinear mapping learning capabilities. Commonly used activation functions include the Sigmoid function and the Tanh function, which cause the gradient to disappear. The present invention uses the ReLU function to output 0 when the input value is less than 0, which increases the sparsity of the network, helps the convergence of the stochastic gradient descent method, and speeds up the training speed of the network. When the input value is greater than 0, the ReLU function is a linear unit, and the output value is equal to the input value. At this time, the derivative is always equal to 1, and the gradient saturation phenomenon will not occur, solving the problem of gradient disappearance of the Tanh function and the Sigmoid function. When the input value is equal to 0, the ReLU function is not differentiable, the function form is simple, the computational complexity is very small, and it is easy to implement in hardware. The calculation of ReLU is shown in formula (6):

[0180] (6);

[0181] Determine the training parameters such as the learning rate of the convolutional neural network damage model, select the cross-entropy loss function and activation function, add labels to the acoustic emission experimental data through cluster analysis, divide the experimental data with existing labels into data sets, establish a sample database, train and test the convolutional neural network, and construct a wind turbine blade main beam damage identification model.

[0182] During the actual calculation process, acoustic emission sample data is input, processed through convolutional layers, pooling layers, and fully connected layers, and the results are output to the output layer. During this process, if the output result does not meet the maximum number of iterations or minimum error requirements, backpropagation is used to iterate the calculation again. The optimizer identifies the error magnitude and uses this as a basis to adjust the weights and thresholds until the error meets the expected value. Through repeated iterative training, a neural network model with recognition and prediction capabilities is formed.

[0183] Convolutional neural network training consists of two phases: forward computation and backpropagation. Forward computation involves extracting features from the input data through a series of operations, such as convolution, pooling, and nonlinear activation, to calculate the corresponding predicted output. Backpropagation compares the predicted output with the expected output to calculate the loss function. This loss function is then propagated back through each layer, using gradient descent to adjust the network parameters. This forward computation and error backpropagation cycle repeats until the network model converges.

[0184] The steps of convolutional neural network training are as follows. Figure 12 shown.

[0185] (1) Dataset division: 80% of the total dataset is divided into a training set and 20% into a test set. The training set is used to train the convolutional neural network, and the test set is used to evaluate the performance of the model until the model recognition accuracy meets the requirements. Finally, the accuracy and error of the trained model are evaluated.

[0186] (2) Network weight initialization, that is, giving each weight parameter in the network an initial value. During the training process of the convolutional neural network, the appropriateness of weight initialization has a significant impact on the results. Inappropriate initialization parameters may cause the convolutional neural network to converge slowly or even fail to converge during training. Common initialization methods include pre-training initialization, random initialization, fixed value initialization, etc.

[0187] (3) The divided training set is used as the input of the convolutional neural network. The peak frequency and amplitude have the greatest impact on damage pattern recognition. This paper selects these two types of data to train the convolutional neural network and obtains the output value of the convolutional neural network through forward calculation.

[0188] (4) Compare the output value of the convolutional neural network with the expected output value and calculate the loss function. If the maximum number of iterations and minimum error requirements are met, save the model and end the operation. If one of the requirements is not met, proceed to the next step.

[0189] (5) Use the optimizer to backpropagate the loss function to update the neural network parameters, and iterate until the maximum number of training rounds is reached or the accuracy requirements are met. Stop training and save the network model parameters.

[0190] The accuracy of the model training process is visualized to timely understand the convergence of the model. The training is terminated when the minimum expected error is met or the maximum number of iterations is reached. According to the change in the proportion of the number of correctly identified samples to the total number of samples, the training accuracy curve of the damage identification model is drawn as follows: Figure 13 shown.

[0191] The cross entropy loss function is used to calculate the error between the recognition prediction result and the true label. The smaller the loss value, the better the recognition prediction effect of the model. Conversely, the worse the recognition prediction effect of the model. Record the loss value during the model training process and draw the loss curve of the damage recognition model as shown in the figure. Figure 14 shown.

[0192] Analysis of the damage identification model's training accuracy reveals that with increasing iterations, the model's accuracy rapidly increases, then slowly grows and gradually levels off, with no significant increases or fluctuations, indicating that the model has converged. High accuracy is maintained throughout the training process, demonstrating that the model can accurately identify the inherent correlations within the data in the sample database. Analysis of the damage identification model's training loss reveals that the loss value is high at the beginning of training, reaching 1.2. As the number of iterations increases, the model's loss value continuously decreases and approaches 0, indicating that the model's recognition predictions are gradually approaching the true labels. Analysis of the model's accuracy and loss curves provides real-time insights into the model's learning progress. The high recognition accuracy and steadily decreasing loss value demonstrate the model's strong generalization capabilities.

[0193] The performance of the blade main beam damage prediction model is evaluated using a confusion matrix. The relationship between the model prediction results and the actual results is displayed in the form of a matrix. The learning curve is used to analyze how the performance of the machine learning model changes with the amount of training data or the number of training rounds.

[0194] The convolutional neural network has excellent recognition and classification capabilities, with an overall recognition accuracy of 0.999945 for this test set. The recognition accuracy for each category is as follows: 3275 data points for matrix cracking in the test set had a recognition accuracy of 0.9999; 7337 data points for interface delamination had a recognition accuracy of 1; and 7712 data points for fiber fracture had a recognition accuracy of 1. To intuitively reflect the classification performance of the damage recognition model, the confusion matrix of the model is shown below. Figure 15 shown.

[0195] For the pre-established wind turbine blade early damage prediction model of this embodiment, the established sample database mainly consists of three types of damage patterns. Therefore, the confusion matrix of this model is a 3×3 matrix, where each row of the matrix represents the true category, each column represents the predicted category, the elements on the diagonal represent the number of correctly predicted samples, and the elements on the off-diagonal represent the number of incorrectly predicted samples. In the sample library for training convolutional neural networks, the amount of matrix cracking, i.e., the first type of AE data, is relatively small. Therefore, when establishing the damage recognition model, one recognition error occurred in the recognition of matrix cracking damage, which is reflected in the identification of matrix cracking damage as fiber fracture damage in the test set. However, the amount of data on interface delamination and fiber fracture is relatively large, which has a better generalization effect, and no recognition errors occurred in the test set.

[0196] Validation loss is typically plotted against the amount of training data or the number of training rounds, and against the model's performance metrics (such as accuracy and loss function) on the training and validation sets. During machine learning model training, validation loss is used to evaluate the model's performance on the validation dataset. Validation loss is calculated by applying the trained model to the validation set and applying a specific loss function. Validation loss is calculated by averaging the model's loss across all samples in the test set.

[0197] In order to analyze the training process and performance of the damage identification model at different stages, the data in the sample database is divided into 10 parts, the amount of data is increased in sequence, the loss value of each round of training is recorded, and the learning curve of the model is drawn as shown in the figure. Figure 16 shown.

[0198] As the amount of training data increases, the model's performance on the training set continues to improve, with the training loss showing a clear downward trend, gradually decreasing from a peak of 0.055 to 0.005. This demonstrates that the damage identification model has sufficient learning ability to fit the patterns and regularities in the training data. Furthermore, the validation loss is close to 0, while the training loss is low, and the model performs stably during training and validation. This indicates that the damage identification model has learned the inherent regularities in the sample database, outperforming the training set on the test set, demonstrating good generalization ability and the ability to accurately identify and predict new data.

[0199] Example 2:

[0200] This embodiment provides a wind turbine blade early damage identification system, comprising: a data acquisition module, a data cleaning module, a preliminary damage identification module, an acoustic emission signal acquisition module, and a damage type judgment module. The data acquisition module is connected to the data cleaning module, the data cleaning module is connected to the preliminary damage identification module, the preliminary damage identification module is connected to the acoustic emission signal acquisition module, and the acoustic emission signal acquisition module is connected to the damage type judgment module.

[0201] The data acquisition module is used to obtain the real-time wind speed and power data of the SCADA system for the wind turbine blade to be identified;

[0202] A data cleaning module is used to perform two data cleanings on the wind speed-power real-time data of the SCADA system;

[0203] The initial damage identification module is used to compare the real-time wind speed-power data after the last cleaning with the standard wind speed-power curve to obtain the initial damage status of the wind turbine blade;

[0204] An acoustic emission signal acquisition module is used to collect real-time acoustic emission signals if the preliminary wind turbine blade damage status is damaged;

[0205] The damage type judgment module is used to input the real-time acoustic emission signal into the pre-established wind turbine blade early damage prediction model to obtain the type and degree of wind turbine blade early damage identification. Wind turbine blade early damage refers to damage identified during the wind turbine blade's operational warranty period. The pre-established wind turbine blade early damage prediction model is established through the following process: collecting acoustic emission signals through fatigue loading tests on blade specimens. The material of the blade specimens is consistent with the wind turbine blade to be identified. Peak frequency and amplitude are used as feature values. The acoustic emission signals collected in the fatigue loading test are classified using an improved K-means algorithm, and a convolutional neural network is constructed to obtain the pre-established wind turbine blade early damage prediction model.

[0206] Example 3:

[0207] This embodiment provides a computer program product, including a computer program or instructions. When the computer program or instructions are executed by a processor, the method for identifying early damage to a wind turbine blade is implemented.

[0208] Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution can be embodied in the form of a computer program product.

[0209] The various embodiments in this application are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

[0210] The scope of protection of this application is not limited to the above-described embodiments. Obviously, those skilled in the art may make various modifications and variations to this disclosure without departing from the scope and spirit of this disclosure. If such modifications and variations fall within the scope of equivalent technologies of this disclosure, then the disclosure is intended to include such modifications and variations.

Claims

1. A method for identifying early damage to wind turbine blades, characterized in that: include: For the wind turbine blades to be identified, obtain the real-time wind speed and power data from the SCADA system; Performing two data cleanings on the wind speed-power real-time data of the SCADA system; Compare the real-time wind speed-power data after the last cleaning with the standard wind speed-power curve to obtain the preliminary damage status of the wind turbine blades; If the initial wind turbine blade damage status is damaged, real-time acoustic emission signals are collected; The real-time acoustic emission signal is input into a pre-established wind turbine blade early damage prediction model to obtain the type and degree of wind turbine blade early damage identification. Wind turbine blade early damage refers to damage identified during the operational warranty period of the wind turbine blade. The pre-established wind turbine blade early damage prediction model is established by collecting acoustic emission signals through fatigue loading tests on blade specimens made of the same material as the wind turbine blade to be identified. Peak frequency and amplitude are used as feature values. The acoustic emission signals collected from the fatigue loading test are classified using an improved K-means algorithm, and a convolutional neural network is constructed to obtain the pre-established wind turbine blade early damage prediction model.

2. The method for identifying early damage to a wind turbine blade according to claim 1, characterized in that: The wind speed-power real-time data of the SCADA system is cleaned twice, including: The first data cleaning of the wind speed-power real-time data of the SCADA system was performed, including: using the threshold setting method to remove the abnormal data accumulated at the bottom of the wind speed-power real-time data of the SCADA system; using the density-based spatial clustering method with noise to remove the abnormal data sparse in the middle of the wind speed-power real-time data of the SCADA system; If after the first data cleaning, the first remaining data volume, the first data rejection rate, the first cleaning time, and the Spearman coefficients of the wind speed-power real-time data before and after the first data cleaning all reach the first data cleaning threshold, then a second data cleaning is performed on the wind speed-power real-time data of the SCADA system, including: using an improved density-based spatial clustering method with noise to perform a second data cleaning on the central accumulated abnormal data in the wind speed-power real-time data of the SCADA system.

3. The method for identifying early damage to a wind turbine blade according to claim 2, characterized in that: The abnormal data accumulated at the bottom is defined as: the difference between the wind power value corresponding to the wind speed in the wind speed-power real-time data of the SCADA system and zero is less than or equal to the wind power threshold, the number of wind speed-power real-time data located at the bottom of the wind speed-power curve is greater than the number threshold, and the density of the wind speed-power real-time data is greater than the density threshold and is distributed in a long strip shape; The middle sparse abnormal data is defined as: the wind speed-power real-time data is sparsely distributed in the middle area of ​​the wind speed-power curve and below the standard wind speed-power curve, and the density of the wind speed-power real-time data is less than or equal to the density threshold; The middle accumulation abnormal data is defined as: the wind speed-power real-time data of the SCADA system is located in the middle area of ​​the wind speed-power curve, and the density of the wind speed-power real-time data is greater than a density threshold.

4. The method for identifying early damage to a wind turbine blade according to claim 2, characterized in that: The method of using density-based spatial clustering method with noise to remove sparse abnormal data in the middle of the wind speed-power real-time data of the SCADA system includes: Step S2.2.1: Initialize the wind speed-power real-time data of the SCADA system, use the wind speed-power real-time data of the SCADA system as the first data set, and mark all data in the first data set as unprocessed; Step S2.2.2: Traverse the first data set and, for each unprocessed data point, perform the following operations: mark the current data point as processed data, check all data points within the neighborhood of the current data point, where the neighborhood is the circle centered on the current data point and the neighborhood radius is the radius of the circle. The resulting circular area is called the neighborhood. If the number of all data points within the neighborhood point is less than the minimum number of points, mark the current data point as a noise point. Step S2.2.3: If the number of all data in the neighborhood is greater than or equal to the minimum number of points, mark the current data as a core point and create a new cluster with the new cluster as the current cluster; Step S2.2.4: For each point in the neighborhood of the core point, perform the following operations: If the current point has not been processed, mark the current point as processed and check the points in the neighborhood of the current point; if the number of points in the neighborhood of the current point is greater than or equal to the minimum number of points, and the current point has not been assigned to any cluster, add the current point to the current cluster and repeat step S2.2.4 until there are no new points that can be added to the current cluster; Step S2.2.5: Repeat steps S2.2.2 to S2.2.4 until all data in the first data set are marked as processed data; Step S2.2.6: Output the clustering results, including: noise clusters composed of all noise points and data clusters composed of each core point. Remove all noise clusters and use the data clusters composed of each core point as the real-time wind speed-power data after the first data cleaning.

5. The method for identifying early damage to a wind turbine blade according to claim 2, characterized in that: The improved density-based spatial clustering method with noise is used to perform a second data cleaning on the central accumulation abnormal data in the wind speed-power real-time data of the SCADA system, including: The wind speed-power real-time data of the SCADA system after the first cleaning is used as the second data set; For each point in the second data set, calculate the distance from any point M to all other points, output the distance graph from any point M to all other points, and find the k nearest neighbor points from any point M. Record the distance from each point in the second data set to the corresponding k-th nearest neighbor point, which is called the k-nearest neighbor distance; Sort the k-nearest-neighbor distances of all points in the second dataset in ascending order, and number all points in the second dataset according to the size of the k-nearest-neighbor distances. Draw a relationship graph between the sorted k-nearest-neighbor distances and the point numbers, which is called a k-nearest-neighbor distance graph. The k-nearest neighbor distance corresponding to the inflection point in the k-nearest neighbor distance graph is used as the neighborhood radius of the density-based spatial clustering method with noise. The data points before the inflection point belong to the data cluster, and the data points after the inflection point belong to the noise cluster. The data points before the inflection point are used as the real-time wind speed-power data of the SCADA system after the second data cleaning.

6. The method for identifying early damage to a wind turbine blade according to claim 1, characterized in that: Acoustic emission signals are collected through fatigue loading tests on blade specimens. The material of the blade specimens is consistent with the wind turbine blade to be identified. Peak frequency and amplitude are used as feature values. The acoustic emission signals collected from the fatigue loading tests are classified using an improved K-means algorithm. A convolutional neural network is constructed to obtain a pre-established wind turbine blade early damage prediction model, including: Acquiring acoustic emission signals through a fatigue loading test of a blade specimen, the material of which is consistent with the wind turbine blade to be identified; Preprocessing the collected acoustic emission signals; Principal component analysis is used to reduce the dimension of the pre-processed acoustic emission signals; The peak frequency and amplitude are used as feature values, and the improved K-means algorithm is used to classify the acoustic emission signals after dimension reduction. The classified acoustic emission signals are used to construct a convolutional neural network to obtain a pre-established wind turbine blade early damage prediction model.

7. The method for identifying early damage to a wind turbine blade according to claim 6, characterized in that: The improved K-means algorithm includes: Step S5.4.1: using the acoustic emission signal after dimension reduction as the third data set; Step S5.4.2: Randomly select a data point from the third data set as the first cluster center; Step S5.4.3: Calculate the shortest distance D(x) from each point in the third data set to the first cluster center; Step S5.4.4: Select a new point from the third data set as the second cluster center, and the probability of selection is the same as D(x) 2 proportional to; Step S5.4.5: Repeat steps S5.4.2 to S5.4.4 until K cluster centers are selected; Step S5.4.6: Using the selected K cluster centers as initial values, calculate the distance from each data point in the third data set to the K cluster centers, assign each data point to the cluster with the nearest cluster center, update the cluster center, and repeat step S5.4.6 until the clustering results converge.

8. The method for identifying early damage to a wind turbine blade according to claim 6, characterized in that: The activation function of the convolutional neural network is the ReLU function. When the input value of the ReLU function is less than zero, the output of the ReLU function is 0. When the input value of the ReLU function is greater than or equal to zero, the output value of the ReLU function is equal to the input.

9. A wind turbine blade early damage identification system, characterized in that: include: The data acquisition module is used to obtain the real-time wind speed and power data of the SCADA system for the wind turbine blade to be identified; A data cleaning module is used to perform two data cleanings on the wind speed-power real-time data of the SCADA system; The initial damage identification module is used to compare the real-time wind speed-power data after the last cleaning with the standard wind speed-power curve to obtain the initial damage status of the wind turbine blade; An acoustic emission signal acquisition module is used to collect real-time acoustic emission signals if the preliminary wind turbine blade damage status is damaged; The damage type judgment module is used to input the real-time acoustic emission signal into the pre-established wind turbine blade early damage prediction model to obtain the type and degree of wind turbine blade early damage identification. Wind turbine blade early damage refers to damage identified during the wind turbine blade's operational warranty period. The pre-established wind turbine blade early damage prediction model is established through the following process: collecting acoustic emission signals through fatigue loading tests on blade specimens. The material of the blade specimens is consistent with the wind turbine blade to be identified. Peak frequency and amplitude are used as feature values. The acoustic emission signals collected in the fatigue loading test are classified using an improved K-means algorithm, and a convolutional neural network is constructed to obtain the pre-established wind turbine blade early damage prediction model.

10. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the method for identifying early damage to a wind turbine blade according to any one of claims 1 to 8 is implemented.