Wind turbine abnormality detection method and system based on weighted kernel density estimation
Patent Information
- Application Number
- CN202610815638.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-08
- Publication Date
- 2026-08-28
AI Technical Summary
[0004]有鉴于此,本发明提供一种基于加权核密度估计的风电机组异常检测方法及系统,主要目的在于解决现有风电机组异常检测准确性低的问题
本发明提供了一种基于加权核密度估计的风电机组异常检测方法及系统,本发明实施例通过获取待检测周期内的风电机组监测数据,并对所述风电机组监测数据进行基于风电机组的规格参数和/或工作环境的聚类划分,得到至少一个风电机组集群的数据集合,所述数据集合包括对应风电机组集群中各风电机组的监测数据;分别对各风电机组集群的数据集合中的数据进行加权补偿,并基于加权补偿后的数据集合进行概率密度分布建模,得到不同风电机组集群各自对应的概率密度分布模型;依据所述监测数据在所述概率密度分布模型下的密度值,提取所述风电机组集群的核心运行域;针对各风电机组集群,依据集群内各风电机组的监测数据与所述核心运行域的位置分布关系,确定出各所述风电机组集群中的异常风电机组,大大减少了环境与规格差异导致的误判,降低了模型对数据分布假设的依赖,同时,又确保了核心运行域提取的物理可解释性,从而大大提高集群异常风电机组识别的可信度与检测的准确性。
Smart Images

Figure CN122649972A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wind power detection technology, and in particular to a method and system for detecting anomalies in wind turbine units based on weighted kernel density estimation. Background Technology
[0002] Currently, the wind power industry is developing rapidly, with the cumulative installed capacity and new installed capacity of wind turbines increasing year by year, and the scale of wind farms also expanding accordingly. This increases the difficulty of operation and maintenance monitoring of a large number of turbines in wind farms. Larger wind farms have hundreds of turbines. How to quickly identify abnormal turbines from a large number of turbines and improve the efficiency of wind farm operation and maintenance is of significant engineering value for ensuring the safe and stable operation of wind farms, reducing operation and maintenance costs, and rationally scheduling operation and maintenance time.
[0003] Current methods for identifying wind turbine anomalies primarily rely on deviation analysis of the guaranteed power curve. This involves cleaning the raw data acquired by the real-time data acquisition and monitoring system and calculating the deviation between the actual power value and the guaranteed power curve at the corresponding wind speed. If the deviation exceeds a set absolute threshold, such as 5% or 10%, the turbine is deemed to be malfunctioning. However, this method has significant limitations when dealing with large-scale, aging wind farms. For example, for turbines that have been in operation for more than 5-10 years, their intrinsic performance has deviated due to mechanical wear, control system upgrades, or partial technical modifications. The original guaranteed power curve no longer reflects the turbine's health status. Using a failure-based evaluation would lead to an extremely high false alarm rate, failing to meet the accuracy requirements for monitoring wind turbine anomalies. Summary of the Invention
[0004] In view of this, the present invention provides a wind turbine anomaly detection method and system based on weighted kernel density estimation, the main purpose of which is to solve the problem of low accuracy in existing wind turbine anomaly detection.
[0005] According to one aspect of the present invention, a method for anomaly detection of wind turbine generators based on weighted kernel density estimation is provided, comprising: Acquire wind turbine monitoring data within the testing period, and perform clustering based on the wind turbine specifications and / or operating environment to obtain at least one wind turbine cluster data set, the data set including monitoring data of each wind turbine in the corresponding wind turbine cluster; The data in the datasets of each wind turbine cluster are weighted and compensated respectively, and the probability density distribution model is modeled based on the weighted and compensated datasets to obtain the probability density distribution models corresponding to each wind turbine cluster. Based on the density values of the monitoring data under the probability density distribution model, the core operating domain of the wind turbine cluster is extracted; For each wind turbine cluster, based on the monitoring data of each wind turbine within the cluster and the location distribution relationship of the core operating domain, abnormal wind turbines in each wind turbine cluster are identified.
[0006] According to another aspect of the present invention, a wind turbine anomaly detection system based on weighted kernel density estimation is provided, comprising: The data preprocessing layer is used to preprocess the wind turbine monitoring data acquired during the detection period, and to perform clustering based on the wind turbine specifications and / or working environment to obtain a data set of at least one wind turbine cluster, wherein the data set includes the monitoring data of each wind turbine in the corresponding wind turbine cluster. The weighted kernel density estimation modeling layer is used to perform weighted compensation on the data set of each wind turbine cluster, and to perform probability density distribution modeling based on the weighted compensation data set to obtain the probability density distribution model corresponding to each wind turbine cluster. The core domain adaptive extraction layer is used to extract the core operating domain of the wind turbine cluster based on the density value of the monitoring data under the probability density distribution model. Anomaly quantification scoring layer is used to identify abnormal wind turbines in each wind turbine cluster based on the location distribution relationship between the monitoring data of each wind turbine in the cluster and the core operating domain.
[0007] According to another aspect of the present invention, a storage medium is provided, wherein at least one executable instruction is stored therein, the executable instruction causing a processor to perform operations corresponding to the wind turbine anomaly detection method based on weighted kernel density estimation described above.
[0008] According to another aspect of the present invention, a terminal is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction, which causes the processor to perform the operation corresponding to the wind turbine anomaly detection method based on weighted kernel density estimation described above.
[0009] By employing the above technical solutions, the technical solutions provided by the embodiments of the present invention have at least the following advantages: This invention provides a method and system for detecting anomalies in wind turbines based on weighted kernel density estimation. The embodiments of this invention acquire monitoring data of wind turbines within the detection period and cluster the monitoring data based on the specifications and / or operating environment of the wind turbines to obtain a dataset of at least one wind turbine cluster. The dataset includes monitoring data of each wind turbine in the corresponding cluster. Weighted compensation is applied to the data in the datasets of each wind turbine cluster, and probability density distribution modeling is performed based on the weighted compensation datasets to obtain probability density distribution models corresponding to different wind turbine clusters. Based on the density values of the monitoring data under the probability density distribution models, the core operating domain of the wind turbine cluster is extracted. For each wind turbine cluster, based on the positional distribution relationship between the monitoring data of each wind turbine in the cluster and the core operating domain, the abnormal wind turbines in each wind turbine cluster are determined. This significantly reduces misjudgments caused by environmental and specification differences, reduces the model's dependence on data distribution assumptions, and ensures the physical interpretability of the extracted core operating domain, thereby greatly improving the credibility and accuracy of identifying abnormal wind turbines in the cluster.
[0010] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention, it can be implemented according to the contents of the specification. Furthermore, in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0011] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 The flowchart of a wind turbine anomaly detection method based on weighted kernel density estimation provided by an embodiment of the present invention is shown. Figure 2 A flowchart illustrating a method for constructing a probability density distribution model according to an embodiment of the present invention is shown; Figure 3 This invention provides a power curve and a data-intensive region diagram according to an embodiment of the invention. Figure 4 This invention provides a data density heatmap according to an embodiment of the invention. Figure 5 This diagram illustrates the distribution of abnormal and normal generating units according to an embodiment of the present invention. Figure 6The diagram shows a block diagram of a wind turbine anomaly detection system based on weighted kernel density estimation provided by an embodiment of the present invention. Figure 7 A schematic diagram of the structure of a terminal provided in an embodiment of the present invention is shown. Detailed Implementation
[0012] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0013] To address the issue of low accuracy in existing wind turbine anomaly detection methods, this invention provides a wind turbine anomaly detection method based on weighted kernel density estimation. This method does not rely on any prior guaranteed power curves and utilizes the "collective consensus" of wind turbines with similar operating characteristics to construct a dynamic benchmark space. Addressing the underlying physical characteristics of wind turbines operating in high wind conditions (near rated power), where the naturally sparse sample size leads to a high false alarm rate with traditional kernel density estimation, this invention creatively introduces a power frequency inverse weighting correction coefficient. Without altering the physical time series framework, it applies mathematical density weighting compensation to the low-frequency, high-power range, reducing statistical bias caused by uneven operating condition distribution and ensuring the physical interpretability and detection accuracy of the core operating domain extraction across the entire operating condition range. Figure 1 As shown, the method includes steps 101-104: 101. Obtain the monitoring data of wind turbine units within the testing period, and perform clustering on the monitoring data of wind turbine units based on the specifications and / or working environment of the wind turbine units to obtain a data set of at least one wind turbine unit cluster.
[0014] In this embodiment of the invention, the detection period can be one week, half a month, or one month, and the data set includes monitoring data of each wind turbine in the corresponding wind turbine cluster. The monitoring data includes active power, reactive power, rotor speed, generator speed, and high-speed shaft speed. After acquiring the wind turbine monitoring data for the detection period, the data needs to be clustered based on specifications and / or operating environment to form data sets for different wind turbine clusters. Clustering groups turbines with similar operating characteristics and conditions into the same cluster, eliminating the interference of inherent differences between turbines on anomaly detection. This allows subsequent modeling and comparative analysis to be conducted within the same type of turbine, improving the targeting and accuracy of the detection. Specifications include rated power, blade diameter, and hub height; operating environment includes geographical location, climate conditions, and terrain features. For example, turbines with the same rated power within the same wind farm can be grouped into one category, or they can be aggregated across different wind farms, that is, turbines of the same model from multiple wind farms that are geographically close and have similar climate conditions can be grouped into the same cluster; they can also be aggregated across different models, that is, turbines of different manufacturers or different models with the same rated power and similar blade diameter can be grouped into the same cluster.
[0015] In one embodiment of the present invention, for further explanation and limitation, the process of dividing a wind turbine cluster includes: Extract the static specification characteristics, environmental parameter characteristics, and dynamic health index characteristics of each wind turbine unit; The data for each feature dimension are normalized, the information entropy of each feature dimension is calculated, and the weight value of each feature dimension is determined based on the difference coefficient calculated from the information entropy. For each dimension of features after weighting, a weighted K-means clustering algorithm is performed using weighted Euclidean distance as a similarity measure, and the wind turbines in each cluster are treated as a wind turbine cluster.
[0016] In this embodiment of the invention, static specifications include rated power, blade diameter, and hub height; environmental parameters include geographical location (longitude and latitude), annual average wind speed, and air density; and dynamic health indicators include power curve deviation, average vibration amplitude, pitch frequency, and frequency of fault codes. The raw data for each feature dimension are normalized to eliminate the influence of dimensions. The information entropy of each feature dimension is calculated, and then the difference coefficient for each dimension is calculated based on the information entropy. The difference coefficient is calculated as 1 - information entropy. After normalizing the difference coefficients of each dimension, the final weight value for each feature dimension is obtained. Finally, the feature values of each dimension are multiplied by their corresponding weights, and a weighted K-means clustering algorithm is executed using weighted Euclidean distance as a similarity measure: a preset number of clusters is used, cluster centers are randomly initialized, each turbine is assigned to the nearest cluster center, and then the cluster center is updated to the weighted mean of the samples within the cluster. This assignment and update process is repeated until the cluster centers no longer change. Finally, all wind turbines in each cluster are considered as a single wind turbine cluster.
[0017] In one embodiment of the present invention, for further explanation and limitation, before partitioning the data set of the wind turbine cluster, the method further includes: Data cleaning is performed on the monitoring data of the wind turbine based on the rationality of the wind speed domain across the entire range, the physical boundary of the power domain, and the matching of operating parameters.
[0018] In this embodiment of the invention, the data cleaning process based on the reasonableness of the entire wind speed domain includes: removing sensor failure data caused by transient jamming or zero-point drift of the anemometer, i.e., data with wind speed values less than or equal to 0 m / s; simultaneously, removing data with wind speeds exceeding the limit survival wind speed of most wind turbine units, i.e., data with wind speed values greater than 50 m / s. This removes data from anemometer malfunctions or extreme weather conditions that prevent power generation, thus eliminating data with performance benchmarking value. Data cleaning based on the physical boundaries of the power domain is performed because, in actual wind power operation, negative power usually indicates that the unit is in standby yaw power absorption, transient over-cable disconnection, or off-grid power generation state, which belongs to abnormal power generation and consumption conditions. Therefore, data with negative real-time active power is removed. At the same time, the design rated power of each wind turbine unit is statistically analyzed, and data with real-time active power significantly exceeding the upper limit of the reasonable range, i.e., greater than 1.15 times the design rated power, is removed. The matching of operating parameters involves collaboratively verifying the time-series correlation between wind speed, power, turbine speed, and generator speed, eliminating storm shutdown data where wind speeds are extremely high (e.g., exceeding the cut-off wind speed) but power and speed are zero. These operations eliminate interference from single-unit transient faults and improve data validity.
[0019] 102. Perform weighted compensation on the data set of each wind turbine cluster, and perform probability density distribution modeling based on the weighted compensation data set to obtain the probability density distribution model corresponding to each wind turbine cluster.
[0020] In this embodiment of the invention, weighted compensation is applied to the data sets of each wind turbine cluster. This involves assigning different weighting coefficients or performing correction calculations to each monitoring data point based on turbine specifications such as rated power and blade diameter, and operating environment factors such as air density, turbulence intensity, and wind speed distribution. This eliminates data discrepancies between turbines within the cluster caused by factors beyond their control, such as installation location, local wakes, and sensor bias, thus highlighting the true operating characteristics of the turbines. Then, probability density distribution modeling is performed based on the weighted compensation data sets to obtain probability density distribution models corresponding to different wind turbine clusters. Before probability density distribution modeling, the monitoring data in the data sets needs to be normalized, and then probability density distribution modeling is performed based on the normalized monitoring data. The probability density distribution model can characterize the data distribution characteristics and fluctuation range of the turbines within the cluster under normal operating conditions, serving as a benchmark reference for subsequent anomaly detection.
[0021] In one embodiment of the present invention, for further illustration and limitation, such as Figure 2 As shown, the modeling process for the probability density distribution model of any wind turbine cluster includes steps 201-204: The weighted compensation process for the dataset of any wind turbine cluster includes: 201. Based on the frequency of all wind turbine monitoring data appearing in different power ranges, determine the weight of the corresponding power range to obtain the range compensation weight corresponding to each power range.
[0022] 202. The monitoring data within the power range is weighted and compensated according to the interval compensation weight to obtain the compensated monitoring data.
[0023] In this embodiment of the invention, the power range includes either a wind speed range or a power range. That is, the power range can be divided based on rated power or wind speed, with different power ranges corresponding to different power ranges or different wind speed ranges. Based on the frequency of monitoring data from all wind turbines within the cluster appearing in different power ranges, weights are determined for the corresponding power ranges. Ranges with higher frequencies are assigned lower weights, and ranges with lower frequencies are assigned higher weights to balance the uneven distribution of data, thus obtaining the range compensation weights corresponding to each power range. Then, the monitoring data within the corresponding power range is weighted and compensated according to the range compensation weights, so that the originally sparse low-frequency ranges gain greater influence in the modeling, thereby avoiding the model from being overly biased towards high-density conventional operating conditions.
[0024] In the process of weighted compensation of monitoring data within a corresponding power range based on the interval compensation weight, it is not necessary to compensate all monitoring data, but only a portion of the data strongly correlated with wind turbine faults. The specific extraction process of strongly correlated data includes: extracting monitoring data from a period of time before each fault occurrence from historical fault records, such as 10-30 minutes before the fault, and using this data as fault-correlated data. Then, statistically analyzing which power range contains the most fault-correlated data. Selecting the top few power ranges with the highest proportion of fault-correlated data, preferably the top 30%, as high-correlation fault ranges. Within each high-correlation fault range, comparing the fault precursor data with normal data, identifying the data with the most significant changes, such as significant power fluctuations or abnormal speed fluctuations, and selecting the top few variables with the largest changes, such as five, as the strong fault-correlation data for that range. Through the above screening, only monitoring data that simultaneously meets the criteria of being located in a high-correlation fault range and belonging to strongly correlated fault variables can proceed to the subsequent weighted compensation steps. The remaining data is not compensated or participated in subsequent density probability distribution modeling, thus enabling the accurate location of data subsets strongly correlated with faults from massive monitoring data, avoiding computational redundancy and noise interference caused by indiscriminate compensation. Correspondingly, when dividing the core operating area and identifying anomalies, density values can be calculated only for the aforementioned fault-correlated data points.
[0025] It should be noted that the power range refers to the statistical intervals divided along the wind speed axis or power axis in the two-dimensional data space of wind speed and power. Inverse weighting compensation is performed based on the frequency of monitoring data from all wind turbines within the cluster falling into different intervals. The core physical and mathematical mechanism is that the meteorological wind speed of a real wind farm naturally follows a Weibull or Rayleigh distribution, resulting in extreme non-uniformity of the monitoring data of wind turbines across the entire operating range. Data accounting for more than 70% of the total operating time is highly concentrated in the conventional power generation area with low and medium wind speeds, constituting the high-frequency conventional interval. For the high-wind operating condition range, such as the rated power high-wind speed area from 10 m / s to a cutoff wind speed of 25 m / s, the probability of high winds is inherently low, resulting in naturally sparse monitoring data frequency within this interval, belonging to the low-frequency operating condition interval. Although the data frequency in the high-wind operating condition area is low, this area is the core area where the core benefits of wind turbines are the highest, the mechanical load is the greatest, and the pitch system and control strategy are most prone to potential performance degradation such as pitch angle offset and power limiting failure. This area has irreplaceable asset monitoring value in the batch detection of performance anomalies and should not be discarded.
[0026] If traditional kernel density estimation is used for direct modeling, the absolute probability density value of even completely normal power generation points in the low-frequency high-wind area will be severely diluted due to the extremely high data density in the normal area. When extracting the core operating domain, these normal low-frequency high-wind data points will be incorrectly judged as performance anomalies because their density values are below the global threshold, causing serious system-level false alarms. Therefore, by calculating the frequency of data points within each power / wind speed interval and assigning an interval inverse weighting compensation coefficient to the data within the corresponding interval, the total weighted integral of the data in each interval is mathematically made to approach equal weighting. The essence of this amplification mechanism is not to artificially amplify any specific abnormal operating condition, but rather to use mathematical means to remove the statistical bias caused by the uneven distribution of natural wind energy resources, leveling the benchmark comparability across all operating conditions. This ensures that when the kernel density probability distribution model delineates the core operating domain, its density distribution can truly and objectively reflect the deviation between the wind turbine's own power generation efficiency and the collective consensus, rather than the probability of wind speed occurrence, thus eliminating systematic false alarms of normal power generation data in the high-wind-speed rated power range.
[0027] In one embodiment of the present invention, for further explanation and limitation, based on the frequency of monitoring data of all wind turbine units appearing in different power ranges, a weight is determined for the corresponding power range to obtain the range compensation weight corresponding to each power range, including: After the active power of each wind turbine in the generator cluster is adaptively normalized across the entire range, it is divided into multiple power sub-intervals with equal width according to a preset interval in the per-unit value space. For each power range, the data compensation weight for that power range is calculated based on the reciprocal of the frequency of the monitoring data within the cluster appearing in that power range.
[0028] In this embodiment of the invention, the data compensation weight is used to compensate for the asymmetry in the probability distribution of monitoring data within different power ranges. A preset range width is selected based on the rated power parameter; for example, for a 3MW unit, each range can be 50kW, i.e., the preset range width is 50kW. The full-range power of the unit is divided into multiple equally wide power ranges based on the preset range width. Then, the total number of monitoring data points whose power falls within each power range is counted to obtain the frequency. Finally, the compensation weight for different power ranges is calculated according to the weighting formula. The weighting formula is expressed as: ; in, This represents the compensation weight for the l-th power interval. This represents the total number of data points in the monitoring data whose power falls within the l-th power interval. This represents the total number of monitoring data points. Through the above weighted compensation, the contribution of each power interval to the final density estimate tends to be equal, that is, a point in the high-power segment will have a higher component in the probability distribution model than a point in the low-power segment, thus compensating for the asymmetry of the probability distribution.
[0029] Specifically, the method of dividing the power sub-intervals into equal-width sub-intervals within the per-unit value space includes: extracting the real-time maximum active power of all wind turbines in the current cluster from the dataset, dividing the absolute active power of each wind turbine by the real-time maximum active power to obtain a dynamic power per-unit value mapped to the closed interval [0,1], thereby eliminating dimensional interference when aggregating units with different rated capacities; then, setting the number of segments for the power sub-intervals, which is a discrete integer between [15-30]. Finally, within the per-unit value space of [0,1], the entire range is divided into multiple continuous power sub-intervals with equal width. The number of segments is set to 20. If this number is too small, it will cause confusion between low-frequency high-wind conditions and medium-to-high-frequency conventional conditions within the same interval, making it impossible to achieve refined inverse weight compensation; if this number is too large, it will cause the sample frequency of the extreme high-wind interval at the tail end to degenerate to single digits or even zero, leading to mathematical divergence of the weighting coefficients and excessive noise amplification.
[0030] It should be noted that, in addition to power, wind speed can also be divided into intervals in the same way as described above. The wind speed interval with the longest wind turbine operation time can be downweighted, while the intervals near extreme wind speed cut-in / cut-out can be upweighted. This can also achieve the goal of balancing the data distribution.
[0031] In one embodiment of the present invention, for further explanation and limitation, the process of modeling the probability density distribution based on the weighted and compensated dataset for any wind turbine cluster includes: 203. Based on the dispersion of the compensated monitoring data, calculate the adaptive bandwidth corresponding to wind speed and power respectively based on Gaussian approximation.
[0032] 204. Map the compensated monitoring data onto the wind speed-power two-dimensional plane, calculate the density estimate of any grid node in the wind speed-power two-dimensional plane based on the adaptive bandwidth and kernel function, and obtain the probability density distribution model of the wind turbine cluster.
[0033] In this embodiment of the invention, the adaptive bandwidth is first calculated based on the dispersion of the compensated monitoring data, using the Gaussian approximation, i.e., the Silverman rule. Specifically, the standard deviation and interquartile range of wind speed and power in the compensated dataset are calculated first, and the smaller of the two is taken as a robust estimate of the data dispersion. Then, the Silverman empirical formula is substituted to calculate the globally optimal bandwidth. The bandwidth is proportional to the dispersion; the larger the dispersion, the larger the bandwidth to enhance smoothness, and the smaller the dispersion, the narrower the bandwidth to preserve details. This avoids oversmoothing or underfitting while ensuring that the constructed probability density distribution model more accurately reflects the true distribution characteristics of the weighted compensated monitoring data. The compensated monitoring data is then mapped onto a two-dimensional wind speed-power plane, and the density estimate of any grid point x(v,p) in the wind speed-power two-dimensional plane is calculated based on the density estimation formula. The formula for the density estimate is expressed as: ; Where n is the total number of monitoring data points in the dataset, This represents the weight of the i-th monitoring data point. The adaptive bandwidth representing wind speed. The adaptive bandwidth represents the power, which typically fluctuates between 0.1 and 0.3; K represents the kernel function; x(v,p) represents the power at monitoring data point x, where v represents wind speed and p represents power at monitoring data point x. This represents the wind speed at the i-th monitoring data point. Let represent the power of the i-th monitoring data point. The wind speed-power plane is divided into fine grids, and the probability density value of each grid intersection is calculated using the above formula to construct a continuous density surface, thus obtaining the probability density distribution model.
[0034] It should be noted that the kernel function is preferably a Gaussian kernel function, but it can also be replaced with an Epanechnikov kernel function, an exponential kernel function, or a trigonometric kernel function. These substitutions of mathematical functions do not change the core logic of establishing a dynamic benchmark through probability density. Of course, the distribution of the power curve can also be fitted based on a Gaussian mixture model or the expectation-maximization algorithm; this embodiment of the invention does not impose specific limitations.
[0035] In one embodiment of the present invention, for further explanation and definition, the grid division process in the wind speed-power two-dimensional plane includes: Extract the extreme values of wind speed and power from the compensated monitoring data respectively; Based on the extreme wind speed and the extreme power, multiple discrete sampling points are selected at equal intervals on the wind speed axis and on the power axis using a linear spatial interpolation operator. The discrete sampling points of the wind speed axis and the power axis are cross-interlocked and projected to generate a two-dimensional grid array containing multiple topological nodes, wherein the number of discrete sampling points on the wind speed axis and the power axis is the same.
[0036] In this embodiment of the invention, after weighted compensation of the monitoring data, the compensated dataset is scanned to find the minimum and maximum values on the wind speed axis and the power axis, respectively. These two extreme value ranges jointly determine the boundary of the subsequent grid construction, ensuring that all compensated data points fall within the grid coverage area. Then, based on the extracted wind speed and power extreme value ranges, multiple discrete sampling points are selected at equal intervals on the wind speed and power axes using a linear spatial interpolation operator. Equal interval selection means taking points at equal intervals within the extreme value range, with the same number of sampling points on each axis. This number ranges from 50 to 200, preferably set to 100. The number of sampling points determines the fineness of the grid: the more points, the higher the grid resolution. The discrete sampling points selected on the wind speed axis and the discrete sampling points selected on the power axis are cross-interlocked projected, that is, each wind speed sampling point is paired with each power sampling point to form a two-dimensional topological node matrix. This matrix is a two-dimensional grid array, and the total number of nodes it contains is equal to the number of wind speed sampling points multiplied by the number of power sampling points. When the number of sampling points for both axes is preferably set to one hundred, the grid array contains a total of 10,000 topology nodes.
[0037] By discretizing the continuous two-dimensional wind speed-power space into a grid matrix composed of finite nodes, the subsequent kernel density estimation is transformed from high-dimensional continuous computation to discrete matrix operations. On the one hand, the high-density discrete grid ensures the fitting accuracy and edge smoothness of the core operating domain boundary, avoiding jagged distortion. On the other hand, the nearest neighbor node indexing operator enables millisecond-level fast matching in subsequent dynamic verification, thereby reducing the memory overhead and computation time of full-field batch anomaly detection.
[0038] 103. Based on the density values of the monitoring data under the probability density distribution model, extract the core operating domain of the wind turbine cluster.
[0039] In this embodiment of the invention, the core operating domain corresponds to the most concentrated and typical operating state space of the unit under normal operating conditions. By extracting the core operating domain, the high-dimensional operating parameter space of each cluster can be divided into a core area and an edge area, providing clear probability boundaries for subsequent anomaly detection.
[0040] In one embodiment of the present invention, for further explanation and limitation, the process of determining the core operating domain of any of the wind turbine clusters includes: Calculate the density value of each monitoring data point in the cluster under the probability density distribution model; All density values are sorted in descending order to obtain a density sequence. The density quantile positions are determined from the density sequence according to preset density quantile parameters, and the density values at the density quantile positions are used as critical density values. The connected domain formed by all monitoring data points in the wind speed-power two-dimensional plane whose density value is greater than the critical density value is determined as the core operating domain.
[0041] In this embodiment of the invention, each monitoring data point within the generator cluster is substituted into a probability density distribution model to calculate the density values of different monitoring data points. All density values are then sorted in descending order to obtain a density sequence. A density quantile parameter, such as 87% or 90%, is introduced. The position of the density quantile parameter in the density sequence is determined, for example, the 87th position from the beginning, i.e., the density quantile position. The density value at this position is then used as the critical density value to delineate the core operating region in the wind speed-power two-dimensional plane. This region automatically encompasses the operating zones of the vast majority of generators. If the entire site is under power curtailment, the core operating region will automatically adopt a curtailment shape; if the entire site is operating well, the core operating region will closely resemble the standard curve shape.
[0042] It should be noted that the core operating domain can also be defined using density contour lines or support vector machine regression. Density contour lines: First, high-density points are selected, then the operating boundary is delineated using Alpha-shape (concave hull algorithm) or Convex Hull (convex hull algorithm). Support vector machine regression: One-class SVM is used to learn the boundaries of the weighted population data, identifying the smallest hyperplane region containing 90% of the monitored data points as the core operating domain.
[0043] In this embodiment of the invention, the density quantile parameter introduced during the core operating domain division process and the individual abnormality point ratio threshold used in subsequent single-unit abnormality judgment, i.e., the individual abnormality coefficient threshold, form a dual threshold decoupling buffer mechanism to mitigate the interference of normal low-frequency / discrete operating conditions on the detection results. The reason for this is that in actual wind turbine operation, there are numerous normal low-frequency transient operating conditions. However, conditions such as timing-sequence grid connection near the cut-in wind speed, yaw wind transient adjustment periods, constant pitch testing, and safety chain transient triggering do not consider falling outside the domain as a necessary and sufficient condition for determining an anomaly. Because the cumulative proportion of these normal low-frequency transient operating conditions over the long-term operating cycle of the wind turbine has a very strong physical boundary limitation, statistics show that the proportion of the aforementioned transient points under normal operating conditions remains stable between 2% and 6% of the total sampling points. By setting the threshold for individual anomaly coefficients to a spatial span significantly larger than the aforementioned transient proportion, preferably with a preset threshold of 14%, these normal low-frequency discrete points generated by natural physical characteristics can be completely absorbed and physically eliminated by a 14% global tolerance buffer. This decoupled design with dual thresholds ensures that the core operating domain boundary can compactly and accurately encompass the core benchmark band of high power generation efficiency for the group, while also giving the model tolerance for the inherent normal discrete operating conditions of wind turbines at the mechanism level. This fundamentally avoids the technical defect of traditional probabilistic modeling that misjudges low-frequency normal points as anomalies.
[0044] 104. For each wind turbine cluster, based on the monitoring data of each wind turbine in the cluster and the location distribution relationship of the core operating domain, the abnormal wind turbines in each wind turbine cluster are identified.
[0045] In this embodiment of the invention, for each wind turbine cluster, abnormal wind turbines are identified based on the location distribution relationship between the monitoring data of each wind turbine within the cluster and the core operating domain of the cluster. Specifically, the monitoring data points of each turbine are projected onto a pre-constructed probability density distribution model, and the distance between the data points and the boundary of the core operating domain or the proportion falling outside the domain is calculated. If the monitoring data of a turbine continuously or significantly deviates from the core operating domain, the turbine is determined to be an abnormal turbine. This method enables comparable anomaly detection within the same cluster, effectively identifying turbines that deviate from normal behavior patterns due to blade wear, gearbox degradation, propeller angle deviation, sensor drift, etc., while avoiding misjudgments caused by environmental differences or different turbine specifications, thus improving the accuracy of anomaly detection.
[0046] In one embodiment of the present invention, for further explanation and limitation, the method for determining abnormal wind turbines for any wind turbine cluster includes: For each wind turbine in the wind turbine cluster, the monitoring data of the wind turbine is mapped to the core operating domain, abnormal monitoring data points outside the core operating domain are counted, and the ratio of the abnormal monitoring data points to the total number of monitoring data points contained in the monitoring data is used as the individual abnormality coefficient of the wind turbine. For any wind turbine, if the individual anomaly coefficient is less than or equal to a preset individual anomaly coefficient threshold, the wind turbine is determined to be a normal turbine. If the individual anomaly coefficient is greater than the preset individual anomaly coefficient threshold, the wind turbine is considered a candidate abnormal turbine, and the proportion of all candidate abnormal turbines in the wind turbine cluster is calculated. If the proportion is greater than the first preset anomaly proportion threshold, the wind turbine is determined to be a normal turbine. If the proportion is less than the second preset anomaly proportion threshold, the wind turbine is determined to be an abnormal turbine. If the proportion is between the first preset anomaly proportion threshold and the second preset anomaly proportion threshold, anomaly identification is performed based on the distance from the monitoring data point to the boundary of the core operating domain.
[0047] In this embodiment of the invention, mapping and threshold comparison are performed for each wind turbine in the wind turbine cluster, that is, for the first... The number of points belonging to the Taiwanese server group that fall outside the core operating domain is counted. This accounts for the total number of points of the unit. The proportion of this is denoted as the individual abnormality coefficient: Set individual abnormality coefficient thresholds. .when When this happens, the unit is marked as an abnormal unit.
[0048] After obtaining the individual anomaly coefficient, preliminary screening of candidate anomalous wind turbines is performed by comparing the individual anomaly coefficient with a preset individual anomaly coefficient threshold. When the individual anomaly coefficient of a wind turbine exceeds the preset threshold, it is not directly determined to be anomalous, but rather marked as a candidate anomalous turbine, and the proportion of this candidate anomalous turbine in the entire cluster is calculated. If this proportion is greater than the first preset anomaly proportion threshold, meaning that the vast majority of turbines in the cluster are simultaneously exhibiting deviation behavior, then the wind farm is determined to be experiencing a site-wide external disturbance. In this case, the individual turbine is not considered anomalous; conversely, if the proportion is lower than the corresponding second preset anomaly proportion threshold, meaning that only a few turbines are deviating, indicating that the deviation behavior is an individual-specific manifestation, then the wind turbine is ultimately confirmed as a true anomalous turbine, and an operation and maintenance dispatch order is issued. The first preset anomaly proportion threshold can be 0.75, and the second preset anomaly proportion threshold can be 0.25.
[0049] If the proportion falls between the first and second preset abnormal proportion thresholds, the distance from the monitoring data points to the boundary of the core operating domain can be further used to identify candidate abnormal units. Specifically, the geometric distance from each monitoring data point of the candidate abnormal unit to the boundary of the core operating domain is calculated, and a cubic wind speed inverse weighted correction is applied to obtain the physical weighted deviation. This deviation is compared with a preset deviation threshold; if it exceeds the threshold, the unit is determined to be a true abnormal unit; otherwise, it is determined to be a normal unit. (Data monitoring points outside the core operating domain) Physical weighted deviation distance to the domain boundary Defined as: ; in, For point The geometric distance to the boundary of the core operating domain. The real-time wind speed corresponding to this point. This refers to the rated wind speed of this model. By using inverse weighting of the cube of the wind speed, slight physical deviations in the low wind speed region are given a significant distance gain weight, while the saturation control band in the high wind speed region is given a reasonable physical smoothness. This greatly improves the microscopic detection sensitivity of wind turbine pitch angle drift and early blade degradation from a physical mechanism perspective.
[0050] By introducing a companion pool cluster based on the comparison of individual anomaly coefficients, dual verification of individual and companion pools is achieved, which can effectively identify group anomalies caused by concentrated power rationing due to the execution of grid commands by the entire field or by collective micro-icing of blades caused by the occurrence of extreme low temperatures across the entire field, thereby improving the accuracy of anomaly identification.
[0051] It should be noted that distance-weighted deviation can also be used, which calculates the sum of the geometric distances from every point outside the core operating domain to the boundary of the core domain. This means that the farther a point is from the core domain, the greater its contribution to the anomaly score, thus more sensitively detecting severe deviations. Alternatively, distributional difference measures can be used, such as Kullback-Leibler Divergence or Wasserstein distance, to directly measure the overall distance between the probability distribution of an individual wind turbine and the group-weighted distribution, thus serving as an indicator for anomaly detection.
[0052] In one embodiment of the present invention, for further explanation and limitation, after identifying the abnormal wind turbines in each wind turbine cluster, the monitoring data of any wind turbine cluster is visualized, specifically including: The population probability density distribution calculated based on the probability density distribution model is rendered in the form of a contour map onto the wind speed-power two-dimensional plane coordinate system; The monitoring data points of each wind turbine in the wind turbine cluster are distinguished by their respective unit identifiers and displayed as a scatter plot overlaid on the wind speed-power two-dimensional plane coordinate system to obtain a visual display of the wind turbine cluster.
[0053] In this embodiment of the invention, the population probability density distribution calculated based on the probability density distribution model is rendered as a contour map onto a two-dimensional wind speed-power plane coordinate system. Contour regions of different densities correspond to different density levels of the core operating domain. The population probability density distribution is a contour projection of a continuous probability density surface calculated based on a Gaussian kernel density estimation model onto a plane, used to characterize the overall operating behavior of all wind turbine units in the wind speed-power plane. Then, the monitoring data points of each wind turbine in the wind turbine cluster are distinguished by their respective unit identifiers and superimposed on the same coordinate system in the form of a scatter plot to obtain a visualization. The scatter plots represent the measured monitoring data points of a single wind turbine, i.e., wind speed-power pairs, to describe the specific operating status recorded by a single unit at the sampling time.
[0054] During visualization output, a scatter plot with a contour background is automatically generated, where the background is the population density distribution contour lines. In the scatter plot: points matching the population characteristics are displayed in dark blue, while outliers deviating from the population are displayed in a specific highlighted color based on their deviation direction, such as low power or abnormal wind speed measurements. This allows for a direct comparison between population characteristics and individual unit operating status. Figure 3 As shown, this is the power curve and data-dense region map generated when there are 200 monitored data points, an outlier threshold of 0.14, a bandwidth of 0.2, and a density quantile parameter of 87%. The red lines in the figure represent density distribution contours. Further data density heatmaps corresponding to the power curve and data-dense region map can be generated, such as... Figure 4 As shown. Figure 5 As shown, the abnormal and normal units can be clearly presented through the image. Gray data points in the image represent normal units, while colored data points represent abnormal wind turbine units. Of course, colored data points can also be used as candidate abnormal units for further screening.
[0055] In a specific application example, the above method is implemented using the Wind TurbineAnomaly Detector class: pandas is used for multi-threaded data reading and cleaning. scipy.stats.gaussian_kde is used for weighted density core calculations. matplotlib is used to automatically draw diagnostic plots containing kernel density contour lines, normal point bands, and abnormal protrusions, facilitating intuitive verification by operations and maintenance personnel.
[0056] This invention provides a wind turbine anomaly detection method based on weighted kernel density estimation. The method involves acquiring wind turbine monitoring data within a specified detection period and clustering this data based on wind turbine specifications and / or operating environment to obtain at least one wind turbine cluster dataset. Each dataset includes monitoring data from each wind turbine within the cluster. Weighted compensation is applied to the data in each cluster dataset, and probability density distribution modeling is performed based on the weighted compensation dataset to obtain probability density distribution models for each cluster. The core operating domain of each wind turbine cluster is extracted based on the density values of the monitoring data under the probability density distribution model. For each wind turbine cluster, the abnormal wind turbines are identified based on the positional distribution relationship between the monitoring data of each wind turbine within the cluster and the core operating domain. This significantly reduces misjudgments caused by environmental and specification differences, lowers the model's dependence on data distribution assumptions, and ensures the physical interpretability of the extracted core operating domain, thereby greatly improving the reliability and accuracy of identifying abnormal wind turbines within the cluster.
[0057] Furthermore, as a response to the above Figure 1 The implementation of the method shown in this embodiment of the invention provides a wind turbine anomaly detection system based on weighted kernel density estimation, such as... Figure 6 As shown, the system includes: The data preprocessing layer 31 is used to preprocess the wind turbine monitoring data acquired during the detection period, and to perform clustering based on the wind turbine specifications and / or working environment to obtain a data set of at least one wind turbine cluster, wherein the data set includes the monitoring data of each wind turbine in the corresponding wind turbine cluster. The weighted kernel density estimation modeling layer 32 is used to perform weighted compensation on the data in the dataset of each wind turbine cluster, and to perform probability density distribution modeling based on the weighted compensation dataset to obtain the probability density distribution model corresponding to each wind turbine cluster. The core domain adaptive extraction layer 33 is used to extract the core operating domain of the wind turbine cluster based on the density value of the monitoring data under the probability density distribution model. Anomaly quantification scoring layer 34 is used to identify abnormal wind turbines in each wind turbine cluster based on the location distribution relationship between the monitoring data of each wind turbine in the cluster and the core operating domain.
[0058] Furthermore, the weighted kernel density estimation modeling layer 32 includes: The weight determination module is used to determine the weight for the corresponding power range based on the frequency of the monitoring data of all wind turbine units appearing in different power ranges, so as to obtain the range compensation weight corresponding to each power range. The compensation module is used to perform weighted compensation on the monitoring data within the power range according to the range compensation weight, so as to obtain the compensated monitoring data.
[0059] Furthermore, the weight determination module is specifically used to perform full-range adaptive normalization of the active power of each wind turbine in the wind turbine cluster, and then divide it into multiple power sub-intervals with equal width according to the preset interval width in the per-unit value space. For each power range, the data compensation weight of the power range is calculated based on the reciprocal of the frequency of the monitoring data in the power range. The data compensation weight is used to compensate for the asymmetry of the probability distribution of monitoring data in different power ranges.
[0060] Furthermore, the weighted kernel density estimation modeling layer 32 also includes: The bandwidth calculation module is used to calculate the adaptive bandwidth corresponding to wind speed and power based on Gaussian approximation. The construction module is used to map the compensated monitoring data onto a wind speed-power two-dimensional plane, calculate the density estimate of any grid node in the wind speed-power two-dimensional plane based on the adaptive bandwidth and kernel function, and obtain the probability density distribution model of the wind turbine cluster. The kernel function includes a Gaussian kernel function, an Epanechnikov kernel function, an exponential kernel function, or a triangular kernel function. The grid partitioning process in the wind speed-power two-dimensional plane includes: Extract the extreme values of wind speed and power from the compensated monitoring data respectively; Based on the extreme wind speed and the extreme power, multiple discrete sampling points are selected at equal intervals on the wind speed axis and on the power axis using a linear spatial interpolation operator. The discrete sampling points of the wind speed axis and the power axis are cross-interlocked and projected to generate a two-dimensional grid array containing multiple topological nodes, wherein the number of discrete sampling points on the wind speed axis and the power axis is the same.
[0061] Furthermore, the core domain adaptive extraction layer 33 includes: The density calculation module is used to calculate the density value of each monitoring data point in the cluster under the probability density distribution model. The extraction module is used to sort all density values in descending order to obtain a density sequence, determine the density quantile position from the density sequence according to a preset density quantile parameter, and use the density value at the density quantile position as the critical density value. The domain partitioning module is used to determine the connected domain formed by all monitoring data points with density values greater than the critical density value in the wind speed-power two-dimensional plane as the core operating domain.
[0062] Furthermore, the anomaly quantification scoring layer 34 includes: The statistics module is used to map the monitoring data of each wind turbine in the wind turbine cluster to the core operating domain, count the abnormal monitoring data points outside the core operating domain, and use the ratio of the abnormal monitoring data points to the total number of monitoring data points contained in the monitoring data as the individual abnormality coefficient of the wind turbine. The first anomaly determination module is used to determine the wind turbine as a normal unit when the individual anomaly coefficient is less than or equal to a preset individual anomaly coefficient threshold. The second anomaly determination module is used to identify the wind turbine as a candidate anomalous unit when the individual anomaly coefficient is greater than the preset individual anomaly coefficient threshold, and to calculate the proportion of all candidate anomalous units in the wind turbine cluster. If the proportion is greater than the first preset anomaly proportion threshold, the wind turbine is identified as a normal unit. If the proportion is less than the second preset anomaly proportion threshold, the wind turbine is identified as an anomalous unit. If the proportion is between the first preset anomaly proportion threshold and the second preset anomaly proportion threshold, anomaly identification is performed based on the distance from the monitoring data point to the boundary of the core operating domain.
[0063] Furthermore, the system also includes a visualization layer, which is used to calculate the population probability density distribution based on the probability density distribution model and render it in the form of a contour map onto the wind speed-power two-dimensional plane coordinate system, wherein contour regions of different densities correspond to different density levels of the core operating domain. The monitoring data points of each wind turbine in the wind turbine cluster are distinguished by their respective unit identifiers and displayed as a scatter plot overlaid on the wind speed-power two-dimensional plane coordinate system to obtain a visual display of the wind turbine cluster.
[0064] Furthermore, the data preprocessing layer 31 is also used to extract the static specification features, environmental parameter features, and dynamic health indicator features of each wind turbine. The data for each feature dimension are normalized, the information entropy of each feature dimension is calculated, and the weight value of each feature dimension is determined based on the difference coefficient calculated from the information entropy. For each dimension of features after weighting, a weighted K-means clustering algorithm is performed using weighted Euclidean distance as a similarity measure, and the wind turbines in each cluster are treated as a wind turbine cluster.
[0065] This invention provides a wind turbine anomaly detection system based on weighted kernel density estimation. In this embodiment, monitoring data of wind turbines within a detection period is acquired, and the monitoring data is clustered based on the specifications and / or operating environment of the wind turbines to obtain a data set of at least one wind turbine cluster. The data set includes monitoring data of each wind turbine within the corresponding cluster. Weighted compensation is applied to the data set of each wind turbine cluster, and probability density distribution modeling is performed based on the weighted compensation data set to obtain probability density distribution models corresponding to different wind turbine clusters. Based on the density values of the monitoring data under the probability density distribution model, the core operating domain of the wind turbine cluster is extracted. For each wind turbine cluster, based on the positional distribution relationship between the monitoring data of each wind turbine within the cluster and the core operating domain, abnormal wind turbines in each wind turbine cluster are identified. This significantly reduces misjudgments caused by environmental and specification differences, lowers the model's dependence on data distribution assumptions, and ensures the physical interpretability of the extracted core operating domain, thereby greatly improving the credibility and accuracy of identifying abnormal wind turbines in the cluster.
[0066] According to one embodiment of the present invention, a storage medium is provided, the storage medium storing at least one executable instruction, which is capable of executing the wind turbine anomaly detection method based on weighted kernel density estimation in any of the above method embodiments.
[0067] Figure 7 The diagram shows a structural schematic of a terminal according to an embodiment of the present invention. The specific implementation of the present invention is not limited to the specific implementation of the terminal.
[0068] like Figure 7 As shown, the terminal may include: a processor 402, a communication interface 404, a memory 406, and a communication bus 408.
[0069] The processor 402, communication interface 404, and memory 406 communicate with each other via communication bus 408.
[0070] Communication interface 404 is used for network communication with other devices such as clients or other servers.
[0071] The processor 402 is used to execute program 410, which can specifically execute the relevant steps in the above embodiment of the wind turbine anomaly detection method based on weighted kernel density estimation.
[0072] Specifically, program 410 may include program code that includes computer operation instructions.
[0073] Processor 402 may be a central processing unit (CPU), a specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. The terminal may include one or more processors of the same type, such as one or more CPUs; or it may include processors of different types, such as one or more CPUs and one or more ASICs.
[0074] Memory 406 is used to store program 410. Memory 406 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0075] Specifically, program 410 can be used to cause processor 402 to perform the following operations: Acquire wind turbine monitoring data within the testing period, and perform clustering based on the wind turbine specifications and / or operating environment to obtain at least one wind turbine cluster data set, the data set including monitoring data of each wind turbine in the corresponding wind turbine cluster; The data in the datasets of each wind turbine cluster are weighted and compensated respectively, and the probability density distribution model is modeled based on the weighted and compensated datasets to obtain the probability density distribution models corresponding to each wind turbine cluster. Based on the density values of the monitoring data under the probability density distribution model, the core operating domain of the wind turbine cluster is extracted; For each wind turbine cluster, based on the monitoring data of each wind turbine within the cluster and the location distribution relationship of the core operating domain, abnormal wind turbines in each wind turbine cluster are identified.
[0076] It is obvious to those skilled in the art that the modules or steps of the present invention described above can be implemented using general-purpose computing systems. They can be centralized on a single computing system or distributed across a network of multiple computing systems. Optionally, they can be implemented using program code executable by a computing system, thereby storing them in a storage system for execution by the computing system. In some cases, the steps shown or described can be performed in a different order than those presented herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0077] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for detecting anomalies in wind turbine generators based on weighted kernel density estimation, characterized in that, include: Acquire wind turbine monitoring data within the testing period, and perform clustering based on the wind turbine specifications and / or operating environment to obtain at least one wind turbine cluster data set, the data set including monitoring data of each wind turbine in the corresponding wind turbine cluster; The data in the datasets of each wind turbine cluster are weighted and compensated respectively, and the probability density distribution model is modeled based on the weighted and compensated datasets to obtain the probability density distribution models corresponding to each wind turbine cluster. Based on the density values of the monitoring data under the probability density distribution model, the core operating domain of the wind turbine cluster is extracted; For each wind turbine cluster, based on the monitoring data of each wind turbine within the cluster and the location distribution relationship of the core operating domain, abnormal wind turbines in each wind turbine cluster are identified.
2. The method according to claim 1, characterized in that, The weighted compensation process for the dataset of any wind turbine cluster includes: Based on the frequency of all wind turbine monitoring data appearing in different power ranges, a weight is determined for the corresponding power range, and the range compensation weight corresponding to each power range is obtained. The monitoring data within the power range is weighted and compensated according to the interval compensation weight to obtain the compensated monitoring data.
3. The method according to claim 2, characterized in that, Based on the frequency of all wind turbine monitoring data appearing in different power ranges, weights are determined for the corresponding power ranges to obtain the range compensation weights for each power range, including: After the active power of each wind turbine in the generator cluster is adaptively normalized across the entire range, it is divided into multiple power sub-intervals with equal width according to a preset interval in the per-unit value space. For each power range, the data compensation weight of the power range is calculated based on the reciprocal of the frequency of the monitoring data in the power range. The data compensation weight is used to compensate for the asymmetry of the probability distribution of monitoring data in different power ranges.
4. The method according to claim 3, characterized in that, For any wind turbine cluster, the process of modeling the probability density distribution based on the weighted and compensated dataset includes: The adaptive bandwidth corresponding to wind speed and power is calculated based on the Gaussian approximation. The compensated monitoring data is mapped onto a two-dimensional wind speed-power plane. The density estimate of any grid node in the two-dimensional wind speed-power plane is calculated based on the adaptive bandwidth and kernel function to obtain the probability density distribution model of the wind turbine cluster. The kernel function is any one of the Gaussian kernel function, Epanechnikov kernel function, exponential kernel function, and trigonometric kernel function. The process of dividing the grid in the wind speed-power two-dimensional plane includes: Extract the extreme values of wind speed and power from the compensated monitoring data respectively; Based on the extreme wind speed and the extreme power, multiple discrete sampling points are selected at equal intervals on the wind speed axis and on the power axis using a linear spatial interpolation operator. The discrete sampling points of the wind speed axis and the power axis are cross-interlocked and projected to generate a two-dimensional grid array containing multiple topological nodes, wherein the number of discrete sampling points on the wind speed axis and the power axis is the same.
5. The method according to claim 1, characterized in that, The process of determining the core operating domain of any of the aforementioned wind turbine clusters includes: Calculate the density value of each monitoring data point in the cluster under the probability density distribution model; All density values are sorted in descending order to obtain a density sequence. The density quantile positions are determined from the density sequence according to preset density quantile parameters, and the density values at the density quantile positions are used as critical density values. The connected domain formed by all monitoring data points in the wind speed-power two-dimensional plane whose density value is greater than the critical density value is determined as the core operating domain.
6. The method according to claim 5, characterized in that, For any wind turbine cluster, the methods for identifying abnormal wind turbines include: For each wind turbine in the wind turbine cluster, the monitoring data of the wind turbine is mapped to the core operating domain, abnormal monitoring data points outside the core operating domain are counted, and the ratio of the abnormal monitoring data points to the total number of monitoring data points contained in the monitoring data is used as the individual abnormality coefficient of the wind turbine. For any wind turbine, if the individual anomaly coefficient is less than or equal to a preset individual anomaly coefficient threshold, the wind turbine is determined to be a normal turbine. If the individual anomaly coefficient is greater than the preset individual anomaly coefficient threshold, the wind turbine is considered a candidate abnormal turbine, and the proportion of all candidate abnormal turbines in the wind turbine cluster is calculated. If the proportion is greater than the first preset anomaly proportion threshold, the wind turbine is determined to be a normal turbine. If the proportion is less than the second preset anomaly proportion threshold, the wind turbine is determined to be an abnormal turbine. If the proportion is between the first preset anomaly proportion threshold and the second preset anomaly proportion threshold, anomaly identification is performed based on the distance from the monitoring data point to the boundary of the core operating domain.
7. The method according to claim 1, characterized in that, The process of dividing the wind turbine cluster includes: Extract the static specification characteristics, environmental parameter characteristics, and dynamic health index characteristics of each wind turbine unit; The data for each feature dimension are normalized, the information entropy of each feature dimension is calculated, and the weight value of each feature dimension is determined based on the difference coefficient calculated from the information entropy. For each dimension of features after weighting, a weighted K-means clustering algorithm is performed using weighted Euclidean distance as a similarity measure, and the wind turbines in each cluster are treated as a wind turbine cluster.
8. A wind turbine anomaly detection system based on weighted kernel density estimation, characterized in that, include: The data preprocessing layer is used to preprocess the wind turbine monitoring data acquired during the detection period, and to perform clustering based on the wind turbine specifications and / or working environment to obtain a data set of at least one wind turbine cluster, wherein the data set includes the monitoring data of each wind turbine in the corresponding wind turbine cluster. The weighted kernel density estimation modeling layer is used to perform weighted compensation on the data set of each wind turbine cluster, and to perform probability density distribution modeling based on the weighted compensation data set to obtain the probability density distribution model corresponding to each wind turbine cluster. The core domain adaptive extraction layer is used to extract the core operating domain of the wind turbine cluster based on the density value of the monitoring data under the probability density distribution model. Anomaly quantification scoring layer is used to identify abnormal wind turbines in each wind turbine cluster based on the location distribution relationship between the monitoring data of each wind turbine in the cluster and the core operating domain.
9. A storage medium, characterized in that, The storage medium stores at least one executable instruction that causes the processor to perform the operation corresponding to the wind turbine anomaly detection method based on weighted kernel density estimation as described in any one of claims 1-7.
10. A terminal, characterized in that, include: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction that causes the processor to perform the operation corresponding to the wind turbine anomaly detection method based on weighted kernel density estimation as described in any one of claims 1-7.