Wind power plant fault detection method based on historical power generation data

By generating wake field data to identify highly disturbed turbine units and upstream dominant turbine units, and combining power residuals with state correlation features, clustering and isolated forest algorithms are used for wind farm fault detection. This solves the problem of high false alarm rate under the influence of wake interference and achieves efficient fault detection and operation and maintenance guidance.

CN121901907APending Publication Date: 2026-04-21DATANG TONGXIN NEW ENERGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DATANG TONGXIN NEW ENERGY CO LTD
Filing Date
2025-12-19
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing fault detection methods do not fully consider the complex effects of wake interference between wind turbines in wind farms, resulting in a high false alarm rate and an inability to adapt to the differences in operating conditions under different wind speeds and directions, making it difficult to meet the needs of efficient operation and maintenance of wind farms.

Method used

By collecting historical and real-time SCADA data streams, wake field data that is spatiotemporally matched with wind turbine operation data is generated to identify highly disturbed units and their upstream dominant units. Fault detection features are established by combining power prediction regression models and clustering algorithms, and anomaly detection is performed using the isolated forest algorithm to distinguish between wake pseudo-anomalies and real unit faults.

Benefits of technology

It effectively distinguishes between wake pseudo-anomalies and real unit faults, improving the accuracy of fault detection and operation and maintenance guidance. It adapts to the differences in operating conditions under different wind speeds and directions, meeting the needs of efficient operation and maintenance of wind farms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121901907A_ABST
    Figure CN121901907A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of wind power plant monitoring, in particular to a wind power plant fault detection method based on historical power generation data. The method comprises the following steps: collecting historical and real-time SCADA data streams of a wind power plant, and preprocessing the data streams to obtain fan operation data; based on the terrain and the fan layout parameters, wake flow field data in space-time matching is generated through computational fluid mechanics simulation; identifying a highly disturbed unit and a corresponding upstream dominant unit; calculating a power residual error in combination with related data and extracting state correlation features to form fault detection features; a health condition mode is established through a clustering algorithm, and health indexes are calculated; and when the health index is lower than a threshold value, performing wake flow data verification, and performing anomaly detection by adopting an isolated forest algorithm. According to the method, wake flow influence is captured, fault detection features are constructed, wake flow pseudo anomalies and unit real faults are effectively distinguished, and the problem that a traditional fault detection method cannot adapt to the complex operation environment of a wind power plant is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wind farm monitoring technology, and more specifically, to a method for detecting wind farm faults based on historical power generation data. Background Technology

[0002] Fault detection in wind farms is a core technology for ensuring the stable operation of wind power systems and improving power generation efficiency. It primarily involves analyzing wind turbine operating data to identify equipment anomalies and potential faults. Existing fault detection methods often rely on monitoring single operating parameters or general algorithm models, failing to fully consider the complex impact of wake interference between wind turbines within the wind farm. This can easily lead to misjudging operational fluctuations caused by wake interference as faults. Furthermore, the lack of refined modeling of historical normal operating conditions makes it difficult to adapt to differences in operating conditions under varying wind speeds and directions, resulting in low fault detection accuracy and high false alarm rates, thus failing to meet the practical needs of efficient wind farm operation and maintenance. Therefore, this invention provides a wind farm fault detection method based on historical power generation data. Summary of the Invention

[0003] The purpose of this invention is to provide a wind farm fault detection method based on historical power generation data, in order to solve the problem that existing fault detection methods mentioned in the background art mostly rely on monitoring of single operating parameters or general algorithm models, and do not fully consider the complex influence of wake interference between wind turbines in the wind farm.

[0004] To achieve the above objectives, the present invention aims to provide a wind farm fault detection method based on historical power generation data, comprising the following steps:

[0005] S1. Collect historical and real-time SCADA data streams from the wind farm, and preprocess the SCADA data streams to obtain wind turbine operation data;

[0006] S2. Based on the topography and wind turbine layout parameters of the wind farm, wake field data that is spatiotemporally matched with the wind turbine operation data is generated through computational fluid dynamics simulation.

[0007] S3. Based on wake field data, identify the highly disturbed units that are severely affected by the wake and their corresponding upstream leading units;

[0008] S4. For highly disturbed units, the power residual of the highly disturbed units is calculated by combining the wind turbine operation data and wake field data through the power prediction regression model, and its state correlation features with the upstream dominant units are extracted to form fault detection features.

[0009] S5. Based on the fault detection features under historical normal data, several healthy operating condition modes are established through clustering algorithms, and the deviation of the real-time fault detection features from the corresponding healthy operating condition modes is calculated to obtain health indicators.

[0010] S6. When the health index is less than the preset health threshold, first verify it by combining real-time wake field data. For the abnormal fault detection features that pass the verification, use the isolated forest algorithm to detect the anomalies.

[0011] As a further improvement to this technical solution, the specific steps involved in preprocessing the SCADA data stream to obtain the wind turbine operating data in step S1 are as follows:

[0012] Based on the physical reasonable range of each wind turbine operating parameter in the SCADA data stream, outliers and invalid data are removed, and short-term missing data is filled by interpolation; among them, wind turbine operating parameters include wind speed, wind direction, power, and rotational speed;

[0013] The timestamps of data with different sampling frequencies in the cleaned SCADA data stream are aligned by means of aggregation.

[0014] The minimum and maximum values ​​of the operating parameters of each wind turbine are calculated based on historical normal data, and the time-aligned data are normalized using the minimum-maximum normalization method.

[0015] The cleaned, aligned, and normalized data are integrated into a structured data table in chronological order, with each row representing a point in time and each column representing a wind turbine operating parameter, thus forming wind turbine operating data.

[0016] As a further improvement to this technical solution, the specific steps involved in S2, which involve generating wake field data that spatiotemporally matches the wind turbine operating data through computational fluid dynamics simulation, are as follows:

[0017] A numerical simulation model was established in CFD software based on the topography and wind turbine layout parameters of the wind farm.

[0018] Wind speed and direction are obtained from wind turbine operation data and used as inlet boundary conditions to input into the numerical simulation model. Simulation conditions corresponding to the data acquisition time points are also set.

[0019] For each combination of wind speed and direction, the flow field distribution of the entire wind farm area is calculated using a numerical simulation model, and the simulation results are output. The simulation results include the wind speed, turbulence intensity, wind pressure at each wind turbine location, and the wake influence area between wind turbines.

[0020] For each wind turbine, within its upstream influence area, calculate the wake velocity ratio and wake turbulence increment at different locations in the simulation results;

[0021] The simulation results, wake velocity ratio, wake turbulence increment, and wind turbine operating data are matched and aligned according to time and wind turbine number to construct corresponding wake features for each wind turbine at each moment, forming structured wake field data. Among them, the wake features include the wind speed, wind direction, wake velocity ratio, and wake turbulence increment at the location of the wind turbine.

[0022] As a further improvement to this technical solution, the specific steps involved in S3 for identifying the highly disturbed turbine units severely affected by wake ripples and their corresponding upstream main turbine units are as follows:

[0023] Based on the wake velocity ratio and wake turbulence increment of each wind turbine in the wake field data, for each pair of wind turbines in the wind farm, it is determined whether the upstream wind turbine is located in the upstream influence area of ​​the downstream wind turbine.

[0024] If yes, the wake velocity ratio and wake turbulence increment are combined according to a preset weight to calculate the comprehensive influence intensity value; if no, the comprehensive influence intensity value is recorded as zero.

[0025] For each downstream wind turbine The total disturbance degree of the wind turbine is obtained by summing the combined impact intensity values ​​received from all upstream wind turbines. ;

[0026] For all wind turbines, the total disturbance level at each time point is set above the preset disturbance threshold. The wind turbines are marked as currently highly disturbed.

[0027] The proportion of time points marked for each wind turbine in historical data is counted. If the proportion exceeds a preset frequency threshold, the wind turbine is finally identified as a highly disturbed unit.

[0028] For each highly disturbed turbine unit, find the one with the largest combined influence intensity value from all its upstream turbines. Each upstream wind turbine is designated as the upstream leading turbine of the highly disturbed unit.

[0029] As a further improvement to this technical solution, the specific steps involved in calculating the power residual of highly disturbed units through the power prediction regression model in step S4 are as follows:

[0030] Based on wind turbine operating data and wake field data, a power prediction regression model is constructed for each highly disturbed unit using a gradient boosting regression tree, with the current wind turbine state vector as the input. The output is the predicted power value of the wind turbine. The wind turbine state vector includes the current wind speed, wind direction, wake velocity ratio, wake turbulence increment, and turbulence intensity.

[0031] The power prediction regression model is trained using historical normal data, with the actual power generation as the target value, the mean squared error as the loss function, and a regularization term is added to prevent overfitting.

[0032] For highly disturbed turbine units in real-time data, their current turbine state vector is used. Input the trained power prediction regression model to obtain the power prediction value. By comparing the actual power output of the wind turbine with the actual power output of the wind turbine, the difference between the two is calculated to obtain the power residual. ;

[0033] The power residuals at each time point are processed by an exponentially weighted moving average to obtain a smoothed power residual sequence.

[0034] As a further improvement to this technical solution, the specific steps involved in S4, which involve extracting the state correlation features between the highly disturbed unit and the upstream dominant unit to jointly form the fault detection features, are as follows:

[0035] For highly disturbed units and its upstream main units The power correlation is obtained by calculating the Pearson correlation coefficient between their power sequences;

[0036] Based on the comprehensive impact intensity value, the fluctuation variance of the wake impact intensity of the upstream dominant unit on the highly disturbed unit is calculated;

[0037] The smoothed power residual, each power correlation, each fluctuation variance, as well as the current wake velocity ratio and wake turbulence increment are sequentially concatenated to obtain a multidimensional feature vector;

[0038] Each feature in the multidimensional feature vector is standardized using z-score to obtain the fault detection features. .

[0039] As a further improvement to this technical solution, the specific steps involved in establishing several health condition modes through clustering algorithms in step S5 are as follows:

[0040] Fault detection features are extracted from all fault-free moments in historical data to form a training sample set;

[0041] The K-means clustering algorithm is used to cluster the training sample set, dividing the samples into groups. Classes, each corresponding to a health condition mode;

[0042] The goal of the K-means clustering algorithm is to minimize the sum of squared distances from all samples to the center of their respective clusters.

[0043] The center and sample assignment for each category are updated iteratively until the center no longer changes significantly, yielding the final result. One healthy operating condition mode;

[0044] For each healthy operating condition mode, calculate the mean of all its fault detection features, and use it as the center vector of that healthy operating condition mode. And calculate all its fault detection features to the center vector. The maximum Euclidean distance is used as the distribution radius of this health condition mode. ;

[0045] Record the frequency of each health condition pattern in historical data. .

[0046] As a further improvement to this technical solution, in step S5, the deviation of the real-time fault detection characteristics from the corresponding healthy operating condition mode is calculated, and the specific steps involved in obtaining the health indicators are as follows:

[0047] For real-time fault detection features Calculate its center vector to each health condition mode. Euclidean distance ;

[0048] The health condition mode with the smallest Euclidean distance is selected as the health condition mode to which the real-time fault detection feature belongs. ;

[0049] Based on real-time fault detection features The deviation is calculated by dividing the Euclidean distance from the center of the corresponding health condition model by the distribution radius of that model. ;

[0050] Based on deviation If it is not greater than 1, the health indicator is set to 1; if it is greater than 1, it is mapped to a health indicator through an exponential decay function.

[0051] If health indicators Less than the preset health threshold If the current state is abnormal, the verification process will be triggered.

[0052] As a further improvement to this technical solution, the specific steps involved in verification in step S6, which combines real-time wake field data, are as follows:

[0053] From real-time wake field data, obtain the current state of the highly disturbed unit. The wake characteristics, including the wake velocity ratio Wake turbulence increment Turbulence intensity at current location ;

[0054] Based on the current time and the previous The wake characteristics at each time point are calculated, and the current wake velocity ratio, wake turbulence increment, and absolute difference between the turbulence intensity and the mean of the corresponding parameters are calculated. These values ​​are then weighted and summed according to preset weights to obtain the wake abrupt change index. ;

[0055] If the wake mutation index Less than the preset wake mutation threshold If the current wake environment is relatively stable, the abnormal characteristics originate from a fault in the unit itself, and this is verified.

[0056] If the wake mutation index Greater than or equal to the preset wake mutation threshold If the error is not detected, it is determined to be a false anomaly caused by wake disturbance, and it is only recorded and not proceeded to subsequent fault detection.

[0057] As a further improvement to this technical solution, in step S6, the specific steps involved in anomaly detection using the isolated forest algorithm for the verified abnormal fault detection features are as follows:

[0058] Using fault detection features from historical normal data as the training set, a system is constructed... An isolated forest model consisting of isolated trees;

[0059] For real-time fault detection features The sample is then input into an isolation forest model, and its anomaly score is calculated based on the path length of the sample in each isolation tree. ;

[0060] If abnormal scores Greater than or equal to the significant fault threshold This is identified as a significant fault, triggering an advanced alarm.

[0061] If abnormal scores Below the significant fault threshold And greater than or equal to the warning threshold This is identified as a potential anomaly, triggering an alert.

[0062] If abnormal scores Below the warning threshold If the current state is determined to be a normal fluctuation, monitoring will continue.

[0063] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0064] 1. In this wind farm fault detection method based on historical power generation data, wake field data is generated through computational fluid dynamics simulation to identify highly disturbed units and upstream dominant units. Fault detection features are constructed by combining power residuals and unit status correlation characteristics. Compared with traditional detection methods that ignore wake interference, this method effectively distinguishes between wake pseudo-anomalies and real unit faults, and solves the problems of high false alarm rate and inability to adapt to the complex wake environment of wind farms in traditional methods.

[0065] 2. In this wind farm fault detection method based on historical power generation data, multiple healthy operating condition modes are established based on historical normal data through clustering algorithms. Anomaly classification detection is achieved by combining wake mutation verification and isolated forest algorithm. Compared with the traditional detection method based on a single operating condition benchmark, it can adapt to the differences in operating conditions under different wind speeds and directions. Furthermore, the fault type is assisted in localization through feature contribution quantification, which improves the accuracy of fault detection and operation and maintenance guidance, and meets the needs of efficient operation and maintenance of wind farms. Attached Figure Description

[0066] Figure 1 This is a flowchart of the overall method of the present invention. Detailed Implementation

[0067] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0068] Example: Please refer to Figure 1 As shown in the figure, this embodiment provides a wind farm fault detection method based on historical power generation data, including the following steps:

[0069] S1. Collect historical and real-time SCADA data streams from the wind farm, and perform data cleaning, alignment and normalization preprocessing on the SCADA data streams to obtain wind turbine operation data;

[0070] In this embodiment, outliers and invalid data are removed based on the physical reasonable range of each wind turbine operating parameter in the SCADA data stream, and short-term missing data is filled by interpolation; wherein, the wind turbine operating parameters include wind speed, wind direction, power, and rotational speed;

[0071] Specifically, for sampling points with numerical anomalies, linear interpolation is used to correct them by combining the data trends of the preceding and following times; for missing data, if the missing duration is less than a set threshold (e.g., 5 minutes), the average of adjacent time points is used to fill the missing data; if the missing duration is longer, the time period is marked as invalid data and will not be included in subsequent modeling.

[0072] The timestamps of data with different sampling frequencies in the cleaned SCADA data stream are aligned by means of aggregation.

[0073] Specifically, a fixed time window (e.g., 10 minutes) mean aggregation method is used to downsample high-frequency data (e.g., second-level vibration) to the same time granularity as low-frequency data (e.g., minute-level power); during the alignment process, it is necessary to ensure that all parameters have corresponding observation values ​​at the same time point to avoid modeling bias caused by time misalignment;

[0074] The minimum and maximum values ​​of the operating parameters of each wind turbine are calculated based on historical normal data, and the time-aligned data are normalized using the minimum-maximum normalization method.

[0075] The cleaned, aligned, and normalized data are integrated into a structured data table in chronological order, with each row representing a point in time and each column representing a wind turbine operating parameter, thus forming wind turbine operating data.

[0076] S2. Based on the topography and wind turbine layout parameters of the wind farm, wake field data that is spatiotemporally matched with the wind turbine operation data is generated through computational fluid dynamics simulation.

[0077] In this embodiment, a numerical simulation model is established in CFD software based on the topography of the wind farm and the wind turbine layout parameters (such as wind turbine coordinates, hub height, and rotor diameter).

[0078] The numerical simulation model should include actual environmental elements such as terrain undulations, vegetation cover, and buildings to realistically reflect the characteristics of the wind field.

[0079] Wind speed and direction are obtained from wind turbine operation data and used as inlet boundary conditions to input into the numerical simulation model. Simulation conditions corresponding to the data acquisition time points are set to ensure that the simulation time is synchronized with the actual operation time.

[0080] For each combination of wind speed and direction, the flow field distribution of the entire wind farm area is calculated using a numerical simulation model, and the simulation results are output. The simulation results include the wind speed, turbulence intensity, wind pressure at each wind turbine location, and the wake influence area between wind turbines.

[0081] For each wind turbine, within its upstream influence area (e.g., a fan-shaped area centered on the wind turbine with a radius of 10 times the impeller diameter), calculate the wake velocity ratio and wake turbulence increment at different locations in the simulation results.

[0082] The formula for calculating the wake velocity ratio is:

[0083]

[0084] in, This is the wake velocity ratio; a value less than 1 indicates that the wind speed is reduced due to the wake. The simulated wind speed at the downstream location. The upstream reference wind speed (usually taken as the wind speed at the wind farm inlet or the wind speed in front of the upstream wind turbine);

[0085] The formula for calculating the wake turbulence increment is:

[0086]

[0087] in, This is the turbulence increment; a value greater than 0 indicates that the wake causes enhanced turbulence. The turbulence intensity in the wake region, The turbulence intensity in free flow;

[0088] The simulation results, wake velocity ratio, wake turbulence increment, and wind turbine operating data are matched and aligned according to time and wind turbine number to construct corresponding wake features for each wind turbine at each moment, forming structured wake field data. Among them, the wake features include parameters such as wind speed, wind direction, wake velocity ratio, and wake turbulence increment at the location of the wind turbine.

[0089] S3. Based on wake field data, identify the highly disturbed units that are severely affected by the wake and their corresponding upstream leading units;

[0090] In this embodiment, based on the wake velocity ratio and wake turbulence increment of each wind turbine in the wake field data, it is determined whether the upstream wind turbine is located in the upstream influence area of ​​the downstream wind turbine for each pair of wind turbines in the wind farm.

[0091] If yes, the wake velocity ratio and wake turbulence increment are combined according to a preset weight to calculate the comprehensive influence intensity value, which is used to represent the magnitude of the wake influence of the upstream fan on the downstream fan; if no, the comprehensive influence intensity value is recorded as zero.

[0092] Specifically, for wind farms Typhoon generator, building a Influence intensity matrix , where matrix elements This indicates the upstream wind turbine at a specific wind speed and direction. For downstream wind turbines The intensity of the overall impact;

[0093] For any pair of wind turbines and ,like exist Calculate the overall impact intensity within the upstream influence area:

[0094]

[0095] in, This is the weighting coefficient for wind speed attenuation. For turbulence enhancement weighting coefficients (e.g., all taken as 0.5); wake velocity ratio The smaller the value (the greater the wind speed reduction), the greater the increase in wake turbulence. The larger the value (the more significant the turbulence enhancement), the stronger the overall influence. The larger the value, the stronger the influence.

[0096] like Not here Within the upstream influence area, ;

[0097] For each downstream wind turbine The total disturbance degree of the wind turbine is obtained by summing the combined impact intensity values ​​received from all upstream wind turbines. ;

[0098]

[0099] in, The total disturbance degree is represented by its value; the larger the value, the stronger the disturbance degree of the wind turbine. The more severe the impact of the wake from the upstream wind turbine on the location;

[0100] For all wind turbines, the total disturbance level at each time point is set above the preset disturbance threshold. The wind turbines are marked as currently highly disturbed.

[0101] in, The disturbance threshold is based on the total disturbance of all wind turbines. Historical statistical distribution (e.g., all wind turbines at all time points) The value is used to determine this; for example, the disturbance threshold is set to a historical value. The upper quartile (75th percentile) of the value distribution;

[0102] The proportion of time points marked for each wind turbine in historical data is counted. If the proportion exceeds a preset frequency threshold (e.g., 50%), the wind turbine is finally identified as a highly disturbed unit.

[0103] For each highly disturbed turbine unit, find the one with the largest combined influence intensity value from all its upstream turbines. The upstream wind turbines are designated as the upstream dominant turbines of the highly disturbed unit, i.e., the turbines that contribute the most to its wake impact.

[0104] Specifically, in highly disturbed units The corresponding wake influence vector ( Find the largest value in the list. Each element is recorded with its corresponding upstream turbine number. These upstream turbines are all considered as the main unit set to more comprehensively reflect the source of wake impact. This refers to the number of upstream main generating units (usually 2-3 units).

[0105] S4. For highly disturbed units, the power residual of the highly disturbed units is calculated by combining the wind turbine operation data and wake field data through the power prediction regression model, and its state correlation features with the upstream dominant units are extracted to form fault detection features.

[0106] In this embodiment, based on wind turbine operating data and wake field data, a power prediction regression model is constructed for each highly disturbed unit using a gradient boosting regression tree, with the wind turbine state vector at the current moment as the input. The output is the predicted power value of the wind turbine. The wind turbine state vector includes the current wind speed, wind direction, wake velocity ratio, wake turbulence increment, and turbulence intensity.

[0107] The prediction formula for the power prediction regression model is:

[0108]

[0109] in, For a moment The wind turbine state vector; Current wind speed (unit: m / s), from wind turbine operating data; Current wind direction (unit: °), from wind turbine operation data; The current wake velocity ratio is derived from wake field data; This represents the current increase in wake turbulence, derived from wake field data. The current turbulence intensity is derived from wake field data; The number of trees; For the first The prediction function of a regression tree; This is the predicted output power value;

[0110] The power prediction regression model is trained using historical normal data (data marked as fault-free periods), with actual power generation as the target value, mean squared error as the loss function, and a regularization term is added to prevent overfitting.

[0111] Loss function:

[0112]

[0113]

[0114] in, The loss function value is used to train a set of model hyperparameters that minimizes the value of the loss function. For a moment Actual power generation (unit: kW); This represents the number of training samples; is the regularization coefficient, a hyperparameter greater than 0; For regularization terms; For the first The complexity penalty term for trees penalizes trees with many leaf nodes (very deep trees) and large absolute values ​​of leaf node weights, encouraging the model to use simpler and smoother trees; A threshold parameter to control the difficulty of splitting the decision tree;

[0115] During training, 5-fold cross-validation is used to optimize hyperparameters (such as tree depth, learning rate, number of trees, etc.).

[0116] For highly disturbed turbine units in real-time data, their current turbine state vector is used. Input the trained power prediction regression model to obtain the power prediction value. By comparing the actual power output of the wind turbine with the actual power output of the wind turbine, the difference between the two is calculated to obtain the power residual. ;

[0117]

[0118] Among them, For a moment Power residual (unit: kW);

[0119] An exponentially weighted moving average is applied to the power residuals at each time point to obtain a smoothed power residual sequence, which can reflect the trend of power deviation.

[0120]

[0121] in, For a moment The smoothed power residuals are arranged in chronological order to form a smoothed power residual sequence. This is a smoothing coefficient, typically taken as 0.2 to 0.3; This is the smoothed power residual from the previous time step (initialized to 0).

[0122] In this embodiment, for highly disturbed units and its upstream main units The power correlation is obtained by calculating the Pearson correlation coefficient between their power series. ;

[0123]

[0124] in, For the unit In recent Power sequences over a time period of 1 hour are typically used. For the unit In recent Power sequence over time; For the unit With the unit The power correlation has a value range of [-1, 1], and the closer the value is to 1, the more consistent the power change trend is. This is the function for calculating the Pearson correlation coefficient;

[0125] Based on the comprehensive impact intensity value, the fluctuation variance of the wake influence intensity of the upstream dominant unit on the highly disturbed unit is calculated. ;

[0126]

[0127] in, For upstream units For highly disturbed units The comprehensive impact intensity value; To calculate the nearest The variance of the combined influence intensity value sequence over a period of time (usually 2 hours); In recent The variance of the fluctuation over time, which reflects the stability of the wake effect;

[0128] The smoothed power residual, each power correlation, each fluctuation variance, as well as the current wake velocity ratio and wake turbulence increment are sequentially concatenated to obtain a multidimensional feature vector;

[0129]

[0130] in, For multidimensional feature vectors, The power residual after smoothing; To correlate with the power output of each upstream main generating unit; To account for the variance of fluctuations in each of the upstream leading generating units; This refers to the number of upstream main generating units (usually 2-3 units).

[0131] Each feature in the multidimensional feature vector is standardized using z-score to obtain the fault detection features. ;

[0132]

[0133] in, This represents the mean of this feature in historical normal data. The standard deviation of this feature in historical normal data; These are fault detection features.

[0134] S5. Based on the fault detection features under historical normal data, several healthy operating condition modes are established through clustering algorithms, and the deviation of the real-time fault detection features from the corresponding healthy operating condition modes is calculated to obtain health indicators.

[0135] In this embodiment, fault detection features of all fault-free moments are extracted from historical data to form a training sample set;

[0136]

[0137] in, For the training sample set, For the first One fault detection feature sample, The number of samples;

[0138] The K-means clustering algorithm is used to cluster the training sample set, dividing the samples into groups. Classes, each corresponding to a health condition mode;

[0139] The goal of the K-means clustering algorithm is to minimize the sum of squared distances from all samples to the center of their respective clusters:

[0140]

[0141] in, The number of preset healthy operating condition modes (generally 3 to 5). Indicates Euclidean distance; The objective function value of the K-means clustering algorithm; For the first The center vector (mean) of the class; For the first A sample set of classes;

[0142] The center and sample assignment for each category are updated iteratively until the center no longer changes significantly, yielding the final result. One healthy operating condition mode;

[0143] For each healthy operating condition mode, calculate the mean of all its fault detection features, and use it as the center vector of that healthy operating condition mode. And calculate all its fault detection features to the center vector. The maximum Euclidean distance is used as the distribution radius of this health condition mode. ;

[0144]

[0145] in, For the first The distribution radius of the class; each class pattern is represented by its center vector. and distribution radius express;

[0146] Record the frequency of each health condition pattern in historical data. The proportion of this type of sample to the total number of samples is used as the weighting parameter for this health condition mode.

[0147] In this embodiment, for real-time fault detection features Calculate its center vector to each health condition mode. Euclidean distance ;

[0148]

[0149] in, For fault detection features to the first Euclidean distance of the center of the health-like working condition model; This is a real-time fault detection feature;

[0150] The health condition mode with the smallest Euclidean distance is selected as the health condition mode to which the real-time fault detection feature belongs. ;

[0151]

[0152] in, The health condition mode to which the real-time fault detection characteristics belong;

[0153] Based on real-time fault detection features The deviation is calculated by dividing the Euclidean distance from the center of the corresponding health condition model by the distribution radius of that model. ;

[0154]

[0155] in, for Euclidean distance to the center of the corresponding healthy working condition model; For the corresponding health condition mode The distribution radius; for Relative to the corresponding health condition mode The degree of deviation;

[0156] Based on deviation If it is not greater than 1, the health indicator is set to 1; if it is greater than 1, it is mapped to a health indicator between 0 and 1 through an exponential decay function.

[0157]

[0158] in, This is the attenuation coefficient (which can be taken as 0.5~1.0) to control the descent rate after exceeding the range; As a health indicator;

[0159] If health indicators Less than the preset health threshold If the value is 0.7, then the current state is considered abnormal, and the verification process is triggered.

[0160] S6. When the health indicators are less than the preset health threshold, the real-time wake field data is first used for verification. For the abnormal fault detection features that pass the verification, the isolated forest algorithm is used for anomaly detection.

[0161] In this embodiment, the highly disturbed unit is obtained from the real-time wake field data at the current moment. The wake characteristics, including the wake velocity ratio (Reflecting the degree of wind speed attenuation), wake turbulence increment (Reflects the degree of turbulence enhancement), turbulence intensity at the current location (Reflects the level of environmental turbulence);

[0162] Based on the current time and the previous A time (usually taken) The wake characteristics (representing the wake over the past hour) are calculated, and the absolute differences between the current wake velocity ratio, wake turbulence increment, and turbulence intensity and the mean values ​​of the corresponding parameters are calculated. These values ​​are then weighted and summed according to preset weights to obtain the wake abrupt change index. ;

[0163]

[0164] in, For the front The average of the wake velocity ratio at each moment For the front The mean value of the wake turbulence increment at each time step. For the front The average turbulence intensity at each moment; As the speed ratio weight, As the turbulence increment weight, The turbulence intensity weight is usually taken as... This reflects the differences in the sensitivity of different wake parameters to sudden changes; The wake abrupt change index indicates that the larger the value, the more drastic the change in the wake environment.

[0165] If the wake mutation index Less than the preset wake mutation threshold (Dimensionless, preferred) If the current wake environment is relatively stable, the abnormal characteristics originate from a fault in the unit itself, and this is verified.

[0166] If the wake mutation index Greater than or equal to the preset wake mutation threshold If the error is not detected, it is determined to be a false anomaly caused by wake disturbance, and it is only recorded and not proceeded to subsequent fault detection.

[0167] In this embodiment, fault detection features from historical normal data are used as the training set to construct a system. An isolated forest model consisting of isolated trees;

[0168] Each isolated tree isolates samples by randomly selecting features and segmentation values; abnormal samples, due to their large differences from normal patterns, have shorter paths in the tree and are easier to isolate.

[0169] For real-time fault detection features The sample is then input into an isolation forest model, and its anomaly score is calculated based on the path length of the sample in each isolation tree. ;

[0170]

[0171] in, For the sample The path length in each isolated tree; Represent the expected function; For the sample The average path length across all isolated trees; For the sample size The standardization factor at time is calculated using the following formula: , It is the harmonic number; The score represents the anomaly score; the closer it is to 1, the more likely the sample is to be anomalous; the closer it is to 0, the closer the sample is to the normal pattern.

[0172] If abnormal scores Greater than or equal to the significant fault threshold (For example This is identified as a significant fault and triggers an advanced alarm.

[0173] If abnormal scores Below the significant fault threshold And greater than or equal to the warning threshold (For example This is identified as a potential anomaly, triggering an alert.

[0174] If abnormal scores Below the warning threshold If so, the current state is determined to be a normal fluctuation, and monitoring continues;

[0175] For samples identified as significant faults and potential anomalies, the contribution of each feature is quantified by replacing the feature value with the historical normal mean and recalculating the anomaly score. The feature with the highest contribution corresponds to the possible fault type: abnormal power residuals indicate power generation efficiency problems, decreased power correlation indicates coordination anomalies, and abnormal wake velocity ratio and wake turbulence increment indicate yaw error. The system can use this information to assist in operation and maintenance location.

[0176] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A wind farm fault detection method based on historical power generation data, characterized in that: Includes the following steps: S1. Collect historical and real-time SCADA data streams from the wind farm, and preprocess the SCADA data streams to obtain wind turbine operation data; S2. Based on the topography and wind turbine layout parameters of the wind farm, wake field data that is spatiotemporally matched with the wind turbine operation data is generated through computational fluid dynamics simulation. S3. Based on wake field data, identify the highly disturbed units that are severely affected by the wake and their corresponding upstream leading units; S4. For highly disturbed units, the power residual of the highly disturbed units is calculated by combining the wind turbine operation data and wake field data through the power prediction regression model, and its state correlation features with the upstream dominant units are extracted to form fault detection features. S5. Based on the fault detection features under historical normal data, several healthy operating condition modes are established through clustering algorithms, and the deviation of the real-time fault detection features from the corresponding healthy operating condition modes is calculated to obtain health indicators. S6. When the health index is less than the preset health threshold, first verify it by combining real-time wake field data. For the abnormal fault detection features that pass the verification, use the isolated forest algorithm to detect the anomalies.

2. The wind farm fault detection method based on historical power generation data according to claim 1, characterized in that: In step S1, the specific steps involved in preprocessing the SCADA data stream to obtain wind turbine operating data are as follows: Based on the physical reasonable range of each wind turbine operating parameter in the SCADA data stream, outliers and invalid data are removed, and short-term missing data is filled by interpolation; among them, wind turbine operating parameters include wind speed, wind direction, power, and rotational speed; The timestamps of data with different sampling frequencies in the cleaned SCADA data stream are aligned by means of aggregation. The minimum and maximum values ​​of the operating parameters of each wind turbine are calculated based on historical normal data, and the time-aligned data are normalized using the minimum-maximum normalization method. The cleaned, aligned, and normalized data are integrated into a structured data table in chronological order, with each row representing a point in time and each column representing a wind turbine operating parameter, thus forming wind turbine operating data.

3. The wind farm fault detection method based on historical power generation data according to claim 1, characterized in that: In step S2, the specific steps involved in generating wake field data that spatiotemporally matches the wind turbine operating data through computational fluid dynamics simulation are as follows: A numerical simulation model was established in CFD software based on the topography and wind turbine layout parameters of the wind farm. Wind speed and direction are obtained from wind turbine operation data and used as inlet boundary conditions to input into the numerical simulation model. Simulation conditions corresponding to the data acquisition time points are also set. For each combination of wind speed and direction, the flow field distribution of the entire wind farm area is calculated using a numerical simulation model, and the simulation results are output. The simulation results include the wind speed, turbulence intensity, wind pressure at each wind turbine location, and the wake influence area between wind turbines. For each wind turbine, within its upstream influence area, calculate the wake velocity ratio and wake turbulence increment at different locations in the simulation results; The simulation results, wake velocity ratio, wake turbulence increment, and wind turbine operating data are matched and aligned according to time and wind turbine number to construct corresponding wake features for each wind turbine at each moment, forming structured wake field data. Among them, the wake features include the wind speed, wind direction, wake velocity ratio, and wake turbulence increment at the location of the wind turbine.

4. The wind farm fault detection method based on historical power generation data according to claim 1, characterized in that: In step S3, the specific steps involved in identifying the highly disturbed generating units severely affected by wake ripples and their corresponding upstream leading generating units are as follows: Based on the wake velocity ratio and wake turbulence increment of each wind turbine in the wake field data, for each pair of wind turbines in the wind farm, it is determined whether the upstream wind turbine is located in the upstream influence area of ​​the downstream wind turbine. If so, the wake velocity ratio and wake turbulence increment are combined according to preset weights to calculate the comprehensive influence intensity value; If not, the overall impact strength value is recorded as zero; For each downstream wind turbine The total disturbance degree of the wind turbine is obtained by summing the combined impact intensity values ​​received from all upstream wind turbines. ; For all wind turbines, the total disturbance level at each time point is set above the preset disturbance threshold. The wind turbines are marked as currently highly disturbed. The proportion of time points marked for each wind turbine in historical data is counted. If the proportion exceeds a preset frequency threshold, the wind turbine is finally identified as a highly disturbed unit. For each highly disturbed turbine unit, find the one with the largest combined influence intensity value from all its upstream turbines. Each upstream wind turbine is designated as the upstream leading turbine of the highly disturbed unit.

5. The wind farm fault detection method based on historical power generation data according to claim 1, characterized in that: In step S4, the specific steps involved in calculating the power residual of highly disturbed units using the power prediction regression model are as follows: Based on wind turbine operating data and wake field data, a power prediction regression model is constructed for each highly disturbed unit using a gradient boosting regression tree, with the current wind turbine state vector as the input. The output is the predicted power value of the wind turbine. The wind turbine state vector includes the current wind speed, wind direction, wake velocity ratio, wake turbulence increment, and turbulence intensity. The power prediction regression model is trained using historical normal data, with the actual power generation as the target value, the mean squared error as the loss function, and a regularization term is added to prevent overfitting. For highly disturbed turbine units in real-time data, their current turbine state vector is used. Input the trained power prediction regression model to obtain the power prediction value. By comparing the actual power output of the wind turbine with the actual power output of the wind turbine, the difference between the two is calculated to obtain the power residual. ; The power residuals at each time point are processed by an exponentially weighted moving average to obtain a smoothed power residual sequence.

6. The wind farm fault detection method based on historical power generation data according to claim 5, characterized in that: In step S4, the specific steps involved in extracting the state correlation features between the highly disturbed unit and the upstream dominant unit to jointly form the fault detection features are as follows: For highly disturbed units and its upstream main units The power correlation is obtained by calculating the Pearson correlation coefficient between their power sequences; Based on the comprehensive impact intensity value, the fluctuation variance of the wake impact intensity of the upstream dominant unit on the highly disturbed unit is calculated; The smoothed power residual, each power correlation, each fluctuation variance, as well as the current wake velocity ratio and wake turbulence increment are sequentially concatenated to obtain a multidimensional feature vector; Each feature in the multidimensional feature vector is standardized using z-score to obtain the fault detection features. .

7. The wind farm fault detection method based on historical power generation data according to claim 1, characterized in that: In step S5, the specific steps involved in establishing several health condition modes using a clustering algorithm are as follows: Fault detection features are extracted from all fault-free moments in historical data to form a training sample set; The K-means clustering algorithm is used to cluster the training sample set, dividing the samples into groups. Classes, each corresponding to a health condition mode; The goal of the K-means clustering algorithm is to minimize the sum of squared distances from all samples to the center of their respective clusters. The center and sample assignment for each category are updated iteratively until the center no longer changes significantly, yielding the final result. One healthy operating condition mode; For each healthy operating condition mode, calculate the mean of all its fault detection features, and use it as the center vector of that healthy operating condition mode. And calculate all its fault detection features to the center vector. The maximum Euclidean distance is used as the distribution radius of this health condition mode. ; Record the frequency of each health condition pattern in historical data. .

8. The wind farm fault detection method based on historical power generation data according to claim 7, characterized in that: In step S5, the deviation of real-time fault detection characteristics from the corresponding healthy operating condition mode is calculated, resulting in the specific steps involved in obtaining health indicators: For real-time fault detection features Calculate its center vector to each health condition mode. Euclidean distance ; The health condition mode with the smallest Euclidean distance is selected as the health condition mode to which the real-time fault detection feature belongs. ; Based on real-time fault detection features The deviation is calculated by dividing the Euclidean distance from the center of the corresponding health condition model by the distribution radius of that model. ; Based on deviation If it is not greater than 1, the health indicator is set to 1; if it is greater than 1, it is mapped to a health indicator through an exponential decay function. If health indicators Less than the preset health threshold If the current state is abnormal, the verification process will be triggered.

9. The wind farm fault detection method based on historical power generation data according to claim 1, characterized in that: In step S6, the specific steps involved in the verification using real-time wake field data are as follows: From real-time wake field data, obtain the current state of the highly disturbed unit. The wake characteristics, including the wake velocity ratio Wake turbulence increment Turbulence intensity at current location ; Based on the current time and the previous The wake characteristics at each time point are calculated, and the current wake velocity ratio, wake turbulence increment, and absolute difference between the turbulence intensity and the mean of the corresponding parameters are calculated. These values ​​are then weighted and summed according to preset weights to obtain the wake abrupt change index. ; If the wake mutation index Less than the preset wake mutation threshold If the current wake environment is relatively stable, the abnormal characteristics originate from a fault in the unit itself, and this is verified. If the wake mutation index Greater than or equal to the preset wake mutation threshold If the error is not detected, it is determined to be a false anomaly caused by wake disturbance, and it is only recorded and not proceeded to subsequent fault detection.

10. A wind farm fault detection method based on historical power generation data according to claim 9, characterized in that: In step S6, the specific steps involved in anomaly detection using the isolated forest algorithm on the verified abnormal fault detection features are as follows: Using fault detection features from historical normal data as the training set, a system is constructed... An isolated forest model consisting of isolated trees; For real-time fault detection features The sample is then input into an isolation forest model, and its anomaly score is calculated based on the path length of the sample in each isolation tree. ; If abnormal scores Greater than or equal to the significant fault threshold This is identified as a significant fault, triggering an advanced alarm. If abnormal scores Below the significant fault threshold And greater than or equal to the warning threshold This is identified as a potential anomaly, triggering an alert. If abnormal scores Below the warning threshold If the current state is determined to be a normal fluctuation, monitoring will continue.