Computer-implemented method for optimizing the operation of a wind turbine drivetrain

The hybrid model enhances fault diagnosis in wind turbine drivetrains by generating synthetic data and condition indicators, addressing dataset imbalances and improving model explainability for better maintenance and operation.

JP2025527167APending Publication Date: 2025-08-20FUNDACION TECNALIA RESEARCH & INNOVATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025504128
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-28
Filing Date
2023-07-25
Publication Date
2025-08-20

AI Technical Summary

Technical Problem

Current data-driven models for wind turbine drivetrain fault diagnosis suffer from imbalanced datasets and lack of explainability, leading to poor accuracy and decision-making confidence in maintenance and operation.

Method used

A hybrid approach combining physics-based and data-driven models to create a digital twin of the drivetrain, generating synthetic data for fault scenarios and enhancing existing classifiers with condition indicators, thereby optimizing operation and maintenance.

Benefits of technology

Improves fault detection and prediction accuracy by expanding training datasets and increasing model explainability, allowing for prescriptive analysis and optimized maintenance strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025527167000001_ABST
    Figure 2025527167000001_ABST
Patent Text Reader

Abstract

A computer-implemented method for training a classifier for fault diagnosis of a wind turbine drivetrain, comprising: training a physics-based model of the wind turbine drivetrain using operational data of the drivetrain under normal conditions, thereby providing a normal hybrid model of the wind turbine drivetrain; modeling at least one abnormal situation in the normal hybrid model and training it using operational data of the drivetrain under fault conditions, thereby providing a fault hybrid model of the wind turbine drivetrain; and applying a set of input data of the drivetrain under normal conditions to the normal hybrid model, thereby providing a fault hybrid model of the wind turbine drivetrain. 10. A computer-implemented method comprising: obtaining a set of synthetic data 280 under normal conditions; applying a set of fault data 150″ and a set of synthetic input data 710 associated with at least one abnormal situation of the drivetrain to a fault hybrid model 500, thereby generating a set of synthetic fault data 750 for the at least one abnormal situation; obtaining a set of condition indicators 8105 for the at least one abnormal situation from the set of drivetrain output data 152, 152′, the obtained set of synthetic data 280 under normal conditions, and the set of synthetic fault data 750 for the at least one abnormal situation; and training a classifier 900 using the set of condition indicators 8105 for the at least one abnormal situation. 11. A computer-implemented method for optimizing operation of a drivetrain of a wind turbine, comprising using the classifier 900 as previously trained.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the technical field of drive trains in wind turbines. In particular, it relates to a method and a system based on computer modeling and digitization for improving the efficiency of operation of and / or improving the maintenance of the power conversion system of a wind turbine. The present invention is particularly applicable to the wind energy sector. [Background technology]

[0002] Wind energy conversion is one of the most promising and reliable energy technologies today. Europe has already installed 220 GW of wind capacity and plans to install an additional 105 GW over the next five years (Wind Energy in Europe: 2020 Statistics and the outlook for 2021-2025. WindEurope, https: / / windeurope.org / intelligence-platform / product / wind-energy-in-europe-in-2020-trends-and-statistics / ). Parties involved in this energy source are continuously researching this technology with the aim of achieving the best levelized cost of energy (LCOE). According to WindEurope, operation and maintenance (O&M) costs account for 25-35% of a wind turbine's LCOE (Wind energy digitalization towards 2030. Cost reduction, better performance, safer operations. Published in November 2021 (https: / / windeurope.org / intelligence-platform / product / wind-energy-digitalization-towards-2030 / )), with repair and maintenance costs accounting for 30-60% of O&M costs (Hansen AD. Wind Energy Engineering: A Handbook for Onshore and Offshore Wind Turbines. Chapter 8 - Wind Turbine Technologies. 2017, 145-160. doi:10.1016 / B978-0-12-809451-8.00008-4).The possibilities offered by today's digitalization and artificial intelligence (AI) can significantly contribute to increasing wind farm energy production, reducing unplanned interruptions, optimizing O&M operations, and extending the lifespan of components.

[0003] Wind turbine systems can be classified according to the type of generator, gearbox, and power converter used. The double-fed induction generator (DFIG) with multi-stage gearbox and part-scale converter is a widespread technology (Hansen AD, Iov F, Blaabjerg F, Hansen LH. Review of Contemporary Wind Turbine Concepts and Their Market Penetration. Wind Engineering. 2004;28(3):247-263. doi:10.1260 / 0309524041590099). In the DFIG topology (S. Muller, M. Deicke, and R.W. De Doncker, "Doubly fed induction generator systems for wind turbines," in IEEE Industry Applications Magazine, vol. 8, no. 3, pp. 26-33, May-June 2002, doi:10.1109 / 2943.999610), the stator windings are directly connected to a constant-frequency utility grid, while the rotor windings are connected to the utility grid through a pulse-width modulation (PWM) power converter using a set of slip rings. Power converters can control rotor circuit current, frequency, and phase angle shift (F. Blaabjerg, M. Liserre, and K. Ma, "Power electronics converters for wind turbine systems," 2011 IEEE Energy Conversion Congress and Exposition, 2011, pp. 281-290, doi:10.1109 / ECCE.2011.6063781). This type of induction generator can operate over a wide slip range, typically ±30% of synchronous speed.As a result, they offer many advantages such as high energy yield, reduced mechanical stress and power fluctuations, and controllability of reactive power. The disadvantage of DFIGs is the unavoidable need for slip rings.

[0004] Wind turbines also include control systems that ensure proper operation of the wind turbine under all operating conditions and keep it within its normal operating range. Wind turbines may include electrical, mechanical, hydraulic, or pneumatic systems and require transducers to sense variables that will determine the control actions required. The most common variables monitored in a control system are wind speed, rotor speed, active and reactive power, voltage, and frequency at the wind turbine's connection points. In addition, the control system must be able to shut down the wind turbine if necessary. One control strategy is pitch angle control, which has proven to be an attractive option for variable-speed operation and for wind turbines larger than 1 MW (M.K. Dhar, M. Thasfiquzzaman, R.K. Dhar, M.T. Ahmed, and A.A. Mohsin, "Study on pitch angle control of a variable-speed wind turbine using different control strategies," 2017 IEEE International Conference on Power, Control, Signals and Instrumentation Engineering (ICPCSI), 2017, pp. 285-290, doi:10.1109 / ICPCSI.2017.8392258). Using this control, the blades can be quickly turned away from or turned towards the wind when the power output becomes too high or too low, respectively. The pitching system can be based on a hydraulic system controlled by a computer system or an electronically controlled electric motor. The pitch control system must be able to adjust the pitch angle a fraction of a degree at a time in response to changes in wind speed to maintain a constant power output.

[0005] Critical failure modes of wind turbine drivetrain systems, and in particular, generators and power conversion systems, have been analyzed (K. Alewine and W. Chen, "A review of electrical winding failures in wind turbine generators," in IEEE Electrical Insulation Magazine, vol. 28, no. 4, pp. 8-13, July-August 2012, doi:10.1109 / MEI.2012.6232004). Based on the generator failure source, (Tavner, PJ, Ran, Li, Penman, Jim, and Sedding, Howard, (2008), Condition Monitoring of Rotating Electrical Machines. Bibliographical OAI Repository, the University of Chicago). Press, doi:10.1049 / PBPO056E), the types of faults are: "thermal faults", caused by the effects of currents and overcurrents circulating through the windings on the insulation, taking into account the maximum temperature allowed depending on the type of insulation and the operating conditions; "electrical faults", caused by the voltages applied to the conductors under normal operating conditions and in abnormal situations such as surges coming from the converter; "environmental faults", caused by environmental conditions and their possible effects such as degradation of the insulating polymer, corrosion phenomena, etc.; "mechanical faults", mainly caused by vibrations; and "thermodynamic faults", caused by cyclic operating conditions with sudden or continuous changes in temperature, which affect differently each material that is part of the cable and its accessories (insulation, shielding, ...). The generator and the power converter are the components that have the greatest influence on the reliability, failure rate, and unavailability of the wind turbine.Their failure rate is 15% per year for generators and 6.8% for power converters in offshore wind farms (E. Byon, L. Ntaimo, C. Singh, and Y. Ding, "Wind energy facility reliability and maintenance", Handbook of Wind Power Systems, pp. 639-672; Shafiee y Dinmohammadi - An FMEA-Based Risk Assessment Approach for Wind Turbine Systems: A Comparative Study of Onshore and Offshore, 2014). These components are equipped with sensors (temperature, vibration, electrical measurements, etc.) and connected to wind turbine supervisory control and data acquisition (SCADA) and condition-based monitoring (CBM) systems. Therefore, a long history of real-world operational data sets exists for each turbine in a wind farm. Sometimes this data set includes recorded anomalies or faults in the operation of the turbine.

[0006] Computer modeling and digitalization are gaining importance in the wind energy sector because they provide tools for improving wind turbine design and performance to reduce both capital and operating costs. The introduction of large-scale sensors and increased big data processing capabilities over the last decade has made it possible to efficiently collect and analyze data, enabling the development and use of digital models of wind turbines (sometimes referred to as digital twins). Data-driven models apply AI techniques to extract knowledge from real-world measurements, which can help identify meaningful patterns in the large amounts of data collected from complex systems. Several approaches for this type of model exist in the wind energy generation technology field. For example, spectral analysis of measured currents has been used for health monitoring of main bearings and stator and rotor windings in wind turbines (T. Gerber, N. Martin, and C. Mailhes, "Time-Frequency Tracking of Spectral Structures Estimated by a Data-Driven Method," in IEEE Transactions on Industrial Electronics, vol. 62, no. 10, pp. 6616-6626, October 2015, doi:10.1109 / TIE.2015.2458781).In S. Yin, W. Guang, and H.R. Karimi, "Data-driven design of robust fault detection system for wind turbines," Mechatronics, vol. 24, no. 4, pp. 298-306, December 2014, a data-driven model was directly constructed to detect and isolate sensor and actuator faults in wind turbines. In contrast, in a study by E. Alizadeh, N. Meskin, and K. Khorasani, "A negative selection immune system inspired methodology for fault diagnosis of wind turbines," IEEE Trans Cybern, PP(99):1-15, 2016, a hierarchical bank of negative selection algorithms (NSAs) was developed to detect and isolate common faults in wind turbines. The research, M. Li, D. Yu, Z. Chen, K. Xiahou, T. Ji, and Q.H. Wu, "A Data-Driven Residual-Based Method for Fault Diagnosis and Isolation in Wind Turbines," in IEEE Transactions on Sustainable Energy, vol. 10, no. 2, pp. 895-904, April 2019, doi:10.1109 / TSTE.2018.2853990, uses a data-driven fault diagnosis and isolation (FDI) method for wind turbines, implementing a long short-term memory (LSTM) network for residual generation and applying a random forest algorithm for decision making.The main drawback of data-driven FDI methods is that they are designed based on experimental and historical data acquired from systems under normal and fault conditions. Such processes require a well-developed database containing labeled anomaly / failure data. Furthermore, the amount of anomaly / failure data is inherently scarce, resulting in an imbalanced dataset. As a result, the accuracy of data-driven methods is generally poor for cases not included in the training dataset. In addition, black-box models (e.g., deep learning models) exhibit low explainability, making it difficult for domain experts to interpret results and gain the confidence needed to make decisions based on the model's output.

[0007] Therefore, there is a need to develop new methods and systems for supervising and / or optimizing the operation and / or maintenance of, and / or predicting faults in, wind turbines that overcome the shortcomings of the prior art. Summary of the Invention [Means for solving the problem]

[0008] The present invention overcomes the drawbacks of prior methods and systems for condition monitoring of wind turbines by providing a computer-implemented method and system for optimizing the operation and / or maintenance of a wind turbine's drivetrain and improving its maintenance workflow. In the context of the present invention, a drivetrain should be understood as the set of components involved in wind energy conversion. Optimization is achieved using information obtained from a classifier used by the proposed method.

[0009] The method is based on different models of the drivetrain, including a physics-based model of the drivetrain trained using data analysis techniques with real operating data. The physics-based model of the drivetrain is a mathematical model that represents physical phenomena. The physics-based model includes multiple model parameters that describe the relationships between physical quantities that characterize the transformation of wind energy into mechanical power and mechanical power into electrical power. The physics-based model is trained using data analysis techniques with historical and real operating data to generate a so-called normal hybrid model.

[0010] Training a physics-based model using a data-driven algorithm that uses operational data under normal conditions allows for the calibration of values for at least some of the model parameters that define the physics-based model, resulting in a trained physics-based model referred to as a normal hybrid model. In other words, when a physics-based model is trained using normality data (data that represents normal operation), the model becomes a normal hybrid model. In the context of the present invention, the term "hybrid" refers to the combination of a mathematical model involving physical parameters and data analysis involving real data (measurements). Selected values for the parameters are obtained by minimizing the difference between predictions obtained by a physics-based model using a data-driven optimization algorithm trained with real operational data inputs and their corresponding known real operational data outputs.

[0011] Once the normality hybrid model is constructed, it can be extended or adapted to represent fault conditions, resulting in a so-called fault model. This new model is then trained using historical and ultimately actual operating data, both normal and faulty. This is done with the actual fault operating data inputs being fed to the fault model. In other words, when the normal hybrid model is adapted to represent faults and trained with fault data (data representing faulty operating behavior), it becomes a fault hybrid model. Feeding the fault data to the fault model allows for the calibration of the values of the fault model parameters that define the fault model. The selected values of these fault model parameters are obtained by minimizing the difference between the predictions obtained by the fault model using the fault operating data inputs and their corresponding known real operating data fault outputs. As a result, a so-called fault hybrid model of the wind turbine power conversion system (drivetrain) is obtained that considers data from both the drivetrain during normal and faulty operating conditions.

[0012] A digital twin can be defined as a virtual representation of a real system (in this case, the drivetrain of a wind turbine) that has the same behavior as the real one. Thus, a digital twin of a drivetrain is made up of physics-based models, fault models, and their association with the real system. This association refers to the information transferred (automatically or manually) from the wind turbine to the digital twin, as well as the information that can potentially be transferred from the digital twin to the system and its operators. Digital twins can be adapted for different applications.

[0013] On the one hand, the fault hybrid model can be used to generate synthetic data of failure scenarios and even events that have never occurred before. This makes it possible to gain knowledge of the drivetrain's behavior in specific conditions that would not otherwise be obtainable. Thus, the digital twin makes it possible to simulate realistic scenarios that would be difficult or costly to create in a real system. These scenarios can be used for prescriptive analysis of new operating conditions or to test responses to extreme conditions and anomalies or failures. This synthetic data is generated using real fault operating data and the fault hybrid model.

[0014] On the other hand, the digital twin allows existing data-driven models (e.g., machine learning classifiers) for estimating the status of the drivetrain to be improved, aiding in the decision-making process throughout its lifecycle (e.g., for diagnosing abnormal operation of the drivetrain or anomalies in measurement sensors and triggering alarms). In practice, the synthetic data generated by the digital twin (especially by the fault hybrid model) can be used to expand the training dataset and improve the generalization ability of existing data-driven models (e.g., classifiers). In addition, the digital twin can be used to create new features (also called condition indicators) that can be used as new inputs to the classifier, for example. These features are calculated by comparing real operating data (e.g., SCADA data) and / or synthetic data with normal data (predictions) generated by the normal hybrid model. In practice, the normal data generated by the normal hybrid model is used as a baseline for comparing the real asset with each digital replica in a normal (non-faulty) state under the same operating conditions (wind speed, pitch angle, torque, nacelle temperature, etc.).

[0015] In summary, the proposed methodology combines a physics-based model and a data-driven model for a wind turbine drivetrain to represent specific operating conditions, both under normal and abnormal or fault conditions. By enhancing the physics-based model with data-driven techniques, the physics-based model is optimized for the specific design and operating characteristics of the wind turbine in such a way that it can be used to generate synthetic data, classify faults and predict anomalies, and optimize the operation and maintenance of the wind turbine's power conversion system (drivetrain).

[0016] Developing a data-driven algorithm for diagnosing normal or fault conditions is a complex task that involves: i) defining condition indicators, ii) labeling correct and faulty operational data, iii) conceptualizing a classification model (supervised or unsupervised), iv) validating the model and analyzing its accuracy (e.g., number of false positives or false negatives), and v) analyzing whether it represents a machine or a set of machines (e.g., whether the model is valid only for a single wind turbine or for a set of different wind turbines). The proposed method and model can enhance this effort by providing additional synthetic data to enrich the dataset subsequently used by the classification algorithm.

[0017] In a first aspect of the present invention, a computer-implemented method for training a classifier for fault diagnosis of a drivetrain of a wind turbine is provided. The method includes training a physics-based model of the wind turbine drivetrain using operational data of the drivetrain under normal conditions, thus providing a normal hybrid model of the wind turbine drivetrain; modeling at least one abnormal condition in the normal hybrid model and training it using operational data of the drivetrain under fault conditions, thus providing a fault hybrid model of the wind turbine drivetrain; applying a set of input data of the drivetrain under normal conditions to the normal hybrid model, thereby obtaining a set of synthetic data under normal conditions for particular parameters of the drivetrain; applying a set of fault data and a set of synthetic input data associated with the at least one abnormal condition of the drivetrain to the fault hybrid model, thus generating a set of synthetic fault data for the at least one abnormal condition; obtaining a set of condition indicators for the at least one abnormal condition from the set of output data of the drivetrain, the obtained set of synthetic data under normal conditions, and the set of synthetic fault data for the at least one abnormal condition; and training a classifier using the set of condition indicators for the at least one abnormal condition.

[0018] In embodiments of the present invention, a physics-based model of a wind turbine drivetrain includes a first module representing an aerodynamic conversion stage of the drivetrain and a second module representing an electromechanical conversion stage of the drivetrain.

[0019] In embodiments of the present invention, the normal hybrid model is obtained by: providing a physics-based model of a drivetrain of a wind turbine, the physics-based model being represented by a plurality of model parameters and a set of equations that describe energy conversions in the drivetrain of the wind turbine; feeding the physics-based model with operational data of the drivetrain under normal conditions, the operational data including real operational data inputs and real operational data outputs; obtaining a prediction of at least one of the following parameters: generated power, phase currents, phase voltages, and electromagnetic torque; and selecting values of at least some of the model parameters that minimize the difference between the prediction corresponding to the real operational data inputs and the real operational data outputs, thus obtaining a physics-based model trained with the operational data.

[0020] In embodiments of the present invention, the fault hybrid model is obtained by: modeling at least one abnormal situation in a normal hybrid model; providing operational data of the drivetrain under fault conditions to the normal hybrid model including the at least one modeled abnormal situation; obtaining predictions of parameters indicative of the fault situation; and selecting values of at least some of the model parameters defining the normal hybrid model modeled with the at least one abnormal situation that minimize the difference between the predictions corresponding to the actual operational data input and the actual operational data output, thus obtaining a calibrated fault hybrid model.

[0021] In embodiments of the present invention, the at least one abnormal situation is modeled using at least one fault model.

[0022] In embodiments of the present invention, the at least one fault model is represented by a set of fault parameters and equations (mathematical formulas) that describe faults in the energy conversion of the drive train of the wind turbine.

[0023] In a second aspect of the invention, there is provided a computer-implemented method for optimizing the operation of a drivetrain of a wind turbine, the computer-implemented method comprising using a classifier as trained in the first aspect of the invention.

[0024] In embodiments of the present invention, the method includes obtaining a condition indicator from actual operational output data and normal synthetic data obtained from the normal hybrid model when the normal hybrid model is given the actual operational input data; providing the condition indicator to a trained classifier; and outputting, by the trained classifier, an output indicating whether the condition indicator corresponds to a fault condition or a non-fault condition.

[0025] In embodiments of the present invention, the method further comprises triggering an alarm when the classifier determines that the drivetrain is operating in an abnormal condition.

[0026] In embodiments of the present invention, the method further comprises acting on a drive train of the wind turbine when the condition indicator corresponds to a fault condition.

[0027] In a third aspect of the present invention, there is provided a device comprising processing means configured to perform the method according to the first aspect of the present invention.

[0028] In a fourth aspect of the present invention, there is provided a device comprising processing means configured to perform the method according to the second aspect of the present invention.

[0029] In a fifth aspect of the present invention there is provided a computer program product comprising computer program instructions / code for performing a method according to the first aspect of the invention and / or according to the second aspect of the invention.

[0030] In a sixth aspect of the present invention, there is provided a computer readable memory / medium storing program instructions / code for performing a method according to the first aspect of the invention and / or according to the second aspect of the invention.

[0031] The proposed methodology and model address the previously described limitations of current data-driven models. On the one hand, they allow the generation of synthetic data for abnormal / fault conditions and for never-occurring events, which can be used to expand and balance existing datasets and improve the generalization ability of data-driven models. On the other hand, the hybrid model preserves physical-design relationships, which promotes model explainability and increases end-user confidence.

[0032] The proposed method and system have been validated against real-world wind farm operating data, including labeled fault conditions. In particular, the use case of a drivetrain based on a doubly-fed induction generator (DFIG) and its corresponding part-scale power converter is proposed. However, the proposed method and system are also applicable to other wind power conversion technologies (e.g., permanent magnet generators).

[0033] Additional advantages and features of the present invention will become apparent from the following detailed description and are particularly pointed out in the appended claims.

[0034] To complete the description and to provide a deeper understanding of the invention, a set of drawings are provided. These drawings form an integral part of the description and illustrate one embodiment of the present invention. The embodiments should not be construed as limiting the scope of the invention, but merely as examples of how the invention may be practiced. The drawings include the following figures: [Brief explanation of the drawings]

[0035] [Figure 1]FIG. 1 is a diagram illustrating a physics-based model of a particular system and a process for training the physics-based model with real-world operating data, according to an embodiment of the present invention. [Figure 2A] FIG. 1 is a block diagram illustrating a physics-based model including two modules of a wind turbine power conversion drivetrain according to an embodiment of the present invention. [Figure 2B] FIG. 2B illustrates the physics-based model of FIG. 2A in more detail. [Figure 2C] FIG. 1 is a block diagram illustrating a normal hybrid model of a power conversion drivetrain of a wind turbine according to an embodiment of the present invention. [Figure 3A] 1 is a block diagram illustrating a fault hybrid model of a wind turbine power conversion drivetrain, according to an embodiment of the present invention, which is based on a properly calibrated normal hybrid model with additional models or states representing abnormal or fault conditions. [Figure 3B] FIG. 2 is a diagram illustrating an example of a normal hybrid model with an additional model, specifically a thermal model, representing an abnormal or fault condition, according to an embodiment of the present invention. [Figure 3C] FIG. 1 is a schematic diagram illustrating a fault model of a power conversion drivetrain and a process for training the fault model with real operating data, according to an embodiment of the present invention. [Figure 4A] FIG. 1 is a block diagram illustrating synthetic data generation using an impairment hybrid model according to an embodiment of the present invention. [Figure 4B] FIG. 1 is a block diagram illustrating synthetic data generation using multiple impairment hybrid models according to an embodiment of the present invention. [Figure 5A] FIG. 1 is a block diagram illustrating an overview of a proposed method for training a fault classifier according to an embodiment of the present invention. [Figure 5B] FIG. 1 is a schematic block diagram illustrating a proposed method for optimizing the operation of a wind turbine drivetrain using a trained classifier, according to an embodiment of the present invention. [Figure 6]FIG. 1 illustrates a Matlab-Simulink representation of a physics-based model of a wind turbine drivetrain according to an embodiment of the present invention. [Figure 7] Figure 7 shows the generated real (active) power (kW) versus wind speed (m / s) for the wind turbine drivetrain modeled in Figure 6. Four different power curves are shown: real data (measured in a SCADA system); a simulation using the model without calibration; a simulation using the model after a first calibration (six parameters); and a simulation using the model after a second calibration (seven additional parameters). [Figure 8] FIG. 1 shows a Matlab-Simulink representation of a thermal model (thermal circuit of the generator stator windings) added to a normal hybrid model of a wind turbine drivetrain used to model faults in the thermal circuit of the drivetrain generator stator windings. [Figure 9A] FIG. 1 shows five labeled anomaly cases (wind speed and active power signals) of overheating during real-world operation. [Figure 9B] FIG. 1 shows five labeled anomaly cases (wind speed and active power signals) of overheating during real-world operation. [Figure 9C] FIG. 1 shows five labeled anomaly cases (wind speed and active power signals) of overheating during real-world operation. [Figure 9D] FIG. 1 shows five labeled anomaly cases (wind speed and active power signals) of overheating during real-world operation. [Figure 10A] FIG. 10 illustrates generator stator winding temperature values (real and simulated) during an instance labeled an overtemperature fault. [Figure 10B] FIG. 10 illustrates generator stator winding temperature values (real and simulated) during an instance labeled an overtemperature fault. [Figure 10C]FIG. 10 illustrates generator stator winding temperature values (real and simulated) during an instance labeled an overtemperature fault. [Figure 10D] FIG. 10 illustrates generator stator winding temperature values (real and simulated) during an instance labeled an overtemperature fault. [Figure 10E] FIG. 10 illustrates generator stator winding temperature values (real and simulated) during an instance labeled an overtemperature fault. [Figure 11] FIG. 1 illustrates a methodology for synthetic data generation applied to a generator stator winding overheating scenario formed with a data-driven probabilistic fault model and a normal hybrid model, in accordance with an embodiment of the present invention. [Figure 12] FIG. 1 shows the stator winding temperature (non-bold line) calculated by a fault hybrid model (in this case a thermal model) from synthetic input variables, and as measured by the SCADA system (bold line). DETAILED DESCRIPTION OF THE INVENTION

[0036] The following description is not to be taken in a limiting sense, but is merely given for the purpose of illustrating the broad principles of the invention. The following embodiments of the invention will be described by way of example with reference to the above-mentioned drawings, which show apparatus and results according to the invention.

[0037] Next, different steps of a method for improving / optimizing the operation and / or maintenance of a wind turbine drivetrain are disclosed. Non-limiting examples of drivetrain components involved in wind energy conversion are main bearings, pitch system, gearbox, generator, power converter, bearings, and control system.

[0038] The method is based on computer modeling, simulation, and digitization. In the first stage of the proposed method, a physics-based model 100 of a wind turbine power conversion drivetrain is provided. The physics-based model 100 can be a newly created model or an existing physics-based model. FIG. 1 broadly discloses a physics-based model 100 of something (e.g., a device or system, such as a wind turbine power drivetrain). FIG. 1 also shows an optimization process 10 for training the physics-based model 100. The physics-based model 100 of a wind turbine power drivetrain is represented by a set of model parameters and equations that describe the conversion of wind kinetic energy to electrical energy. In other words, it is a mathematical model that describes the physical phenomena of power conversion occurring within the drivetrain, converting kinetic energy to electrical energy. The physics-based model 100 is constructed taking into account system design parameters. It describes the relationship between the physical magnitudes that characterize the conversion of kinetic energy in the wind to mechanical power and from mechanical power to electrical power.

[0039] FIG. 2A illustrates a block diagram of a non-limiting example of a physics-based model 100 of a wind turbine power conversion drivetrain. The physics-based model is shown in more detail in the example of FIG. 2B, where the modeled drivetrain uses a doubly-fed induction generator (DFIG) with back-to-back power converters. However, other wind energy conversion technologies could alternatively be used, such as a permanent magnet generator (PMG) with a full converter or a squirrel cage induction generator (SCIG), among others. The physics-based model 100 models a set of components involved in wind energy conversion, such as main bearings, pitch system, gearbox, generator, power converter, and their control systems. In FIGS. 2A-2B, the physics-based model 100 is divided into two main modules 101 and 102. These models can be used in conjunction with each other or operate independently, depending on the available operational data. The first module 101 represents the conversion of kinetic energy in the wind into mechanical power, represented schematically in FIG. 2B by the wind turbine blades and gearbox. The wind strikes the wind turbine blades, causing the rotor to rotate at a low speed on the main shaft. This low speed can be increased using the gearbox. The resulting mechanical torque after the gearbox in 101 is applied to the second module 102. The second module 102 represents the electromechanical conversion stage of the drivetrain, represented schematically in FIG. 2B by a doubly-fed induction generator (DFIG), a partial-scale power converter, and a corresponding control system (not shown) that enables optimal operation of the drivetrain. The generator and corresponding power converter enable the conversion of mechanical torque into electrical power.

[0040] In the shifting performed in the first module, the mechanical power extracted from the kinetic energy in the wind is determined by parameters such as air density, rotor swept, wind speed, and power coefficient. Other parameters related to the drivetrain's main shafts also typically play a role in the shifting. Non-limiting examples of these parameters are the wind turbine inertia constant, shaft spring constant, shaft cross-damping, and pitch angle.

[0041] In the conversion carried out in the second module, where mechanical power is converted to electrical power, the main elements involved are the generator and its corresponding power converter. When the modeled drivetrain uses a DFIG, the electrical part of the machine can be modeled, for example, by a fourth-order state-space model, and the mechanical part can be modeled, for example, by a second-order system. It requires the solution of several equations from both the electrical and mechanical systems. Non-limiting examples of parameters involved in this transformation are stator resistance and leakage inductance, magnetizing inductance, total stator inductance, q-axis stator voltage and current, d-axis stator voltage and current, stator q- and d-axis flux, rotor angular speed, rotor angular position, number of pole pairs, reference frame angular speed, mechanical angular speed, electrical angular speed, mechanical rotor angular position, electrical rotor angular position, electromagnetic torque, shaft mechanical torque, rotor and load combined inertia coefficient, rotor and load combined inertia constant, rotor and load combined viscous friction coefficient, total rotor inductance, rotor resistance and leakage inductance, q-axis rotor voltage and current, d-axis rotor voltage and current, rotor q- and d-axis flux, DC bus voltage regulator gain, grid-side converter current regulator gain, rotor-side converter current regulator gain.

[0042] Depending on the nature of the equipment, it may be difficult to obtain values for the complete set of design parameters. In this case, estimations are required, which may affect the performance of the model. While nominal values of DFIG parameters, such as rated power, voltage, current, frequency, number of pole pairs, or rotor angular speed, are usually known, characteristic values corresponding to the stator winding design (stator resistance and leakage inductance, magnetizing inductance, total stator inductance) are rarely known.

[0043] The inputs 151 to the physics-based model 100 (inputs 121, 122 to the first module 101) are actual values (measurements) under normal conditions, such as SCADA values. In the model shown in FIG. 2B , these inputs are actual values of the wind speed 121 (e.g., in meters per second, m / s) hitting the turbine blades and the blade pitch angle 122 (e.g., in degrees, deg), acquired, for example, periodically, at specific time intervals. In embodiments of the present invention, the modules 101, 102 may be modeled independently of each other. In this case, when the second module 102 is modeled independently, the actual value under normal conditions of the mechanical torque 131 in the wind turbine main shaft can also be an input to the model 100. More precisely, although the data 151, 152 used to train the physics-based model 100 are generally considered to be data under normal conditions, in practice, certain input data 151, by its nature, cannot be distinguished between fault and normal. This is the case, for example, for input data corresponding to wind speed 121, which, by itself, cannot be considered to represent a fault condition. Other input data 151, such as machine torque 131, can actually represent either a normal or fault condition. Similarly, output data 152, 141 can generally represent either a normal or fault condition. Here, the data 151, 152 used to train the physics-based model 100 is normal-state data. The data used to train the physics-based model 100 (particularly the specific input data 131 and output data 152, 141, since the other input data 121, 122 are essentially normal-state data) has preferably been previously labeled by a human operator as "normal-state data." Therefore, the physics-based model 100 is trained using normal-state data. The input values 151, 121, 122, 131 are discrete values of historical normal-state data collected at multiple times within a particular operational period. Typically, the available data corresponds to SCADA data registered, for example, every 10 minutes. However, data with a higher acquisition frequency could also be available. The acquisition frequency is in the order of microseconds (1 μs=10-6 The acquisition frequency may vary from every 100 μs or 1 ms, for example, to every second or minute, for example, every 1 second, every 1 minute, or every 10 minutes. The acquisition frequency may depend on different circumstances, such as the type of parameter being acquired.

[0044] Referring again to the exemplary model of FIG. 2B , the output 131 of the first module 101 is a prediction of the generated mechanical torque acting on the shaft of a doubly-fed induction generator (DFIG) (belonging to the second module 102). This prediction is obtained by calculating the equations of which the model is constructed. The second module 102 takes as input the output 131 of the first module 101, i.e., the mechanical torque acting on the shaft of the DFIG. The output 141 of the second module 102 includes a prediction of the generated power and its associated signals, such as phase current, phase voltage, frequency, and / or electromagnetic torque, for a specific time interval to be sent to the power grid. In a preferred embodiment, the output 141 is a prediction of the generated power and its corresponding phase current. This prediction is obtained by calculating the equations of which the model is constructed. This prediction is referenced as 180 in FIG. 1 .

[0045] Thus, the real operating data 150 under normal conditions includes real inputs (discrete historical values of real operating data, typically measurements) 151 to the physics-based model 100 under normal conditions (e.g., measurements of specific parameters such as wind speed, pitch angle, or mechanical torque, as in the case of the exemplary model of FIG. 2B ), as well as real outputs 152 of the drivetrain of the modeled wind turbine (e.g., measured outputs of specific parameters, e.g., generated power and phase current, that correspond to or are associated with the real inputs 151). The real operating data 150 are discrete values of historical data collected at multiple times within a particular operational period. Reference numeral 180 represents a prediction of a parameter, such as a prediction of generated power (prediction 141 in FIG. 2B ).

[0046] Once the physics-based model 100 is constructed, it is optimized (trained) by applying one or more data-driven optimization algorithms 190 using historical, real operating data 150 under normal conditions to it, thereby obtaining a model augmented with real operating data, the normal hybrid model 200, as shown in FIGS. 1 and 2C. The optimization / training process is diagrammed in FIG. 1, which shows a training process or algorithm 10 for calibrating the physics-based model 100. In FIG. 1, the real operating data 150 under normal conditions is provided to the physics-based model 100 as real input 151 or provided to further stages of the calibration process as real output 152. The calibration process involves optimizing design parameters (also called model parameters), whose accurate values may be estimated given certain physical constraints (e.g., within a given realistic interval), using an objective function to minimize residuals 160. Preferably, among the design parameters to be optimized are parameters associated with the electromechanical conversion stage (e.g., but not limited to, the generator, power converter, and wind turbine control), parameters associated with the aerodynamic conversion, parameters associated with the control strategy, and parameters associated with the pitch system and mechanical parts of the drivetrain. Non-limiting examples of parameters associated with the electromechanical conversion stage that may be selected for optimization are the stator winding resistance, the rotor winding resistance, the generator inertia constant and the generator friction coefficient, the power converter grid-side coupling resistance, the power converter grid-side coupling inductance, and the converter line filter capacitor. Non-limiting examples of parameters associated with the aerodynamic conversion stage that may be selected for optimization are the wind turbine inertia constant, the shaft mutual damping, and the shaft spring constant. Non-limiting examples of parameters associated with the control strategy that may be selected for optimization are the DC bus voltage regulator gain, the speed regulator gain, and the wind speed at nominal speed and at Cp max.

[0047] Thus, the training process 10 allows for the estimation of values of specific design variables (parameters) of the physics-based model 100 that are not known a priori from a set (dataset) of historical operating data 151, 152, 121, 122, 131 under normal conditions that contains a sufficient representation of the system behavior, in order to minimize the difference between the output values (predictions) 141, 180 of the variables of the physics-based model 100 and the actual system values (outputs) 152. The training process 10 shown in FIG. 1 aims to minimize one or more cost functions (or objective functions). In a particular embodiment, there are two cost functions: the first is the difference between the active power estimated (predicted) by the physics-based model 100 and the actual active power (generally represented by reference numeral 152), typically provided by SCADA; and the second is the difference between the phase current estimated (predicted) by the physics-based model 100 and the actual phase current (generally represented by reference numeral 152), typically provided by SCADA; however, a different number of cost functions may be present. In this example, the first cost function aims to minimize the power difference, and the second cost function aims to minimize the phase current difference. The objective function of the process is the minimization of the residual 160, defined as the difference between the physics-based model output (e.g., power prediction) 180 and the actual operating data 152 (e.g., output power) for a given real-world input 151 (e.g., wind speed or torque). The optimization algorithm aims to find a combination of parameter values that minimizes the difference between the physics-based model 100 output (prediction 180) and the measured data 152.

[0048] In embodiments of the present invention, the optimization algorithm used by the training process to calibrate (adjust or tune) the parameters is a multivariate input variable x subject to linear and nonlinear constraints, and a cost function x with some finite bounds.

number

[0049] In other words, the training process is performed as follows: Using real operating data 150 (real input data 151) under normal conditions, an initial prediction 180 of one or more parameters (e.g., generated active power and / or at least one related signal or parameter (phase current, phase voltage, and / or electromagnetic torque) is made. By comparing real outputs 152 (operating data under normal conditions, e.g., measurements of output power and / or at least one related signal or parameter provided by the drive train of the wind turbine for a given real input 151) with the obtained prediction 180, a residual 160 is obtained. Depending on this residual 160, one or more model parameters defining the physics-based model 100 are modified (stage 170). Training The process is repeated n times, as shown schematically at the bottom of FIG. 1 . After n iterations, the physics model 100 is considered a normal hybrid model 200. Generally, the training process continues for a given number of iterations or until a point is found where the objective function falls below a threshold. The normal hybrid model 200 thus consists of a physics-based model augmented with operational data (such as SCADA data) under normal scenario conditions. In other words, the normal hybrid model is the resulting calibrated physics model 200. Here, the term hybrid refers to the need for both the physics-based model 100 and the data-driven optimization algorithm 190 trained with real-world data in its construction.

[0050] Once the normal hybrid model 200 is developed, it can be extended or adapted to include abnormal or fault conditions to study the behavior of the system (wind turbine drivetrain) under such conditions. This is illustrated in FIG. 3A . The normal hybrid model 200, together with one or more additional models 301-306 representing possible fault conditions, forms an extended or fault model 400. The additional models 301-306, also referred to as fault models, can represent a wide variety of models representing different malfunctions of the wind turbine drivetrain, including, but not limited to, aerodynamic and / or mechanical malfunctions in the dynamics conversion stage 101 and / or magnetic, electrical, mechanical, and / or thermal malfunctions in the electromechanical conversion stage 102 of the wind turbine drivetrain, such as overheating or short circuits in the stator or rotor windings, demagnetization of permanent magnets, mechanical faults in the shaft or bearings, among others. The fault models 301-306 are represented by a number of model parameters and a set of equations representing the fault conditions or anomalies. Thus, the fault model allows for the creation or modeling of at least one abnormal or fault condition. The fault model can be generated using new, separate, additional physics-based modules 301-306 that represent the abnormal or fault condition, or it can represent a specific condition (typically, a fault state) integrated into the main body's normal hybrid model 200. For example, the fault model can be represented by adding a separate block to the normal hybrid model 200 or by simply modifying certain parameters of this final normal hybrid model 200. Alternatively, the fault model can be inherently integrated into the physics-based model 100 or the normal hybrid model 200, where it adds a potential fault condition. Exemplary additional models (fault models) implemented as separate blocks shown in FIG. 3A are the aerodynamic model 301 and mechanical model 302 associated with the dynamics conversion module 101, and the electrical model 303, magnetic model 304, mechanical model 305, and thermal model 306 associated with the electromechanical conversion module 102.Other non-limiting examples of additional models inherently integrated within the normal hybrid model 200 that can be used to enhance the normal hybrid model 200 with other fault conditions include short circuit models (e.g., coil-to-coil, phase-to-phase, or to ground) in the stator or rotor, specific sensor readings (indicative of a fault), partial slot discharge models, core fault models, structural fault models, permanent magnet demagnetization, among others.

[0051] The extended model 400 formed by the normal hybrid model 200 plus one or more additional models 301-306 (fault models) representing possible fault conditions may be referred to as the fault model 400. FIG. 3B illustrates an example of the fault model 400 formed by the normal hybrid model 200 and a thermal model 306 representing overheating of the DFIG stator windings. The thermal model 306 is required to represent an overheating abnormal condition. In the illustrated example, the inputs to the fault model (thermal model) 306 are predictions 306-2 of one or more parameters (e.g., predictions of stator phase currents) as predicted by the normal hybrid model 200 (prediction 280 in FIG. 1 ) and actual values 306-1 of one or more parameters (e.g., measurements collected, for example, by a SCADA system), such as the actual value of the nacelle temperature. The fault model 306 provides a predicted output 350, which in this illustrative case corresponds to the temperature of the DFIG stator windings. In this case, an anomaly in the cooling system of the DFIG can be identified when the actual value 306-1 of the stator winding temperature for a particular current differs from, e.g., is higher than, the value estimated by the DT (prediction 350) by a particular amount.

[0052] 3B is an example of an additional model (fault model) independent of the normal hybrid model 200, since the thermal model 306 is only linked to the normal hybrid model 200 through the prediction 306-2. Depending on the fault conditions to be modeled, the fault models may be interrelated with the normal hybrid model 200, e.g., may be inherently integrated within the normal hybrid model 200. For example, the resulting fault model 400 may be based on the normal hybrid model 200, modified to represent the fault conditions. In this case, there are no external fault models 301-306, and the fault conditions are internally integrated within the normal hybrid model 200, forming the extended model 400 (fault model).

[0053] Once the fault model 400 (normal-hybrid model 200 plus a representation of one or more possible fault conditions) is constructed, it is optimized (trained) by applying one or more data-driven optimization algorithms 450 trained with real operating data 150' under fault conditions to obtain the fault hybrid model 500, which is a model enhanced (calibrated) with real fault data, as shown in Figures 3A and 3C. The training process diagrammed in Figure 3C is similar to the training process applied to the physics-based model 100 (Figure 1) to obtain the normal hybrid model 200. The new calibrated model (also referred to as the fault hybrid model) is referred to as 500.

[0054] Calibration or training of the fault model 400 is performed using real operating data 150′ under fault conditions as follows. The real operating data 150′ under fault conditions includes real inputs (discrete historical values, typically measurements, of real operating data that cannot be considered to refer to a fault condition due to their nature (e.g., wind speed)) 151 (e.g., measurements of specific parameters such as wind speed, pitch angle, or machine torque) to the fault model 400 under normal conditions, as well as real outputs 152′ under fault conditions of the modeled wind turbine drivetrain (e.g., measured outputs of specific parameters, e.g., generated power and phase current, that correspond to or are related to the real inputs 151 but are affected by the fault model). The real operating data 150′ are discrete values of historical data (e.g., SCADA data) collected at multiple times within a specific operational period. The output data 152′ used to train the fault model 400 has preferably been previously labeled as “data under fault conditions” by a human operator. Therefore, the fault model 400 is trained using data under fault conditions. Using real operating data 150' (real input data 151, such as the nacelle temperature where the generator is installed or the real room temperature), an initial prediction 180', e.g., a prediction 180' of the stator winding temperature, is made. A residual 160' is obtained by comparing a real output 152' (measured operating data provided by the wind turbine drivetrain and corresponding to the given real input 151, e.g., the stator winding temperature in this example) under fault conditions with the obtained prediction 180'. Depending on this residual 160', at least some model parameters defining the extended model 400 are modified (step 170'). The parameters to be modified may be parameters of the additional models 301-306, parameters of the normal hybrid model 200, or both (the complete fault model 400), depending on how the fault condition is represented. For example, some model parameters defining and representing the fault may be optimized. The exact values of these parameters have been previously estimated by imposing certain physical constraints (eg, by imposing values within a given realistic interval).Non-limiting examples of parameters to be optimized in the fault model (e.g., in particular, in the thermal model) are material thermal conductivity, area perpendicular to the heat flow direction, thickness of the layer where conduction occurs, heat transfer coefficient, and surface area in contact with air. In FIG. 3C, reference numeral 155 indicates a prediction of one or more parameters (e.g., 306-2 in FIG. 3B) (e.g., prediction of stator phase current) as predicted by the normal hybrid model 200 and provided to the additional models 301-306. The calibration process is repeated n' times, as shown schematically at the bottom of FIG. 3C. After n' iterations, the fault model 400 is considered to be the fault hybrid model 500. Generally, the calibration process continues for a given number of iterations or until a point is found where the objective function (cost function) is below a threshold.

[0055] The fault model 400 is trained by applying one or more data-driven optimization algorithms 450 trained with real operating data 150′, as represented schematically in FIG. 3A. The optimization algorithm allows for the estimation of values of certain design variables (parameters) of the fault model 400 that are not known a priori from a set of historical operating data 150′ that contains a sufficient representation of the system behavior in order to minimize the difference between the output values (predictions) 180′ of the variables of the fault model 400 and the real system values (outputs) 152′. The optimization algorithm shown in FIG. 3C aims to minimize one or more cost functions (or objective functions). The optimization algorithm aims to find a combination of parameter values that minimizes the difference between the output (predictions) of the augmented model 400 and the measured operating data 150′. In embodiments of the present invention, the optimization algorithm 450 used to calibrate (adjust or tune) the parameters is a multivariate input variable x, subject to linear and nonlinear constraints, as well as a cost function x with some finite bounds.

number

[0056] The previously described training process of the fault model 400, resulting in the fault hybrid model 500, allows for simulating fault conditions present in the training operating data 150′ (i.e., the collection of historical data, such as SCADA data, used to train the fault model 400).

[0057] So far, a fault hybrid model 500 has been obtained that considers an abnormal or fault condition. There can be one or more fault hybrid models 500, including one or several faults. In particular, there can be as many fault hybrid models 500 as there are different fault conditions to be simulated. Therefore, the thermal model used as an example in this description should not be considered limiting. Nor should the thermal model be considered the only possible fault model within the fault hybrid model 500. A single fault hybrid model 500 will represent (model) one fault condition or model. Alternatively, there can be as many fault hybrid models 500 as there are fault conditions to be modeled.

[0058] The fault hybrid model 500 was calibrated to simulate fault conditions present in the training operational data (e.g., training SCADA data). However, this does not guarantee that these model runs are representative of the entire anomaly feature space. In reality, the frequency of abnormal conditions and faults is relatively low in the actual operational data (e.g., SCADA data 150, 150′), and often these instances are not annotated (labeled). Therefore, relying solely on a deterministic model to generate synthetic fault scenarios would result in a narrow data sample constraint to previously seen patterns.

[0059] To address this limitation, the proposed method incorporates a synthetic data generation stage, as shown schematically in FIGS. 4A-4B. This process uses the previously obtained fault hybrid model 500 to generate synthetic data 750. Synthetic data generation involves a probabilistic model 705, also referred to as a probabilistic synthetic data generation model, which estimates a probability distribution of the occurrence of the failure data and, accordingly, generates synthetic input data 710 associated with each fault. FIG. 4A shows a single fault condition or situation, labeled as "Failure Mode 1." Therefore, two sets of failure data are provided to the fault hybrid model 500: real data 150" (e.g., measured historical data) representing or associated with the failure situation, and the generated synthetic data 710 associated with the fault condition. The synthetic input data 710 and the real data 150" for each fault are used as inputs for the fault hybrid model 500 to generate the synthetic failure data 750. The actual labeled failure data 150'' is historical real-world operational data (typically SCADA data) that has been identified and labeled as failure data (150' in FIG. 3C). The synthetic failure data 750 is generated using the fault hybrid model 500 activated by a predetermined probability distribution 705 of occurrence of the failure data 710. This probability distribution of occurrence of the failure data is calculated from observations of the labeled failure data 150''. In FIG. 4B , the synthetic data generation process is completed incrementally relative to the process of FIG. 4A . It adds as many fault hybrid models 500 as needed to model a set of failure situations (labeled as “Failure Mode 1”, …, “Failure Mode n”), as well as actual labeled failure data 150″ (if present) and / or a corresponding set of synthetic data 710 representing each failure situation. As in FIG. 4A , the synthetic data 710 is obtained from a probabilistic model 705 that estimates a probability distribution of occurrence of failure data associated with each failure situation (“Failure Mode 1”, …, “Failure Mode n”) and accordingly generates synthetic input data 710 associated with each fault. Each fault hybrid model 500 can generate an unlimited number of synthetic failure data 750 for a particular failure type (“Failure Mode 1”, …, “Failure Mode n”) based on the actual failure data 150″ and using the probabilistic model 705.

[0060] Figures 5A and 5B show a schematic block diagram of the proposed method for optimizing the operation of a wind turbine drivetrain: Figure 5A schematically illustrates training of a fault classifier; and Figure 5B schematically illustrates optimizing operation and / or maintenance of the wind turbine drivetrain using the trained classifier.

[0061] FIG. 5A shows different stages for training a classifier 900 based on different models of the wind turbine drivetrain. Specifically, both a normal hybrid model 200 (trained under normal conditions) and a fault hybrid model 500 (trained under fault conditions) are used (the fault hybrid model 500 is not shown, but it is implicitly included in the stages 700, 700′ for synthetic data generation, see FIGS. 4A-4B ). There may be different fault hybrid models 500 for representing different fault conditions (“Failure Mode 1,” …, “Failure Mode n” in FIGS. 4A-4B ). Depending on the fault to be modeled, it may be possible to include different fault scenarios within a single fault hybrid model 500, as broadly shown in FIG. 3A . The stages of normalization / feature creation 800 and training of the fault classifier 900 are also shown schematically. FIG. 5B shows a schematic representation of the use of the trained fault classifier 900 to perform fault detection from actual real-world operating data of one or more drivetrains.

[0062] In a first stage, a set of actual operating data (typically SCADA data) 151 under normal conditions is provided to the normal hybrid model 200, previously generated according to the process shown in Figure 1. The normal hybrid model 200 then provides predictions 280 (a corresponding set of predictions). In other words, the predictions 280 are synthetic data associated with the normal conditions. The predictions 280 are normal synthetic data representing at least one relevant signal or parameter, such as generated power and phase current.

[0063] Another stage, which may be concurrent with the first stage, is the synthetic fault data generation process 700, 700′. One or more hybrid fault models 500 are used to generate synthetic fault data 750 as described with reference to FIGS. 4A-4B. For example, without limitation, if the fault model is a thermal model 306, the synthetic data 750 may be a collection of synthetic fault data representing stator winding temperatures. The synthetic fault data 750 is obtained from a set of actual labeled fault data 150″ as shown in FIG. 5A and synthetic data 710 representing respective fault conditions (disclosed in FIGS. 4A-4B but not shown in the schematic diagram of FIG. 5A).

[0064] Previously, a set of normal synthetic data 280 representing a wide range of normal conditions and a set of fault synthetic data 750 for each type of fault being modeled (see, for example, FIGS. 3A-3C ) have been generated. Next, a new set of features, called condition indicators, is created for each type of fault. In the context of the present invention, a condition indicator is any feature from which an abnormal condition can be detected. Features such as difference, division, slope, root mean square (RMS), peak, peak-to-peak, kurtosis, and crest factor, among others, are widely used for wind turbine fault diagnosis. A condition indicator 8105 is generated from a labeled dataset composed of the real output data 152, 152′ (which can correspond to both normal or fault conditions), the augmented normal synthetic data 280 generated using the methodology already described, and the augmented fault synthetic data 750.

[0065] Both the normal hybrid model 200 and the fault hybrid model(s) 500 are used as a baseline for creating condition indicators 8105 that can improve the accuracy of the fault classifier 900. The fault classifier 900 is a supervised classifier for fault diagnosis. To train the classifier 900, a stage of classification of a normal or faulty state of the wind turbine (block 900) is included. The classifier 900 is trained using the corresponding condition indicators 8105 previously generated (block 800). In other words, the classification is trained using the condition indicators 8105, whose value is driven by the relationship between the actual operating output data 152, 152′, the predictions (synthetic data under normal conditions) 280, and the synthetic fault data 750.

[0066] In stage 800, condition indicators 8105 are calculated by establishing a comparison between the different data input to this stage (stored output measurements 152, 152′ (such as SCADA data), normal composite data 280, and fault composite data 750). For example, depending on the type of fault represented by the corresponding data, the condition indicators 8105 can be, for example, the result of comparing the normal composite data 280 with the operational output data 152, 152′, or the result of comparing the synthetic fault data 750 with the operational output data 152, 152′, or the result of comparing the synthetic fault data 750 with the normal composite data 280, or the result of comparing the synthetic fault data 750 with the normal composite data 280 and with the operational output data 152, 152′. In addition, different types of comparison operations can be established, such as difference, division, slope, or others, to be applied to the data to be compared. Thus, in stage 800, a set of condition indicators 8105 is obtained: one for each type of fault being modeled. For example, under ideal conditions (a fault-free system), when the comparison is between the actual operating data and the normal composite data (minimum difference), the condition indicator 8105 will be close to 0, while when the actual operating data 152' in a fault condition is compared to the normal composite data 280, the condition indicator 8105 will have a high value in the fault condition. A fault type may have one or more associated condition indicators.

[0067] For example, the condition indicator 8105 can be the difference between the actual current 152, 152′ measured in the SCADA system and the current 280 simulated by the normal hybrid model 200. Under no-fault conditions, this difference will be near zero. If this difference exceeds a threshold, it indicates an anomaly to be classified in the next stage of training the fault classifier 900. Another example of the condition indicator 8105 can be the difference between the actual stator temperature and the estimated temperature value determined under fault conditions. In this case, a high value of the actual stator winding temperature can be compared with its corresponding fault synthetic data. If the result of this comparison is within a certain threshold, it indicates an anomaly to be classified in the next stage of training the fault classifier 900.

[0068] The above-mentioned condition indicators 8105 can also improve the generalization ability of the classifier 900, allowing one single classifier to be trained for multiple turbines (e.g., all of them) of an energy farm. Generally, if only real data were used to train the classifier 900, it would be difficult to obtain a common threshold corresponding to a wide collection of turbines. In fact, each turbine usually has different behavior under the same operating conditions due to differences during manufacturing, different locations, etc. Having a single classifier 900 for multiple wind turbines brings many advantages. On the one hand, it facilitates governance and scalability in cloud / edge environments because it is easier to manage a single model rather than a single model per turbine. On the other hand, having one single classifier 900 allows us to train it with a wider range of normal and fault data in different conditions, making the model more robust.

[0069] The classification process is based on a classifier 900 developed using supervised machine learning techniques. Non-limiting examples of machine learning techniques on which the classifier 900 is based can be k Nearest Neighbor (kNN), Random Forest (RF), Support Vector Machine (SVM), Extreme Learning Machine (ELM), among others. Due to its supervised and machine learning nature, the classifier 900 needs to be pre-trained before its further use for classifying normal or fault conditions. The training is performed using labeled data represented by condition indicators 8105 associated with different types of faults / anomalies and thus signaling possible anomalies. The classifier 900 is given a set of condition indicators 8105, each of which is associated with a label indicating "fault" or "no fault" and, in the case of a "fault" label, another label representing the type of fault / anomaly.

[0070] Thus far, the training of the classifier 900 has been disclosed. Once the classifier 900 is trained, it is used to detect faults in the operation of wind turbines of a wind energy farm (step 1000). This is shown schematically in FIG. 5B, which generally refers to the use of the classifier 900 as a fault detector when given condition indicators created from actual real-world data of the wind turbine (versus condition indicators created from historical data used to train the classifier 900). As shown in FIG. 5B, an actual condition indicator 18105 is obtained by comparing actual operating output data 1152 (e.g., SCADA data) with normal synthetic data 1280 obtained from the normal hybrid model 200 when it is given actual operating input data (real-world parameters) 1151. The comparison may involve different mathematical operations, such as difference, division, gradient, or others, applied to the compared data. The obtained condition indicator 18105 is provided to the trained classifier 900, which outputs a label indicating "fault" or "no fault." In case of a "fault" label, another label is output that describes the type of fault / anomaly. In case of a fault, an early warning 1050 is generated by the classifier 900.

[0071] The fault detection of step 1000 can be performed in real time, i.e. while the operational data (discrete values) are being collected, for example by a SCADA system. The operational data is collected in microseconds (1 μs=10 -6 The parameters are acquired at a particular acquisition frequency that may vary from every 100 μs or 1 ms, for example, every 10 seconds, to every 100 μs or 1 ms, to every 1 second or minute, for example, every 1 second, 1 minute, or 10 minutes. The acquisition frequency may depend on different circumstances, such as the type of parameters being acquired.

[0072] Therefore, the condition of the drive train of the wind turbine may be diagnosed using the classifier 900, which receives as input the condition indicators 18105 for different time intervals within the considered operational period. The diagnosis result may indicate efficient operation, inefficient operation, or a fault or abnormality. Depending on the diagnosis result, a message (e.g., an alarm) 1050 may be provided and then action may be taken on the drive train of the wind turbine.

[0073] The proposed method optimizes the operation of a power conversion drivetrain of a wind turbine and improves the efficiency of drivetrain operation and / or drivetrain maintenance. The different models on which the method is based or used, as well as the training of the models and the execution of the trained models, are preferably implemented within an IoT digital architecture, e.g., in the cloud, although it could alternatively be done locally. The different models are fed with field data (real-world operating data) within a specific operational period, including historical data and actual data.

[0074] The method is implemented in one or more processing devices, which may be part of a processing system. The one or more processing devices may comprise at least one memory, which may contain instructions, for example in the form of computer program code, which, when executed, perform the method of the present invention. The processing device may comprise a communications module, at least in embodiments in which it is communicatively coupled to other processing devices.

[0075] The proposed approach was validated using real-world operating data from a doubly-fed induction generator wind farm with a rated power of 1.5 MW. The synthetic data generation methodology was applied to a specific failure mode, namely, stator winding overtemperature. The overall methodology and the obtained results are summarized below.

[0076] Use case: Application of the method to a 1.5MW DFIG wind turbine The method was applied and validated using real SCADA data from an operational wind turbine. The wind turbine drivetrain includes a 1.5 MW DFIG and its corresponding back-to-back power converters. Real operational data collected over a three-year period was imported, organized, and preprocessed. The collected real operational data included wind speed, blade pitch angle, mechanical torque, active power, phase current, stator winding temperature, and nacelle temperature, among others. The available data must be known in detail because it influences the specific model topology. For example, if measurements corresponding to mechanical torque 131 are available, a single module (the second module 102) can be used to model the physics-based model 100. If mechanical torque is not available, the complete physics-based model (module 101 + module 102) must be used.

[0077] During data exploration and preprocessing of the SCADA data, relationships between parameters represented in the collected real-world operating data were analyzed to detect possible outliers. Non-limiting examples of such relationships are wind speed versus generated power and current versus nacelle temperature, among others. For example, in the power curve, anomalies corresponding to outliers were observed and removed to avoid confusing the model. These anomalies were observed to correspond, for example, to periods when there was no power generation despite the presence of a wind source, or periods when the generator had high temperatures in the stator windings while not producing power.

[0078] Once the initial data analysis was performed, a physical model 100 of the wind turbine power conversion drivetrain was developed in Simulink-Matlab R2020b.

[0079] FIG. 6 shows a representation of this model, consisting of a first module 101 and a second module 102. Design information for both the generator and the power converter was used as a starting point to build the model. However, some parameters were calculated or estimated in the absence of this information. Wind speed and pitch angle are input parameters required to operate the physical model 100. In the illustrated case, the output of the model 100 is generated power and related parameters (generated output current, generated output voltage, generated speed, rotor angle, and electromagnetic torque).

[0080] In the first conversion (first module 101), the air density, the rotor wind, the wind speed, and the power coefficient determine the mechanical power extracted from the kinetic energy in the wind. The mechanical power extracted is:

number

[0081] In the second transformation (second module 102), the DFIG block implements a three-phase wound rotor asynchronous machine operating in generator mode. The electrical part of the machine is represented by a fourth-order state-space model, and the mechanical part by a second-order system. All electrical variables and parameters are referenced to the stator, indicated by primes in the following mechanical equations: All stator and rotor quantities are in an arbitrary two-axis dq coordinate system (Table 1). [Table 1]

[0082] The parameters involved in solving the DFIG transformation equation are those indicated in Table 2. [Table 2]

[0083] Other parameters associated with the drivetrain main shafts also typically play a role in the shifting, non-limiting examples of which are the wind turbine inertia constant, shaft spring constant, shaft cross damping, and pitch angle.

[0084] The initial parameters of the physics-based model 100 are generalizations of the true parameters involved in the operation of a given turbine (e.g., stator resistance and leakage inductance, rotor resistance and leakage inductance, magnetizing inductance, total stator inductance, total rotor inductance, friction coefficient, inertia constant). The true values of these parameters can be adjusted using an optimization algorithm (Figure 1). The algorithm aims to find a combination of parameter values that minimizes the difference between the output of the physics-based model 100 and the measured SCADA data. In this case, the parameters were adjusted (or calibrated) using a surrogate optimization algorithm (surrogateopt) in Matlab (Surrogate Optimization Algorithm - MATLAB & Simulink (https: / / es.mathworks.com / help / gads / surrogate-optimization-algorithm.html)). This optimization algorithm is a global solver specifically designed for cases where the objective function is computationally expensive. The algorithm considers multivariate input variables x, which are subject to linear and nonlinear constraints, as well as a cost function with several finite bounds.

number

number

number

[0085] In the case of a wind turbine drivetrain, the objective function is defined as the mean absolute percentage error (MAPE) between the active power estimated by the physics-based model 100 and the active power measured by the SCADA system.

number

[0086] The optimization process in the illustrated case involves 13 parameters: four associated with the electromechanical conversion, three related to the aerodynamic conversion, three parameters of the control strategy, and finally, three parameters associated with the mechanical drivetrain (Table 3). [Table 3]

[0087] Calibration of the physics-based model was performed in two steps: in the first step, six variables were considered, while in the second step, seven further variables were added. The table shows both the initial values defined for each parameter (design values) and the values adopted after the second calibration step (calibration values). [Table 4]

[0088] New values for the calibrated parameters are established, always maintaining their physical meaning. In fact, intervals defining the lower and upper tolerance limits for each parameter have been previously established. The resulting model is an enhanced physics-based model 200.

[0089] As a result, the mean absolute percentage error (MAPE) between the real active power measured in the SCADA system and the value obtained in the simulation using the calibrated model (Model 200) improved from 15% to 2.4%, as shown in Figure 7, which shows active power (kW) versus wind speed (m / s).

[0090] Once the physical model 100 was calibrated, an additional model 300 (in this case, a thermal model) was added to the already developed normal hybrid model (trained physics-based model) 200 in Simulink to estimate the temperature within each phase of the stator winding, resulting in a fault model. Figure 8 shows the thermal model scheme. It must be taken into account that the insulation type of the stator winding is type F, i.e., it is designed to support 155°C. As shown in Figure 8, the thermal circuit considers the heat transfer generated by the stator currents, taking into account the following phenomena: conduction (conduction between the windings of each of the three stator phases) and convection (convection between the windings of each of the three stator phases, convection between each stator winding and the environment, and convection between each stator winding and the rotor). Radiation was ignored.

[0091] The Conduction Heat Transfer Block models heat transfer in a thermal network by conduction through layers of material. The rate of heat transfer is governed by Fourier's law (18) and is proportional to the temperature difference, the thermal conductivity of the material, the area perpendicular to the direction of heat flow, and inversely proportional to the layer thickness:

number

[0092] The inputs to the thermal model are the stator current and the room temperature where the generator is installed (in this case, the temperature of the nacelle), while the outputs are the temperatures of each phase of the stator winding.

[0093] Following the same process used in the normal hybrid model, this fault model is trained using operationally realistic SCADA data (Figure 3C). Similarly, training consists of optimizing the values of certain independent design parameters that represent the fault, with accurate values estimated for a given realistic interval. In this case, these parameters are the material thermal conductivity, the area perpendicular to the direction of heat flow, the thickness of the layer where conduction occurs, the heat transfer coefficient, and the surface area in contact with air. After this training, the fault hybrid model is generated.

[0094] The real-world data available during the analysis and implementation of this use case included five anomaly cases labeled as overtemperatures in the stator windings, as shown in Figures 9A, 9B, 9C, and 9D, which show wind speed (m / s) and active power (kW) at different time intervals. The data during these five anomaly cases was used to validate the fault modeling, yielding the estimated stator winding temperature results shown in Figures 10A–10E, which were compared with the actual SCADA winding temperatures. The MAPE between the actual stator winding temperature values measured in the SCADA system and those obtained in the simulation using the fault hybrid model had a value of 11%, with a maximum percent error of 16% in the worst-case scenario. This value could still be improved if more accurate design data were available for the thermal model.

[0095] Figure 11 shows a schematic diagram of a methodology for synthetic fault data generation applied to a generator stator winding overheating scenario formed with a data-driven probabilistic fault model and a hybrid fault model. Synthetic inputs associated with the faults are randomly generated using learned statistical distributions as shown in Figure 11 and provided as inputs into the developed hybrid fault model 500 (Figures 4A-4B). The hybrid fault model 500 generates the remaining portion of the synthetic fault data (e.g., stator winding temperature).

[0096] Figure 12 shows both the stator winding temperature values calculated by the fault hybrid model, taking the synthetically generated ones (ambient temperature, wind speed, nacelle temperature) as input variables (non-bold lines), and the real values of the stator windings measured by the SCADA system (bold lines).

[0097] A condition indicator was created as the difference between the stator winding temperature calculated using the fault hybrid model and the actual stator winding temperature measured in the SCADA, which allows for classification to detect real cooling problems, which will trigger an alarm. This is a simple rule-based classifier. In addition, benchmarking analyses were performed with other more complex classifiers using multiple features (e.g., wind speed, torque, stator winding temperature difference, nacelle temperature difference, etc.) and different algorithms such as KNN and random forest.

[0098] In summary, a hybrid model-based digital twin has been implemented that combines the advantages of physics-based models with advanced data analysis techniques. The proposed digital twin works very efficiently to control / supervise the operation of the wind turbine's power conversion drivetrain, improve (optimize) the efficiency of the drivetrain operation, and / or improve drivetrain maintenance.

[0099] Among the advantages of the proposed method based on a digital twin are: on the one hand, a procedure for generating synthetic fault data based on real data using different statistical techniques; and, on the other hand, a classification model for detecting anomalies in the operation of wind turbine drivetrains. These two innovations make it possible to overcome the main limitations of current digital twin approaches, which are related to accuracy, explainability, and the lack of sufficient training data.

[0100] In this document, the term "comprises" and its derivatives (e.g., "comprising") should not be understood in an exclusive sense, i.e., these terms should not be interpreted as excluding the possibility that what is being described and defined may include additional elements, steps, etc. In the context of the present disclosure, the term "approximately" and its cognate terms (e.g., "approximate," etc.) should not be understood as indicating a value in the immediate vicinity of that which the term accompanies. That is, deviations from the exact value within reasonable limits should be allowed, as one skilled in the art would understand that such deviations from the indicated value are unavoidable due to measurement imprecision, etc. The same applies to the terms "about," "around," and "substantially."

[0101] The present invention is obviously not limited to the particular embodiment(s) described herein, but also encompasses any modifications that may be considered by any person skilled in the art (e.g., with regard to the selection of materials, dimensions, components, configurations, etc.) to be within the general scope of the invention as defined in the claims.

Claims

1. 1. A computer-implemented method for training a classifier for fault diagnosis of a drive train of a wind turbine, comprising: training a physics-based model (100) of a wind turbine drivetrain using operational data (151, 152) of the drivetrain under normal conditions, thereby providing a normal hybrid model (200) of the drivetrain of the wind turbine; - modelling at least one abnormal situation in the normal hybrid model (200) and training it using operational data (151, 152') of the drivetrain under fault conditions, thus providing a fault hybrid model (500) of the drivetrain of the wind turbine; applying a set of input data (151) of the drivetrain under normal conditions to the normal hybrid model (200), thereby obtaining a set of synthetic data (280) under normal conditions for specific parameters of the drivetrain; applying a set of fault data (150'') associated with the at least one abnormal situation of the drivetrain and a set of synthetic input data (710) to the fault hybrid model (500), thereby generating a set of synthetic fault data (750) for the at least one abnormal situation; obtaining a set of condition indicators (8105) for the at least one abnormal situation from the set of drivetrain output data (152, 152'), the obtained set of composite data (280) under normal conditions, and the set of composite fault data (750) for the at least one abnormal situation; training a classifier (900) using the set of condition indicators (8105) for the at least one abnormal situation; 10. A computer-implemented method comprising:

2. 2. The method of claim 1, wherein the physics-based model (100) of the drivetrain of the wind turbine includes a first module (101) representing an aerodynamic conversion stage of the drivetrain and a second module (102) representing an electromechanical conversion stage of the drivetrain.

3. 3. The method of claim 1 or 2, wherein the normal hybrid model (200) is: providing a physics-based model (100) of the drive train of the wind turbine, the physics-based model (100) being represented by a set of model parameters and equations that describe energy conversion within the drive train of the wind turbine; providing the physics-based model (100) with operational data (151, 152) of the drivetrain under normal conditions, the operational data including real operational data inputs (151) and real operational data outputs (152); Obtaining a prediction (180) of at least one of the following parameters: generated power, phase current, phase voltage, and electromagnetic torque; selecting values for at least some of the model parameters, said values minimizing the difference between the prediction (180) corresponding to the actual motion data input (151) and the actual motion data output (152), thus obtaining the physics-based model (200) trained using motion data; The method is characterized in that it is obtained by

4. 4. The method according to any one of claims 1 to 3, wherein the impairment hybrid model (500) is: modeling at least one abnormal situation in said normal hybrid model (200); providing operational data (151, 152') of the drivetrain under fault conditions to the normal hybrid model (200) including at least one modeled abnormal situation; obtaining a prediction (180') of a parameter indicative of a fault condition; selecting values of at least some of the model parameters defining the normal hybrid model (200) modeled using at least one abnormal situation, the values minimizing the difference between the prediction (180') corresponding to the actual operating data input (151) and the actual operating data output (152'), thereby obtaining the calibrated fault hybrid model (500); The method is characterized in that it is obtained by

5. 5. The method of claim 4, wherein said at least one abnormal situation is modeled using at least one fault model (301-306).

6. 6. The method of claim 5, wherein the at least one fault model (301-306) is represented by a set of fault parameters and equations describing faults in the energy conversion of the drive train of the wind turbine.

7. 7. A computer-implemented method for optimizing the operation of a drivetrain of a wind turbine, comprising using a classifier (900) as trained in any one of claims 1 to 6.

8. 8. The computer-implemented method of claim 7, Obtaining a state indicator (18105) from actual motion output data (1152) and normal synthetic data (1280) obtained from the normal hybrid model (200) when the normal hybrid model (200) is given actual motion input data (1151); providing said state indicator (18105) to said trained classifier (900); outputting, by the trained classifier (900), whether the condition indicator (18105) corresponds to a fault condition or a non-fault condition; 10. A computer-implemented method comprising:

9. 8. The computer-implemented method of claim 7, further comprising: The computer-implemented method, further comprising triggering an alert when the classifier (900) determines that the drivetrain is operating in an abnormal condition.

10. 10. The computer-implemented method of claim 8 or 9, further comprising taking action on the drive train of the wind turbine when the condition indicator (18105) corresponds to a fault condition.

11. A device comprising processing means adapted to perform the method according to any one of claims 1 to 6.

12. A device comprising processing means adapted to perform the method of any one of claims 7 to 10.

13. A computer program product comprising computer program instructions / code for performing the method according to any one of claims 1 to 6 and / or the method according to any one of claims 7 to 10.

14. A computer readable memory / medium storing program instructions / code for performing the method according to any one of claims 1 to 6 and / or the method according to any one of claims 7 to 10.