Methods, devices, equipment and media for identifying anomalies in wind turbine generator sets

By using deep learning models to identify anomalies in wind turbine generators in real time, the problem of poor timeliness in identifying anomalies in wind turbine generators has been solved, thus improving safety.

CN118686748BActive Publication Date: 2026-05-26BEIJING GOLDWIND SCI & CREATION WINDPOWER EQUIP CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING GOLDWIND SCI & CREATION WINDPOWER EQUIP CO LTD
Filing Date
2023-03-21
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Wind turbine generators may experience abnormalities during operation, leading to safety risks. Existing technologies struggle to identify and address these abnormalities in a timely manner, resulting in poor timeliness.

Method used

By combining deep learning models with environmental factors and operational status data, and through a pre-trained anomaly recognition model, abnormal components in wind turbine generators can be identified in real time. The input data is processed and predicted using a deep learning recurrent layer.

Benefits of technology

This improves the timeliness of anomaly identification in wind turbine generators, enabling timely detection and handling of potential safety hazards, and enhancing the safety of wind turbine generators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118686748B_ABST
    Figure CN118686748B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, equipment, and medium for anomaly identification of wind turbine generator sets, belonging to the field of wind power generation. The method includes: acquiring environmental factor data, first-type operating state data, and second-type operating state data of the wind turbine generator set within the current time period; obtaining input data based on the environmental factor data, first-type operating state data, and second-type operating state data within the current time period; inputting the input data into a pre-trained anomaly identification model to obtain predicted data output by the anomaly identification model, wherein the anomaly identification model includes at least one deep learning recurrent layer, iteratively trained based on sample data, and the sample data is obtained based on historical data of the wind turbine generator set with a specificity higher than or equal to a preset specificity; and determining that the component corresponding to the second-type operating state data has an anomaly when the predicted data meets preset anomaly conditions. The embodiments of this application can improve the safety of wind turbine generator sets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of wind power generation, and in particular relates to a method, device, equipment and medium for identifying anomalies in wind turbine generator sets. Background Technology

[0002] A wind turbine is a large device that converts wind energy into electrical energy, consisting of numerous components. These components may malfunction due to prolonged use or inherent defects, leading to abnormalities during operation. Ignoring these abnormalities and continuing operation poses a significant safety risk to the wind turbine.

[0003] On the one hand, wind turbines only generate anomaly files some time after an anomaly occurs, and these files, containing abnormal data, are uploaded to the analysis equipment, meaning the abnormal data cannot be returned in a timely manner. On the other hand, as wind turbine technology continues to improve, the amount of data generated by wind turbines is also increasing. This increased data volume makes it more difficult to identify anomalies in wind turbines, and the identification process takes a significant amount of time, resulting in poor timeliness and reduced safety of wind turbines. Summary of the Invention

[0004] This application provides a method, apparatus, equipment, and medium for identifying anomalies in wind turbine generator sets, which can improve the safety of wind turbine generator sets.

[0005] In a first aspect, embodiments of this application provide a method for identifying anomalies in wind turbine generator sets, comprising: acquiring environmental factor data, first-type operating state data, and second-type operating state data of the wind turbine generator set within the current time period; obtaining input data based on the environmental factor data, first-type operating state data, and second-type operating state data within the current time period; inputting the input data into a pre-trained anomaly identification model to obtain predicted data output by the anomaly identification model, wherein the anomaly identification model includes at least one deep learning recurrent layer, iteratively trained based on sample data, the sample data being obtained based on historical data of the wind turbine generator set whose specificity is higher than or equal to a preset specificity, the historical data including environmental factor data, first-type operating state data, and second-type operating state data within a historical time period; and determining that an anomaly has occurred in a component of the wind turbine generator set corresponding to the second-type operating state data when the predicted data meets preset anomaly conditions.

[0006] Secondly, embodiments of this application provide a wind turbine generator set anomaly identification device, comprising: an acquisition module for acquiring environmental factor data, first type of operating state data, and second type of operating state data of the wind turbine generator set within the current time period; a conversion module for obtaining input data based on the environmental factor data, first type of operating state data, and second type of operating state data within the current time period; a processing module for inputting the input data into a pre-trained anomaly identification model to obtain predicted data output by the anomaly identification model, wherein the anomaly identification model includes at least one deep learning recurrent layer, iteratively trained based on sample data, the sample data being based on historical data of the wind turbine generator set with a specificity higher than a preset specificity, the historical data including environmental factor data, first type of operating state data, and second type of operating state data within a historical time period; and an anomaly determination module for determining that a component corresponding to the second type of operating state data has an anomaly when the predicted data meets preset anomaly conditions.

[0007] Thirdly, embodiments of this application provide an electronic device, including: a processor and a memory storing computer program instructions; the processor executes the computer program instructions to implement the wind turbine generator abnormality identification method of the first aspect.

[0008] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the wind turbine generator set anomaly identification method of the first aspect.

[0009] This application provides a method, apparatus, device, and medium for anomaly identification of wind turbine generator sets. An anomaly identification model can be pre-trained, including at least one deep learning recurrent layer, constituting a deep learning model. The deep learning model is iteratively trained based on sample data using deep learning. The sample data is obtained from environmental factor data, first-type operating state data, and second-type operating state data within a historical time period with specific characteristics of the wind turbine generator set exceeding a preset specificity. This allows the anomaly identification model to quickly output predictive data corresponding to the input data, indicating whether an anomaly has occurred in a component of the wind turbine generator set corresponding to the second-type operating state data. Inputting the input data, obtained from real-time acquired environmental factor data, first-type operating state data, and second-type operating state data of the wind turbine generator set in the current time period, into the anomaly identification model enables it to output predictive data promptly. Based on the predictive data and preset anomaly conditions, it can promptly determine whether an anomaly has occurred in a component of the wind turbine generator set corresponding to the second-type operating state data, improving the timeliness of anomaly identification and enhancing the safety of the wind turbine generator set. Attached Figure Description

[0010] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 A flowchart illustrating a wind turbine generator set anomaly identification method provided in an embodiment of this application;

[0012] Figure 2 A simplified architecture diagram of an example of an anomaly recognition model provided in an embodiment of this application;

[0013] Figure 3 A flowchart of a wind turbine generator set anomaly identification method provided in another embodiment of this application;

[0014] Figure 4 A flowchart of a wind turbine generator set anomaly identification method provided in another embodiment of this application;

[0015] Figure 5 A schematic diagram illustrating an example of the training loop of the anomaly recognition model provided in an embodiment of this application;

[0016] Figure 6 A schematic diagram illustrating an example of the loss function value curve under 20 iterations provided in an embodiment of this application;

[0017] Figure 7 A schematic diagram illustrating another example of the loss function value curve under 20 iterations provided in this application embodiment;

[0018] Figure 8 A schematic diagram illustrating an example of the loss function value curve under 30 iterations provided in an embodiment of this application;

[0019] Figure 9 A schematic diagram illustrating another example of the loss function value curve under 30 iterations provided in an embodiment of this application;

[0020] Figure 10 A schematic diagram illustrating an example of the loss function value curve after 200 iterations, provided in an embodiment of this application;

[0021] Figure 11 A schematic diagram illustrating an example of the loss function value curve and mean absolute error curve under 200 iterations provided in an embodiment of this application;

[0022] Figure 12 A schematic diagram illustrating an example of predicted data and real data from the anomaly detection model provided in this application embodiment;

[0023] Figure 13A flowchart of a wind turbine generator set anomaly identification method provided in another embodiment of this application;

[0024] Figure 14 A box plot showing the distribution of vibration data from multiple wind turbine generator sets provided in an embodiment of this application;

[0025] Figure 15 A schematic diagram illustrating an example of vibration data from multiple wind turbine generator sets provided in an embodiment of this application;

[0026] Figure 16 A comparative schematic diagram illustrating an example of the cumulative values ​​of specific parameters of multiple wind turbine generator sets provided in the embodiments of this application;

[0027] Figure 17 This is a schematic diagram of the structure of a wind turbine generator abnormality identification device provided in an embodiment of this application;

[0028] Figure 18 This is a schematic diagram of the structure of a wind turbine generator abnormality identification device provided in another embodiment of this application;

[0029] Figure 19 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0030] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.

[0031] A wind turbine is a large device that converts wind energy into electrical energy, composed of numerous components. These components may malfunction due to prolonged use or inherent defects, leading to operational anomalies. Ignoring these anomalies and continuing operation poses a significant safety risk to the wind turbine. Firstly, the wind turbine only generates anomaly files some time after an anomaly occurs, uploading these files containing abnormal data to analysis equipment, resulting in delayed data feedback. Secondly, with continuous advancements in wind turbine technology, the amount of data generated is increasing dramatically. This increased data volume makes anomaly identification more difficult and time-consuming, leading to poor timeliness in anomaly detection and ultimately reducing the wind turbine's safety.

[0032] This application provides a method, apparatus, device, and medium for identifying anomalies in wind turbine generator sets. Utilizing a combination of deep learning, machine learning, and data mining techniques, and based on relevant data from the wind turbine generator set, it can promptly identify anomalies. It can acquire relevant data from the wind turbine generator set in real time, and use this real-time data and a pre-trained deep learning model to obtain predicted data output by the deep learning model based on the relevant data. The predicted data is then used to determine whether any components in the wind turbine generator set have malfunctioned. The deep learning model includes at least one deep learning recurrent layer, and the sample data used to train the deep learning model is based on historical data with high specificity to the wind turbine generator set. The trained deep learning model can quickly output predicted data based on the real-time acquired data from the wind turbine generator set, thereby enabling timely identification of anomalies and improving the safety of the wind turbine generator set.

[0033] The following describes the wind turbine generator set anomaly identification method, device, equipment, and medium provided in this application.

[0034] The first aspect of this application provides a method for identifying anomalies in wind turbine generator sets, which can be applied to scenarios involving the identification of anomalies in components of wind turbine generator sets. Specifically, it can be executed by a wind turbine generator set anomaly identification device, electronic equipment, etc., and is not limited thereto. Figure 1 A flowchart of a wind turbine generator anomaly identification method provided in an embodiment of this application is shown below. Figure 1 As shown, the wind turbine generator abnormality identification method may include steps S101 and S104.

[0035] In step S101, environmental factor data, first type of operating status data and second type of operating status data of the wind turbine generator set within the current time period are obtained.

[0036] The current time period refers to the time period to be measured. The length of this period can be set based on the scenario, requirements, experience, etc., and is not limited here. Environmental factor data, Type I operating status data, and Type II operating status data for a single wind turbine generator set within the current time period can be acquired. Alternatively, data from multiple wind turbine generator sets within the current time period can also be acquired. The environmental factor data, Type I operating status data, and Type II operating status data within the current time period can be relevant data obtained in real-time from the wind turbine generator sets. Environmental factor data, Type I operating status data, and Type II operating status data can be collected according to a data acquisition cycle. Specifically, this data can be obtained from data acquisition devices such as sensors related to these data, for example, temperature sensors, wind sensors, and speed sensors installed on the wind turbine generator sets. The data acquisition cycle can be set based on the scenario, requirements, experience, etc., and is not limited here. For example, the data acquisition cycle can be 3 to 7 seconds. Correspondingly, the data volume of a single wind turbine generator set can reach approximately 12,000 data points per day, which is a very large amount of data.

[0037] Environmental factor data includes data characterizing the environment in which the wind turbine operates, and is not limited herein. For example, environmental factor data may include wind speed data, ambient temperature data, etc. The second type of operational status data includes operational status data associated with the component under test (SUT), which includes components of the wind turbine being assessed for abnormalities, corresponding to the second type of operational status data. The types of SUT and the second type of operational status data are not limited herein. For example, if the SUT includes the main bearing of the wind turbine, the second type of operational status data may include vibration data. As another example, if the SUT includes the blades of the wind turbine, the second type of operational status data may include pitch angle data or other blade-related data. The first type of operational status data may include at least some of the other operational status data of the wind turbine besides the second type of operational status data, and is not limited herein. For example, the first type of operational status data may include the power data and operating condition data of the wind turbine.

[0038] In step S102, input data is obtained based on environmental factor data, first type of operating status data and second type of operating status data within the current time period.

[0039] It can process environmental factor data, first-type operational status data, and second-type operational status data within the current time period, converting them into input data suitable for the anomaly detection model. The input data is used to power the anomaly detection model and has format requirements adapted to the model's input requirements. For example, if the anomaly detection model requires tensor data, then the environmental factor data, first-type operational status data, and second-type operational status data within the current time period need to be converted into tensor data to obtain the input data. Tensor data can be implemented in vector, matrix, or other formats.

[0040] In some examples, environmental factor data, first-type operational status data, and second-type operational status data within the current time period can be preprocessed to obtain more accurate data that is easier to process. For example, environmental factor data, first-type operational status data, and second-type operational status data can be divided into multiple groups to achieve data binning; the time data corresponding to environmental factor data, first-type operational status data, and second-type operational status data can be converted to a specified date format; data with abnormal formats in environmental factor data, first-type operational status data, and second-type operational status data can be removed; duplicate and empty data in environmental factor data, first-type operational status data, and second-type operational status data can be removed; noisy data, abnormal data, and non-key data in environmental factor data, first-type operational status data, and second-type operational status data can be removed; when it is necessary to add other data within the current time period, variables including environmental factor data, first-type operational status data, and second-type operational status data can be reconstructed; and numerical standardization and centralization processing can be performed on environmental factor data, first-type operational status data, and second-type operational status data.

[0041] In step S103, the input data is fed into the pre-trained anomaly detection model to obtain the predicted data output by the anomaly detection model.

[0042] The anomaly detection model includes at least one deep learning recurrent layer. A deep learning recurrent layer can be viewed as a data processing module; it takes one or more tensor data points as input and outputs one or more tensor data points. Different deep learning recurrent layers are suitable for different tensor data formats and different types of data processing. Each deep learning recurrent layer only accepts input tensor data in a specified format and returns output tensor data in a specified format. For example, sequence data stored in a three-dimensional tensor with the format (sample, time, feature) can be processed by a deep learning recurrent layer acting as a reproduction layer.

[0043] The anomaly detection model is trained iteratively based on sample data. The sample data is derived from historical data of wind turbine generators with a specificity equal to or higher than a preset specificity. Specificity reflects the degree of uniqueness of environmental factor data, first-type operating state data, and second-type operating state data of wind turbine generators over a period of time, and / or the degree of uniqueness of these data across multiple wind turbine generators. Higher specificity indicates more representative data and a greater role in the anomaly detection model. The preset specificity is used to filter historical data and can be set according to scenarios, needs, experience, etc., but is not limited here. Historical data with a specificity equal to or higher than the preset specificity can be used to obtain sample data; historical data with a specificity lower than the preset specificity can be discarded and will not participate in the training of the anomaly detection model.

[0044] Historical data may include environmental factor data, Category I operating status data, and Category II operating status data within a historical time period. Historical data includes historical data from normal wind turbine generators and historical data from abnormal wind turbine generators. Historical data, as well as the environmental factor data, Category I operating status data, and Category II operating status data within the current time period, can be obtained from the Supervisory Control and Data Acquisition (SCADA) system of the wind turbine generator. The historical time period is prior to the current time period; that is, the historical time period is a period preceding the current time period. The specific content of the environmental factor data, Category I operating status data, and Category II operating status data can be found in the relevant descriptions in the above embodiments, and will not be repeated here. Since a single wind turbine generator can generate a large amount of data per day, the volume of historical data may reach tens of millions of records.

[0045] In some examples, historical data can be preprocessed to obtain sample data. Preprocessing includes tensor transformation, that is, converting historical data into tensor data, and using the transformed tensor data as sample data. Tensor data is data obtained by generalizing vectors and matrices to any number of dimensions. Vectors can be used for one-dimensional tensors, matrices for two-dimensional tensors, and array objects can be used for higher dimensions, which is not limited here. Tensor data is suitable as a basic data structure for model training. During model training, tensor data can be transformed, reset, and processed. Other preprocessing can also be performed to improve the accuracy and applicability of sample data. For example, preprocessing can also include one or more of the following methods: data binning, time data format conversion, outlier data removal, duplicate data removal, empty data removal, numerical outlier data removal, non-critical data removal, variable reconstruction, data standardization, data centralization, etc., which is not limited here. Numerical outlier data removal can use the 3σ principle to determine outliers, which is not limited here.

[0046] The number of deep learning recurrent layers in an anomaly detection model can be determined based on a balance between the computational cost and representational power of the model. Deep learning recurrent layers may include Long Short-Term Memory (LSTM) deep learning recurrent layers, Gated Recurrent Unit (GRU) deep learning recurrent layers, and / or bidirectional recurrent neural network (NRNN) deep learning recurrent layers. When the anomaly detection model includes more than two deep learning recurrent layers, the output of the previous deep learning recurrent layer can be used as the input to the next. Each deep learning recurrent layer has weights, which are related to the tensor data input to the deep learning recurrent layer. The learning of the tensor data in the deep learning recurrent layer can be achieved using stochastic gradient descent. In some examples, the input data to the deep learning recurrent layer has weights, and the weights of the deep learning recurrent layer are related to the weights of the input data. The weights of the deep learning recurrent layer can be adjusted by adjusting the weights of the input data. In some examples, when the anomaly detection model includes more than two deep learning recurrent layers, some deep learning recurrent layers may be LSTM deep learning recurrent layers, and some may be GRU deep learning recurrent layers. For example, Figure 2 A simplified architectural diagram illustrating an example of the anomaly recognition model provided in this application embodiment, as shown below. Figure 2As shown, the anomaly detection model includes a long short-term memory artificial neural network deep learning recurrent layer and a gated recurrent unit deep learning recurrent layer. The output of the long short-term memory artificial neural network deep learning recurrent layer can be used as the input of the gated recurrent unit deep learning recurrent layer. The input data in each deep learning recurrent layer can be configured with weights. The deep learning recurrent layer can be optimized by adjusting the weights of the input data of the deep learning recurrent layer, thereby optimizing the anomaly detection model.

[0047] The output data of the anomaly identification model includes predicted data corresponding to the input data. The predicted data is data obtained by the anomaly identification model based on environmental factor data, first-type operating state data and second-type operating state data within the current time period, which can reflect whether the components in the wind turbine generator corresponding to the second-type operating state data are abnormal.

[0048] In step S104, if the predicted data meets the preset abnormal conditions, it is determined that the component in the wind turbine generator set corresponding to the second type of operating status data has an abnormality.

[0049] Whether a component in a wind turbine generator corresponding to the second type of operating state data is abnormal can be determined by whether the predicted data meets preset abnormality conditions. The component in the wind turbine generator corresponding to the second type of operating state data is the component to be tested in the above embodiment. The type of the second type of operating state data and the type of component corresponding to the second type of operating state data are not limited here. For example, if the second type of operating state data includes vibration data and the component corresponding to the second type of operating state data includes the main bearing of the wind turbine generator, then this wind turbine generator abnormality identification method can be applied to the scenario of identifying abnormalities in the main bearing of the wind turbine generator. As another example, if the second type of operating state data includes pitch angle data and the component corresponding to the second type of operating state data includes the blades of the wind turbine generator, then this wind turbine generator abnormality identification method can be applied to the scenario of identifying abnormalities in the blades of the wind turbine generator.

[0050] The preset anomaly conditions can be set according to the scenario, requirements, type of predicted data, experience, etc., and are not limited here. If the predicted data meets the preset anomaly conditions, it means that the component in the wind turbine corresponding to the second type of operating state data has an anomaly; if the predicted data does not meet the preset anomaly conditions, it means that the component in the wind turbine corresponding to the second type of operating state data has not anomaly.

[0051] In some examples, an early warning signal can be issued when an anomaly is detected in a component of the wind turbine corresponding to the second type of operating status data. This early warning signal indicates the anomaly in the wind turbine, and in some examples, it can also indicate the location of the anomaly within the wind turbine, i.e., the location corresponding to the second type of operating status data. The early warning signal can be an image signal, an audible signal, an indicator light signal, etc., and its form is not limited here.

[0052] In some examples, the predicted data may include predicted anomaly probabilities and / or predicted operational status data. Predicted anomaly probabilities include the probability of an anomaly occurring for the component corresponding to the second type of operational status data, predicted by the anomaly identification model based on environmental factor data, first type of operational status data, and second type of operational status data within the current time period. Predicted operational status data includes the expected second type of operational status data for the component corresponding to the second type of operational status data within the current time period, assuming no anomalies have occurred, predicted by the anomaly identification model based on environmental factor data, first type of operational status data, and second type of operational status data within the current time period.

[0053] Preset anomaly conditions may include: when the predicted data includes a predicted anomaly probability, the predicted anomaly probability is greater than or equal to a preset anomaly probability threshold; when the predicted data includes predicted operating status data, the difference between the predicted operating status data and the second type of operating status data in the current time period exceeds a preset requirement range.

[0054] The preset anomaly probability threshold can be a boundary value for the anomaly probability and can be set according to the scenario, requirements, experience, etc., and is not limited here. The predicted data includes the predicted anomaly probability. If the predicted anomaly probability is greater than the preset anomaly probability threshold, it can be determined that the component in the wind turbine corresponding to the second type of operating state data has an anomaly. For example, if the preset anomaly probability threshold is 85%, and the predicted anomaly probability output by the anomaly identification model is greater than or equal to 85%, it can be determined that the component in the wind turbine corresponding to the second type of operating state data has an anomaly.

[0055] The preset requirement range is the acceptable range for the probability that the component has not experienced an anomaly. It can be set according to the scenario, requirements, experience, etc., and is not limited here. The predicted data includes predicted operating status data. If the difference between the predicted operating status data and the second type of operating status data in the current time period exceeds the preset requirement range, it means that the actual second type of operating status data differs significantly from the predicted second type of operating status data where the component has not experienced an anomaly. This indicates that the component in the wind turbine generator corresponding to the second type of operating status data has experienced an anomaly.

[0056] In this embodiment, an anomaly detection model can be pre-trained. This model includes at least one deep learning recurrent layer, constituting a deep learning model. The deep learning model is iteratively trained based on sample data using deep learning. The sample data is obtained from environmental factor data, first-type operating state data, and second-type operating state data within a historical time period where the wind turbine's specificity exceeds a preset specificity. This allows the anomaly detection model to quickly output predictive data corresponding to the input data, indicating whether any component in the wind turbine corresponding to the second-type operating state data has experienced an anomaly. Inputting the input data—based on real-time acquired environmental factor data, first-type operating state data, and second-type operating state data of the wind turbine in the current time period—into the anomaly detection model enables it to output predictive data promptly. Based on the predictive data and preset anomaly conditions, it can promptly determine whether any component in the wind turbine corresponding to the second-type operating state data has experienced an anomaly, improving the timeliness of anomaly detection and enhancing the safety of the wind turbine.

[0057] The anomaly detection model in this embodiment can be built using the Keras deep learning framework. It can run seamlessly on both Graphics Processing Units (GPUs) and Central Processing Units (CPUs) using the same code. Its user-friendly Application Programming Interface (API) also facilitates the rapid construction of deep learning models and supports built-in support for recurrent networks for sequence processing, convolutional networks for computer vision, and any combination of both, arbitrary network architectures, multi-input or multi-output models, layer sharing, and model sharing. The backend engine of the Keras deep learning framework can be a deep learning platform that provides tensor operation libraries; this is not limited to any particular platform.

[0058] Deep learning models offer simplicity, automating feature engineering and replacing complex, cumbersome feature processing with simple, end-to-end trainable models, thus simplifying anomaly identification in wind turbines. They are also scalable, compatible with both image processors (IPCs) and tensor processing units (TPUs), and can be parallelized. Deep learning models can be trained on multiple batches of sample data, with varying sample sizes. Furthermore, they are versatile and reusable, allowing for continuous online deep learning and reuse for different applications.

[0059] In some embodiments, an anomaly detection model can be pre-trained. Figure 3 A flowchart of a wind turbine generator anomaly identification method provided in another embodiment of this application is shown. Figure 3 and Figure 1 The difference is that, Figure 3 The wind turbine generator set anomaly identification method shown may also include steps S105 to S108.

[0060] In step S105, the acquired historical data is preprocessed to obtain sample data.

[0061] For details on the preprocessing of historical data, please refer to the relevant descriptions in the above embodiments, which will not be repeated here. Sample data obtained based on historical data may include training sample data. The training sample data is used to train the model.

[0062] In some examples, a sample data set comprises two parts: sample input data and sample target data. The sample target data is data of the same type as the prediction data. The model is trained using both the sample input data and the sample target data. Input data of the same type as the sample input data is then fed into the trained model, and the model outputs prediction data of the same type as the sample target data. A generator can be used to convert historical data into sample data. For example, if the prediction data includes predicting anomaly probabilities, then the portion of the sample data related to environmental factor data, first-type operational state data, and second-type operational state data constitutes the sample input data, and the label information representing abnormality or non-abnormality in the sample data constitutes the sample target data. If the prediction data includes predicting operational state data, then the portion of the sample data related to environmental factor data and first-type operational state data constitutes the sample input data, and the second-type operational state data under abnormal or non-abnormal conditions constitutes the sample target data.

[0063] Preprocessed historical data can be converted into floating-point matrices, and then a generator can be used to convert these matrices back into sample data. The generator needs to maintain its internal state, and this conversion can be achieved by calling another function that returns a generation function. The function implementing the generator can return NULL to indicate completion, but the generator function passed to model training can return an unlimited number of values. The number of times the generator function is called can be controlled by certain parameters in the deep learning training method, such as the `epochs` and `steps_per_epoch` parameters. The generator function can use parameters such as `data`, `lookback`, `delay`, `min_index`, `max_index`, `shuffle`, `batch_size`, and `step`, which are not limited here. Sample data includes training sample data and validation sample data, or possibly both validation and test validation sample data. Validation sample data is used to validate the trained model to determine if it meets the requirements. Test sample data is used to test the validated model to determine if it meets the requirements. Generators can correspond to the classification of sample data. For example, three generators can be instantiated using functions that implement generators. The first generator is used to transform the data to obtain training sample data, the second generator is used to transform the data to obtain validation sample data, and the third generator is used to transform the data to obtain test sample data. The three generators can transform historical data within different time periods into sample data. For example, the first generator can transform historical data from time 0 to t1 into training sample data, the second generator can transform historical data from time t1 to t2 into validation sample data, and the third generator can transform historical data from time t2 to t3 into test sample data.

[0064] In step S106, an initial model is established, and at least a portion of the training sample data is used to train the initial model to obtain an intermediate model.

[0065] In the initial model building process, it is difficult to build a suitable anomaly detection model in one go. An initial model can be built first; this initial model is a basic, iteratively built model that can be trained using at least a portion of the training sample data. The training sample data can be divided into multiple batches, and at least a portion of these batches can be used to train the initial model, resulting in an intermediate model. The intermediate model is the model obtained by training the initial model using at least a portion of the training sample data.

[0066] In step S107, if the effect parameters of the intermediate model meet the model effect requirements, an anomaly recognition model is obtained based on the intermediate model.

[0067] After training the intermediate model, its performance parameters can be obtained, which characterize the model's effectiveness. Model performance requirements are used to determine whether the model meets the performance requirements; these can be set based on the scenario, needs, experience, etc., and are not limited here. If the intermediate model's performance parameters meet the model performance requirements, it means the intermediate model meets the performance requirements, and the intermediate model can be used as an anomaly detection model.

[0068] In step S108, if the effect parameters of the intermediate model do not meet the model effect requirements, the intermediate model is optimized until the optimized intermediate model is obtained that meets the model effect requirements. Based on the optimized intermediate model, the anomaly recognition model is obtained.

[0069] If the intermediate model's performance parameters do not meet the required performance conditions, it means the intermediate model's performance does not meet the requirements. In this case, the intermediate model can be optimized. This optimization may change the model's parameters or structure, which is not limited here. After obtaining the optimized intermediate model, its performance parameters can be obtained, and it can be determined whether the optimized intermediate model's performance parameters meet the required performance conditions. If the optimized intermediate model's performance parameters meet the required performance conditions, it can be identified as an anomaly detection model. If the optimized intermediate model's performance parameters do not meet the required performance conditions, it can be optimized again, and so on, until the latest optimized intermediate model's performance parameters meet the required performance conditions. The optimized intermediate model whose performance parameters meet the required performance conditions is identified as an anomaly detection model. In other words, an anomaly detection model can be an intermediate model, or an optimized intermediate model that has undergone at least one optimization process and whose performance parameters meet the required performance conditions. The methods used in different optimization processes can be different or the same, which is not limited here.

[0070] In some examples, the initial model includes at least one deep learning recurrent layer. The optimization processes described above may include one or more of the following:

[0071] Using a preset random dropout probability, the input units in one or more specified deep learning recurrent layers are reset to zero;

[0072] Perform random deactivation regularization on one or more deep learning recurrent layers;

[0073] Stacked deep learning recurrent layers, including long short-term memory artificial neural network deep learning recurrent layers, gated recurrent unit deep learning recurrent layers and / or bidirectional recurrent neural network deep learning recurrent layers;

[0074] Adjust at least some of the model parameters of the deep learning recurrent layer.

[0075] In cases of overfitting of the intermediate model, such as when the training loss of the intermediate model decreases but the validation loss remains stable or increases, it indicates that the intermediate model is overfitting and fits the training sample data too precisely. However, the generalization ability of the intermediate model is greatly weakened, which is not conducive to the model's anomaly identification of large batches of data. According to the random deactivation probability, the input units in one or more deep learning recurrent layers can be randomly reset to zero to break the occasional correlation in the training sample data contacted by the deep learning recurrent layer.

[0076] Random deactivation regularization can also improve the ability of intermediate models to resist overfitting, allowing the optimized intermediate model to balance performance parameters such as training loss, validation loss, and mean absolute error (MAE). Training loss reflects the difference between the model's output corresponding to the training samples and the target data in the training samples. Validation loss reflects the difference between the model's output corresponding to the validation samples and the target data in the validation samples. Mean absolute error reflects the average absolute error between the model's output corresponding to the training samples and the target data in the training samples, and / or the average absolute error between the model's output corresponding to the validation samples and the target data in the validation samples.

[0077] Stacking deep learning recurrent layers can involve stacking one or more of the following: Long Short-Term Memory (LSTM) AINN deep learning recurrent layers, gated recurrent unit (GRU) deep learning recurrent layers, and bidirectional recurrent neural network (NRNN) deep learning recurrent layers; this is not limited to these two types. Stacking deep learning recurrent layers can increase model capacity. The combined use of random deactivation regularization and stacking GRU deep learning recurrent layers can also improve the intermediate model's ability to resist overfitting, allowing the optimized intermediate model to balance performance parameters such as training loss, validation loss, and mean absolute value. Bidirectional NRNN deep learning recurrent layers can be used for reverse training and evaluation of LSM AINN deep learning recurrent layers, thereby improving model accuracy.

[0078] Adjusting at least some of the model parameters of a deep learning recurrent layer can refer to adjusting the values ​​of at least some model parameters, or it can refer to adjusting at least some model parameters of one deep learning recurrent layer in an intermediate model to at least some model parameters of another deep learning recurrent layer. Adjusting at least some model parameters of one deep learning recurrent layer in an intermediate model to at least some model parameters of another deep learning recurrent layer is equivalent to changing the type of the deep learning recurrent layer.

[0079] The following functions and parameters can be used for the definition, compilation, training, and optimization processes in building an anomaly detection model, and are not limited here: [keras_model_sequential(), layer_gru, layer_lstm, layer_dense, [units, input_shape, dropout, recurrent_dropout]], [compile(), optimizer, loss, metrics], [fit_generator(), fit(), epochs, steps_per_epoch].

[0080] In some embodiments, the sample data further includes validation sample data, which is used to validate the trained model to determine whether the model's loss meets the requirements, thereby determining whether further training of the model is necessary. Figure 4 A flowchart illustrating a wind turbine generator anomaly identification method provided in another embodiment of this application. Figure 4 and Figure 3 The difference is that, Figure 3 Step S107 can be further refined as follows: Figure 4 Steps S1071 to S1074 in the process.

[0081] In step S1071, if the effect parameters of the intermediate model meet the model effect requirements, the validation sample data is input into the intermediate model to obtain the validation prediction data output by the intermediate model.

[0082] Once the intermediate model's performance parameters meet the required performance conditions, it needs to be validated again to determine if its loss meets the requirements. Validation prediction data includes the prediction data output by the intermediate model based on the validation sample data.

[0083] In step S1072, the loss parameter value is obtained based on the second type of running state data in the verification prediction data and the verification sample data.

[0084] The second type of operational state data in the validation sample data can be considered as the target data in the validation sample data. The validation prediction data is the prediction data output by the intermediate model based on the validation sample data. By comparing the validation prediction data with the second type of operational state data in the validation sample data, the gap between the prediction data output by the intermediate model and the actual situation can be determined. The loss of the validation prediction data relative to the second type of operational state data in the validation sample data can be calculated using a loss function. This loss can be represented by a loss parameter value; that is, the first loss function value can characterize the loss of the validation prediction data relative to the second type of operational state data in the validation sample data.

[0085] In step S1073, if the loss parameter value is within the preset loss range, the intermediate model is determined as the anomaly identification model.

[0086] The preset loss range is used to define a standard range for the loss function values ​​of a successfully trained model. This range can be determined based on the scenario, requirements, experience, etc., and is not limited here. If the loss parameter values ​​fall within the preset loss range, it indicates that the model has been successfully trained, and the intermediate model can be identified as the anomaly detection model.

[0087] In step S1074, if the loss parameter value exceeds the preset loss range, the weights of the training sample data are adjusted according to the loss parameter value, and the intermediate model is trained using the adjusted training sample data until the loss parameter value is within the preset loss range.

[0088] If the loss parameter value exceeds the preset loss range, it indicates that the model training has not been successful and needs to continue. In this case, the weights of each training sample data can be adjusted according to the loss parameter value. Changing the weights of the training sample data will also change the weights of the deep learning recurrent layers. Using the adjusted training sample data to train the intermediate model can optimize it. Validation sample data can then be input into the trained intermediate model to obtain validation prediction data output by the trained intermediate model. Based on the validation prediction data and the second type of running state data in the validation sample data, the loss parameter value is obtained again. It is then determined whether the loss parameter value is within the preset loss range. If the loss parameter value is within the preset loss range, the trained intermediate model is identified as an anomaly detection model. If the first loss function exceeds the preset loss range, the weights of each training sample data are adjusted according to the loss parameter value. This process is repeated until the first loss function corresponding to the trained intermediate model is within the preset loss range, at which point the trained intermediate model is identified as an anomaly detection model.

[0089] For example, Figure 5 A schematic diagram illustrating an example of the training loop of the anomaly recognition model provided in this application embodiment, as shown below. Figure 5 As shown, the intermediate model includes a deep learning recurrent layer 1 and a deep learning recurrent layer 2. Training sample data and validation sample data can be used as inputs to deep learning recurrent layer 1. The output of deep learning recurrent layer 1 can be used as input to deep learning recurrent layer 2. The output of deep learning recurrent layer 2 is the output of the intermediate model. The validation prediction data output by the intermediate model and the real sample target data are processed by a loss function to obtain the loss parameter value. If the loss parameter value exceeds the preset loss range, the loss parameter value can be transmitted to the optimizer. The optimizer adjusts the weights of deep learning recurrent layer 1 and deep learning recurrent layer 2 according to the loss parameter value. That is, the optimizer adjusts the weights of the training sample data according to the loss parameter value. The above process is repeated until the loss parameter value is within the preset loss range.

[0090] To obtain the anomaly detection model, predefined sample data conforming to the requirements of the deep learning model is provided. This sample data includes both input and target data. The deep learning process for the anomaly detection model is configured by defining and compiling the model, completing its training, validation, testing, evaluation, and optimization. Defining the model may include defining its structure, deep learning recurrent layers, neurons, activation functions, etc. Compiling the model may include comparing loss function values, optimizing the optimizer, and monitoring metrics. The deep learning model learns the correlation between sample data and normal and abnormal wind turbine generators, thus obtaining an anomaly detection model that maps inputs to targets. Through validation, the model can be iteratively optimized, enabling it to accurately identify anomalies in complex operating conditions.

[0091] For example, Figure 6 and Figure 7 The diagram shows the loss function curves for training and validating the model using sample data with the same number of iterations. Figure 6 and Figure 7 Each includes two loss function curves: one for training and one for validation. The training loss function curve represents the loss value generated when the model is input with training sample data, while the validation loss function curve represents the loss value generated when the model is input with validation sample data. Figure 6 and Figure 7 As shown, the horizontal axis represents the number of iterations, and the vertical axis represents the loss function value. Figure 6 and Figure 7 Both models underwent 20 iterations, but the sample data used were different. As the number of iterations increased, the training loss function value and the validation loss function value showed a slow downward trend. Figure 6 The training loss function value corresponding to 20 iterations is 0.2254, and the corresponding validation loss function value is 0.2776.

[0092] Figure 8 and Figure 9 The diagram shows the loss function curves for training and validating the model using sample data with the same number of iterations. Figure 8 and Figure 9 Each includes two loss function value curves, one of which is the training loss function value curve, and the other is the validation loss function value curve. For example... Figure 8 and Figure 9 As shown, the horizontal axis represents the number of iterations, and the vertical axis represents the loss function value. Figure 8 and Figure 9The model was iterated 30 times, but the sample data used were different. As the number of iterations increased, the training loss function value and the validation loss function value showed a slow downward trend. Figure 8 The training loss function value corresponding to 30 iterations is 0.222, and the corresponding validation loss function value is 0.2843.

[0093] Figure 10 The diagram shows the loss function curves for training and validating the model using sample data after 200 iterations. Figure 11 The loss function value curve and mean absolute error curve are shown for training and validating the model using sample data after 200 iterations. Figure 10 Two loss function value curves are shown; one is the training loss function value curve, and the other is the validation loss function value curve. Figure 10 As the number of iterations increases, the training loss function curve and the validation loss function curve become very similar, and the training loss function value and the validation loss function value decrease slowly, indicating that the anomaly recognition model has good stability. Figure 11 The validation loss parameter value curve and mean absolute error curve are shown. As the number of iterations increases, the changes in the validation loss parameter value curve and mean absolute error curve are relatively small, indicating that the anomaly recognition model has good stability.

[0094] Figure 12 This is a schematic diagram illustrating an example of predicted data and real data from the anomaly detection model provided in this application embodiment. The horizontal axis represents the input data number, and the vertical axis represents the data value, as shown below. Figure 12 As shown, the predicted data values ​​are very close to, and almost identical to, the actual data values, indicating that the anomaly detection model has a superior anomaly detection performance. Through experiments, the training loss function value of the anomaly detection model reaches 0.00002277, and the mean absolute error (MAO) corresponding to the training sample data reaches 0.00344. The validation loss function value reaches 0.00001558, and the MAO corresponding to the validation sample data reaches 0.002943. The training loss function value, the MAO corresponding to the training sample data, the validation loss function value, and the MAO corresponding to the validation sample data are all relatively small, indicating that the anomaly detection model has a high anomaly detection accuracy.

[0095] It should be noted that the part of obtaining the anomaly identification model based on the optimized intermediate model in step S108 of the above embodiment can also be implemented by the methods of steps S1071 to S1074 above, and is not limited here.

[0096] In some embodiments, the specificity of historical data can be determined by the value of the second type of operating status data in historical data, the pre-divided value range, and the amount of data in the second type of operating status data. Figure 13 This is a flowchart of a wind turbine generator anomaly identification method provided in another embodiment of this application. Figure 13 and Figure 3 The difference is that, Figure 3 Step S105 can be further refined as follows: Figure 13 Steps S1051 to S1053 in the process.

[0097] In step S1051, if the effect parameters of the intermediate model meet the model effect requirements, the specific parameters of the historical data are determined based on the values ​​of the second type of operating state data in the historical data, multiple preset value ranges, and the amount of second type of operating state data in the historical data.

[0098] The values ​​of a set of environmental factor data, first-type operational status data, and second-type operational status data collected at different times within a historical period may differ. To obtain highly specific historical data for sample data, the value range of the second-type operational status data can be pre-divided into multiple intervals, resulting in value intervals. By observing the different values ​​of the second-type operational status data falling within these value intervals, a specificity parameter for the historical data including that second-type operational status data can be determined. The specificity parameter characterizes specificity. In some examples, the specificity parameter is positively correlated with specificity; the larger the specificity parameter, the higher the specificity. In other examples, the specificity parameter is negatively correlated with specificity; the smaller the specificity parameter, the higher the specificity. In this embodiment, for ease of explanation, a positive correlation between the specificity parameter and specificity is primarily used as an example.

[0099] Historical data on wind turbine generators exhibits specificity in longitudinal and / or lateral comparisons. For example, Figure 14 The box plot shows an example of the distribution of vibration data from multiple wind turbine generator sets provided in an embodiment of this application. The horizontal axis, from 1 to 7, represents the numbers of the wind turbine generator sets, and the vertical axis represents the vibration data. Figure 15 This is a schematic diagram illustrating an example of vibration data from multiple wind turbine generators provided in an embodiment of this application. The horizontal axis represents time, which can be in days, and the vertical axis represents the vibration data. Figure 14 and Figure 15 Therefore, for a wind turbine generator, historical data over a relatively long period will exhibit specificity; this specificity refers to the uniqueness of the wind turbine generator in longitudinal comparisons, such as... Figure 14 The vibration data of wind turbine generator set 1 is not concentrated in one range; the vibration data varies and is specific, such as... Figure 15The vibration data of wind turbine generator 1 over multiple days is not concentrated in one range; points A1 to A5 represent relatively prominent vibration data points, exhibiting specificity. For multiple wind turbine generators, historical data from multiple generators within the same time period will also show specificity; this specificity is a characteristic of horizontal comparison between wind turbine generators, such as... Figure 14 Compared to other wind turbines, wind turbine unit 1 exhibits higher vibration data, demonstrating specificity, such as... Figure 15 Compared to other wind turbines, wind turbine 1 exhibits greater fluctuations in vibration data at certain times, demonstrating specificity. Specific parameters for longitudinal comparison of a single wind turbine can be obtained; that is, specific parameters of historical data for a single wind turbine within a historical time period can be obtained. For ease of explanation, these specific parameters for longitudinal comparison are referred to as the first specific parameters, which include the specific parameters for longitudinal comparison of the wind turbine. Specific parameters for lateral comparison of multiple wind turbines can be obtained; that is, specific parameters of historical data for multiple wind turbines within a preset unit time period within a historical time period can be obtained. For ease of explanation, these specific parameters for lateral comparison are referred to as the second specific parameters, which include the specific parameters for lateral comparison of the wind turbine. Historical data that can be converted into sample data can be determined based on the specific parameters for longitudinal and / or lateral comparisons.

[0100] In some examples, the historical time period may include multiple preset time units, the duration of which is shorter than the duration of the historical time period. The historical data of each wind turbine within a preset time unit can form a historical data set. The form of the historical data set is not limited here; for example, the historical data set can be implemented as a historical data document. For example, if the historical time period is one month and the preset unit time period is one day, and there are seven wind turbine generators, then seven historical data documents can be generated each day. If the month includes thirty days, then the seven wind turbine generators can generate two hundred and ten historical data documents in one month. In order to facilitate the differentiation of different historical data documents, the historical data documents can be named according to "wind turbine generator number-date". For example, the names of the seven historical data documents for the seven wind turbine generators on March 6, 2023 can be "1-20230306", "2-20230306", "3-20230306", "4-20230306", "5-20230306", "6-20230306", and "7-20230306".

[0101] The specific parameters in the above embodiments may include a first specific parameter and / or a second specific parameter. The specific details of the first and second specific parameters can be found in the relevant descriptions above and will not be repeated here. A first frequency and / or a second frequency can be obtained based on the values ​​of the second type of operating state data in historical data, multiple preset value ranges, and the amount of second type of operating state data in historical data. A third frequency is obtained based on the relationship between the values ​​of the second type of operating state data in historical data and the historical data set. A first specific parameter is obtained based on the first and third frequencies, and / or a second specific parameter is obtained based on the second and third frequencies.

[0102] The first frequency includes the frequency with which the value of the second operating state data of each wind turbine falls within each value interval during a historical time period. Specifically, the first number of times the value of the second operating state data of each wind turbine falls within each value interval during a historical time period can be calculated, and the second number of the second operating state data for each wind turbine can be counted. For a wind turbine, the first frequency with which the value of its second operating state data falls within a certain value interval is the ratio of the first number of the wind turbine's second operating state data falling within that value interval to the corresponding second number for that wind turbine. The first frequency reflects the probability of different value intervals of the second operating state data of the wind turbine occurring during a historical time period.

[0103] The second frequency includes the frequency with which the values ​​of the second operating status data of multiple wind turbine generators fall within each value interval in each preset unit time period. Specifically, the third number of times the values ​​of the second operating status data of multiple wind turbine generators fall within each value interval in each preset unit time period can be calculated first, and the fourth number of times the values ​​of the second operating status data of multiple wind turbine generators fall within each preset unit time period can be counted. For a preset unit time period, the second frequency with which the values ​​of the second operating status data of multiple wind turbine generators fall within a certain value interval in that preset unit time period is the ratio of the third number of times the values ​​of the second operating status data of multiple wind turbine generators fall within that value interval in that preset unit time period to the fourth number corresponding to that preset unit time period. The second frequency can reflect the probability of different value intervals of the second operating status data of multiple wind turbine generators occurring in the preset unit time period.

[0104] For example, Table 1 shows the second probability that the values ​​of the second operating status data of multiple wind turbine generators fall into each value range:

[0105] Table 1

[0106]

[0107]

[0108] As shown in Table 1, the proportion of second-class operating state data with values ​​in [0, 0.01) is 52.6%, and the proportion of second-class operating state data with values ​​in [0, 0.03) is 98.92%. The second-class operating state data with values ​​greater than or equal to 0.03 are relatively low-probability data.

[0109] The third frequency includes the frequency with which the value interval of the second type of operating state data appears in multiple historical data sets. Specifically, the fifth number of historical data sets containing the value interval of the second type of operating state data can be obtained, and the total number of historical data sets, i.e., the sixth number, can be obtained. For a certain value interval, the third frequency can be obtained based on the ratio of the fifth number to the sixth number corresponding to that value interval. The third frequency can be obtained by taking the logarithm to the base 10 of this ratio.

[0110] In the above embodiments, if the value interval is divided according to a specific rounding rule, the value of the second type of operating status data can also be rounded, so that the rounded value of the second type of operating status data is directly a fixed point within the value interval. For example, the value interval includes [0, 0.01), [0.01, 0.02), [0.02, 0.03), ..., [0.11, 0.12), and the value of the second type of operating status data can be rounded and retained to two decimal places.

[0111] In some examples, the product of the first frequency and the third frequency can be determined as the first specific parameter of the second type of operating state data of the wind turbine falling into the corresponding value range.

[0112] In some examples, the product of the second frequency and the third frequency can be determined as a second specific parameter of the second type of operating state data of multiple wind turbine generators falling into the corresponding value range within a preset unit time period.

[0113] In some examples, the second frequencies corresponding to the same value interval within each preset unit time period in the historical time period can be added together to obtain the first sum. The product of the first sum and the third frequency is determined as the second specific parameter of the second type of operating status data of multiple wind turbine generators falling into the corresponding value interval in the historical time period.

[0114] A smaller value for the specificity parameter indicates smaller differences between historical datasets, meaning the second type of operational status data is more generalized, with lower specificity and representativeness. Conversely, a larger value for the specificity parameter indicates greater differences in the second type of operational status data between different wind turbine generators or within different preset time periods, resulting in higher specificity and representativeness. In some examples, the specificity of sample data can be directly proportional to its weight; that is, the specificity parameter of sample data can be directly proportional to its weight, and the weight of the sample data can be obtained based on the specificity parameter.

[0115] In step S1052, historical data with specific parameters greater than or equal to a preset specific threshold are determined as target historical data.

[0116] A preset specificity threshold is used to filter target historical data. It can be set according to the scenario, needs, experience, etc., and is not limited here. If the specificity parameter of historical data is greater than or equal to the preset specificity threshold, it means that the historical data is specific enough to be used as sample data; if the specificity parameter of historical data is less than the preset specificity threshold, it means that the specificity of historical data is not prominent, it is relatively common, and it is insufficient to be used as sample data. Target historical data includes historical data with a specificity parameter greater than or equal to the preset specificity threshold.

[0117] In step S1053, sample data is obtained based on the target historical data.

[0118] The method of converting target historical data into sample data is basically the same as the method of converting historical data into sample data in the above embodiments. Please refer to the relevant descriptions in the above embodiments, which will not be repeated here.

[0119] In some examples, the specific parameters of the same wind turbine can be cumulatively calculated to obtain a cumulative value. This cumulative value can then be used to identify multiple wind turbines with relatively higher specificity. For example, Figure 16 A comparative schematic diagram illustrating an example of the cumulative values ​​of specific parameters of multiple wind turbine generator sets provided in the embodiments of this application, as shown below. Figure 16 As shown, among wind turbine generator sets 1 to 7, the cumulative value of the specific parameters of wind turbine generator set 1 is relatively higher, that is, wind turbine generator set 1 is more specific than other wind turbine generator sets.

[0120] In some examples, when the predicted data meets preset anomaly conditions, the second type of operational status data, i.e., the original real data, corresponding to the current time period of the predicted data can be extracted from the predicted data, and the second type of operational status data can be visualized for users to view further.

[0121] In this embodiment, timely anomaly identification and early warning of components in wind turbine generator sets can be achieved. By utilizing a large amount of historical data and combining methods such as deep learning, machine learning, and weight adjustment, the trained anomaly identification model can identify and quantify the characteristics of component anomalies in wind turbine generator sets, as well as the correlation and mapping relationships between specific operating conditions. Moreover, this embodiment integrates and analyzes the model method with anomaly triggering mechanisms and processes, enabling more targeted intervention, optimization, and adjustment of the anomaly identification model, thereby improving the anomaly identification accuracy of the model.

[0122] The second aspect of this application provides a wind turbine generator set anomaly identification device. Figure 17 This is a schematic diagram of the structure of a wind turbine generator abnormality identification device provided in an embodiment of this application, as shown below. Figure 17 As shown, the wind turbine generator abnormality identification 200 may include an acquisition module 201, a conversion module 202, a processing module 203, and an abnormality determination module 204.

[0123] The acquisition module 201 can be used to acquire environmental factor data, first-class operating status data and second-class operating status data of the wind turbine generator set in the current time period.

[0124] The conversion module 202 can be used to obtain input data based on environmental factor data, first-type operating status data and second-type operating status data within the current time period.

[0125] The processing module 203 can be used to input input data into a pre-trained anomaly recognition model to obtain the predicted data output by the anomaly recognition model.

[0126] The anomaly detection model comprises at least one deep learning recurrent layer, trained iteratively based on sample data. The sample data is derived from historical data containing wind turbine generators with a specificity exceeding a preset threshold. This historical data includes environmental factor data, first-type operating status data, and second-type operating status data within a historical time period.

[0127] The anomaly determination module 204 can be used to determine that the component corresponding to the second type of operating status data has an anomaly when the predicted data meets the preset anomaly conditions.

[0128] In this embodiment, an anomaly detection model can be pre-trained. This model includes at least one deep learning recurrent layer, constituting a deep learning model. The deep learning model is iteratively trained based on sample data using deep learning. The sample data is obtained from environmental factor data, first-type operating state data, and second-type operating state data within a historical time period where the wind turbine's specificity exceeds a preset specificity. This allows the anomaly detection model to quickly output predictive data corresponding to the input data, indicating whether any component in the wind turbine corresponding to the second-type operating state data has experienced an anomaly. Inputting the input data—based on real-time acquired environmental factor data, first-type operating state data, and second-type operating state data of the wind turbine in the current time period—into the anomaly detection model enables it to output predictive data promptly. Based on the predictive data and preset anomaly conditions, it can promptly determine whether any component in the wind turbine corresponding to the second-type operating state data has experienced an anomaly, improving the timeliness of anomaly detection and enhancing the safety of the wind turbine.

[0129] In some embodiments, the prediction data includes predicted anomaly probabilities and / or predicted operational status data.

[0130] The preset abnormal conditions include: when the predicted data includes the predicted abnormal probability, the predicted abnormal probability is greater than or equal to the preset abnormal probability threshold; when the predicted data includes the predicted operating status data, the difference between the predicted operating status data and the second type of operating status data in the current time period exceeds the preset requirement range.

[0131] Figure 18 This is a schematic diagram of the structure of a wind turbine generator abnormality identification device provided in another embodiment of this application. Figure 18 and Figure 17 The difference is that, Figure 18 The wind turbine generator set anomaly identification device 200 shown may also include a sample acquisition module 205 and a model generation module 206.

[0132] The sample acquisition module 205 can be used to preprocess the acquired historical data to obtain sample data. The sample data includes training sample data.

[0133] The model generation module 206 can be used to establish an initial model, train the initial model using at least a portion of the training sample data to obtain an intermediate model; if the performance parameters of the intermediate model meet the model performance requirements, an anomaly recognition model is obtained based on the intermediate model; if the performance parameters of the intermediate model do not meet the model performance requirements, the intermediate model is optimized until an optimized intermediate model whose performance parameters meet the model performance requirements is obtained, and an anomaly recognition model is obtained based on the optimized intermediate model.

[0134] In some examples, the initial model includes at least one deep learning recurrent layer.

[0135] Optimization processes include one or more of the following:

[0136] By using a preset random deactivation probability, the input units in one or more deep learning recurrent layers are reset to zero;

[0137] Perform random deactivation regularization on one or more deep learning recurrent layers;

[0138] Stacked deep learning recurrent layers, including long short-term memory artificial neural network deep learning recurrent layers, gated recurrent unit deep learning recurrent layers and / or bidirectional recurrent neural network deep learning recurrent layers;

[0139] Adjust at least some of the model parameters of the deep learning recurrent layer.

[0140] In some examples, the sample data also includes validation sample data.

[0141] The model generation module 206 can be used to: input validation sample data into the intermediate model to obtain validation prediction data output by the intermediate model; obtain loss parameter values ​​based on the validation prediction data and the second type of running state data in the validation sample data; determine the intermediate model as an anomaly identification model if the loss parameter values ​​are within the preset loss range; and adjust the weights of the training sample data according to the loss parameter values ​​if the loss parameter values ​​exceed the preset loss range, and train the intermediate model using the adjusted training sample data until the loss parameter values ​​are within the preset loss range.

[0142] In some examples, the sample acquisition module 205 can be used to: determine the specific parameters of historical data based on the values ​​of the second type of operating state data in historical data, multiple preset value ranges, and the amount of second type of operating state data in historical data, where the specific parameters characterize specificity; determine historical data with specific parameters greater than or equal to a preset specific threshold as target historical data; and convert the target historical data to obtain sample data.

[0143] In some examples, the historical time period includes multiple preset time units. The historical data of each wind turbine within a preset time unit forms a historical data set. Specific parameters include a first specific parameter and / or a second specific parameter.

[0144] The sample acquisition module 205 can be used to: obtain a first frequency and / or a second frequency based on the values ​​of the second type of operating status data in historical data, multiple preset value intervals, and the amount of second type of operating status data in historical data. The first frequency includes the frequency at which the value of the second operating status data of each wind turbine is located in each value interval during a historical time period, and the second frequency includes the frequency at which the value of the second operating status data of multiple wind turbines is located in each value interval during each preset unit time period; obtain a third frequency based on the relationship between the values ​​of the second type of operating status data in historical data and the historical data set. The third frequency includes the frequency at which the value interval to which the value of the second type of operating status data in historical data belongs appears in multiple historical data sets; and obtain a first specific parameter based on the first frequency and the third frequency, and / or obtain a second specific parameter based on the second frequency and the third frequency.

[0145] A third aspect of this application also provides an electronic device. Figure 19 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 19 As shown, the electronic device 300 includes a memory 301, a processor 302, and a computer program stored in the memory 301 and executable on the processor 302.

[0146] In some examples, the processor 302 described above may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits that may be configured to implement the embodiments of this application.

[0147] Memory 301 may include read-only memory (ROM), random access memory (RAM), disk storage media device, optical storage media device, flash memory device, electrical, optical, or other physical / tangible memory storage device. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the wind turbine generator anomaly identification method according to embodiments of this application.

[0148] The processor 302 reads the executable program code stored in the memory 301 to run the computer program corresponding to the executable program code, so as to implement the wind turbine generator abnormality identification method in the above embodiment.

[0149] In some examples, the electronic device 300 may also include a communication interface 303 and a bus 304. For example, Figure 19 As shown, the memory 301, processor 302, and communication interface 303 are connected through bus 304 and complete communication with each other.

[0150] The communication interface 303 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application. Input devices and / or output devices can also be connected through the communication interface 303.

[0151] Bus 304 includes hardware, software, or both, that couples components of electronic device 300 together. For example, and not limitingly, bus 304 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-E) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 304 may include one or more buses. Although specific buses are described and illustrated in the embodiments of this application, this application considers any suitable bus or interconnection.

[0152] A fourth aspect of this application also provides a computer-readable storage medium storing computer program instructions. When executed by a processor, these computer program instructions can implement the wind turbine generator abnormality identification method described in the above embodiments and achieve the same technical effect. To avoid repetition, further details are omitted here. The aforementioned computer-readable storage medium may include non-transitory computer-readable storage media, such as read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, etc., and is not limited thereto.

[0153] This application provides a computer program product. When the instructions in the computer program product are executed by the processor of an electronic device, the electronic device can execute the wind turbine generator abnormality identification method in the above embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0154] It should be clarified that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. For the device embodiments, electronic device embodiments, computer-readable storage medium embodiments, and computer program product embodiments, the relevant parts can be referred to the description section of the method embodiments. This application is not limited to the specific steps and structures described above and shown in the figures. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application. Furthermore, for the sake of brevity, detailed descriptions of known methods and techniques are omitted here.

[0155] The aspects of this application have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0156] Those skilled in the art will understand that the above embodiments are exemplary and not restrictive. Different technical features appearing in different embodiments can be combined to achieve beneficial effects. Based on a study of the drawings, specification, and claims, those skilled in the art should be able to understand and implement other variations of the disclosed embodiments. In the claims, the term "comprising" does not exclude other means or steps; the quantifier "a" does not exclude a plurality; the terms "first" and "second" are used to identify names and not to indicate any particular order. No reference numerals in the claims should be construed as limiting the scope of protection. The functionality of multiple parts appearing in the claims can be implemented by a single hardware or software module. The appearance of certain technical features in different dependent claims does not mean that these technical features cannot be combined to achieve beneficial effects.

Claims

1. A method for identifying anomalies in wind turbine generator sets, characterized in that, include: The system acquires environmental factor data, first-class operating status data and second-class operating status data of the wind turbine generator set within the current time period. The second-class operating status data includes operating status data associated with the component under test. The first-class operating status data may include at least some of the other operating status data of the wind turbine generator set besides the second-class operating status data. The input data is obtained based on the environmental factor data, the first type of operating status data and the second type of operating status data within the current time period; The input data is fed into a pre-trained anomaly detection model to obtain the predicted data output by the anomaly detection model. The anomaly detection model includes at least one deep learning recurrent layer and is obtained through iterative training based on sample data. The sample data is obtained based on historical data of the wind turbine generator set whose specificity is higher than or equal to a preset specificity. The historical data includes environmental factor data, first type of operating state data and second type of operating state data within a historical time period. The specificity reflects the degree of specificity of the environmental factor data, first type of operating state data and second type of operating state data of the wind turbine generator set within a certain period of time, and / or the degree of specificity of the environmental factor data, first type of operating state data and second type of operating state data of the wind turbine generator set among multiple wind turbine generator sets. If the predicted data meets the preset abnormal conditions, it is determined that the component under test in the wind turbine generator set corresponding to the second type of operating status data has an abnormality.

2. The method according to claim 1, characterized in that, The prediction data includes predicted anomaly probability and / or predicted operational status data; The preset abnormal conditions include: When the predicted data includes the predicted anomaly probability, the predicted anomaly probability is greater than or equal to a preset anomaly probability threshold. When the predicted data includes the predicted operating status data, the difference between the predicted operating status data and the second type of operating status data in the current time period exceeds the preset requirement range.

3. The method according to claim 1, characterized in that, Before acquiring the environmental factor data, first-type operating status data, and second-type operating status data of the wind turbine generator set within the current time period, the method further includes: The acquired historical data is preprocessed to obtain the sample data, which includes training sample data. An initial model is established, and the initial model is trained using at least a portion of the training sample data to obtain an intermediate model; If the performance parameters of the intermediate model meet the performance requirements, the anomaly recognition model is obtained based on the intermediate model. If the effect parameters of the intermediate model do not meet the model effect requirements, the intermediate model is optimized until the optimized intermediate model is obtained, and the anomaly recognition model is obtained based on the optimized intermediate model.

4. The method according to claim 3, characterized in that, The initial model includes at least one of the deep learning recurrent layers; The optimization process includes one or more of the following: Using a preset random deactivation probability, the input units in one or more of the deep learning recurrent layers are reset to zero; Perform random deactivation regularization on one or more of the deep learning recurrent layers; Stack the deep learning recurrent layers, which include a long short-term memory artificial neural network deep learning recurrent layer, a gated recurrent unit deep learning recurrent layer, and / or a bidirectional recurrent neural network deep learning recurrent layer; Adjust at least some of the model parameters of the deep learning recurrent layer.

5. The method according to claim 3, characterized in that, The sample data also includes validation sample data. The anomaly detection model obtained based on the intermediate model includes: The validation sample data is input into the intermediate model to obtain the validation prediction data output by the intermediate model. Based on the verification prediction data and the second type of running state data in the verification sample data, the loss parameter value is obtained; If the loss parameter value is within a preset loss range, the intermediate model is determined as the anomaly identification model; If the loss parameter value exceeds the preset loss range, the weights of the training sample data are adjusted according to the loss parameter value, and the intermediate model is trained using the weighted training sample data until the loss parameter value is within the preset loss range.

6. The method according to claim 3, characterized in that, The process of preprocessing the acquired historical data to obtain the sample data includes: Based on the values ​​of the second type of operating status data in the historical data, multiple preset value ranges, and the amount of data of the second type of operating status data in the historical data, specific parameters of the historical data are determined, and the specific parameters characterize specificity. Historical data whose specificity parameter is greater than or equal to a preset specificity threshold are identified as target historical data; The sample data is obtained by converting the target historical data.

7. The method according to claim 6, characterized in that, The historical time period includes multiple preset time periods. The historical data of each wind turbine generator set within one preset time period form a historical data set. The specific parameters include a first specific parameter and / or a second specific parameter. The step of determining the specific parameters of the historical data based on the values ​​of the second type of operational status data in the historical data, multiple preset value ranges, and the amount of the second type of operational status data in the historical data includes: Based on the value of the second type of operating status data in the historical data, multiple preset value intervals, and the amount of second type of operating status data in the historical data, a first frequency and / or a second frequency are obtained. The first frequency includes the frequency at which the value of the second type of operating status data of each wind turbine is located in each of the value intervals in the historical time period. The second frequency includes the frequency at which the value of the second type of operating status data of multiple wind turbines is located in each of the value intervals in each preset unit time period. Based on the relationship between the value of the second type of operating status data in the historical data and the historical data set, a third frequency is obtained. The third frequency includes the frequency with which the value range of the second type of operating status data in the historical data appears in multiple historical data sets. The first specific parameter is obtained based on the first frequency and the third frequency, and / or the second specific parameter is obtained based on the second frequency and the third frequency.

8. A wind turbine generator set anomaly identification device, characterized in that, include: The acquisition module is used to acquire environmental factor data, first type of operating status data and second type of operating status data of the wind turbine generator set in the current time period. The second type of operating status data includes operating status data associated with the component under test. The first type of operating status data may include at least some of the other operating status data of the wind turbine generator set besides the second type of operating status data. The conversion module is used to obtain input data based on environmental factor data, first type of operating status data and second type of operating status data within the current time period; The processing module is used to input the input data into a pre-trained anomaly detection model to obtain the predicted data output by the anomaly detection model. The anomaly detection model includes at least one deep learning recurrent layer, which is obtained through iterative training based on sample data. The sample data is obtained based on historical data of the wind turbine generator set whose specificity is higher than a preset specificity. The historical data includes environmental factor data, first type of operating state data and second type of operating state data within a historical time period. The specificity reflects the degree of specificity of the environmental factor data, first type of operating state data and second type of operating state data of the wind turbine generator set within a certain period of time, and / or the degree of specificity of the environmental factor data, first type of operating state data and second type of operating state data of the wind turbine generator set among multiple wind turbine generator sets. The anomaly determination module is used to determine that the component under test corresponding to the second type of operating status data has an anomaly when the predicted data meets the preset anomaly conditions.

9. An electronic device, characterized in that, include: Processor and memory storing computer program instructions; When the processor executes the computer program instructions, it implements the wind turbine generator set anomaly identification method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the wind turbine generator abnormality identification method as described in any one of claims 1 to 7.