Intelligent fault diagnosis method for gearbox of wind turbine generator, electronic equipment and medium

By integrating Bayesian graph convolutional networks and cooperative game theory, important state variables are screened, which solves the problem of noise interference in wind turbine gearbox fault diagnosis, achieves efficient and accurate fault diagnosis, and improves the safe and stable operation of wind turbines.

CN120804896AActive Publication Date: 2025-10-17NORTH CHINA ELECTRIC POWER UNIV
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
CN202511293839.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-10-17
Estimated Expiration
2045-09-11

AI Technical Summary

Technical Problem

In the existing technology, wind turbine gearbox fault diagnosis methods are easily affected by noise under complex working conditions, which makes it difficult to extract and identify fault features, limits the diagnostic accuracy and robustness, and lacks an accurate diagnostic solution.

Method used

A Bayesian graph convolutional network is used to construct a fault diagnosis model. LightGBM and Spearman correlation analysis are combined to screen important state variables. The accuracy and robustness of fault diagnosis are improved by repeatedly inputting state variable data and fusing multiple initial diagnosis results using cooperative game.

Benefits of technology

It achieves fast, efficient and accurate diagnosis of wind turbine gearbox faults, reduces noise impact, improves diagnosis speed and reliability, reduces human intervention, and realizes the transformation from experience-driven to data-driven.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804896A_ABST
    Figure CN120804896A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent fault diagnosis method for a wind turbine generator gearbox, electronic equipment and a medium. The method comprises the steps that real-time operation data of state variables of the wind turbine generator gearbox are obtained, the state variables comprise environment parameters, power grid parameters and component state parameters of the wind turbine generator gearbox, and the real-time operation data comprise operation data within a preset duration; the real-time operation data of the state variables are repeatedly input into a pre-trained fault diagnosis model for multiple times to obtain multiple initial fault diagnosis results, the fault diagnosis model is constructed based on a Bayesian graph convolutional network, and the initial fault diagnosis results comprise the probability that the wind turbine generator gearbox has each fault type in multiple fault types; and fusing the plurality of initial fault diagnosis results by adopting a cooperative game to obtain a final fault diagnosis result. The method can quickly, efficiently and accurately diagnose the fault of the gearbox of the wind turbine generator.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of fan fault diagnosis, and more particularly to an intelligent fault diagnosis method for a gear box of a wind turbine, an electronic device and a medium. BACKGROUND

[0002] As one of the most economical clean energy, wind power plays an increasingly important role in promoting the green and low-carbon transformation of the energy system, addressing climate change and ensuring energy supply security. As a key component of large-scale wind turbines, the gear box is usually installed in a cabin tens or even hundreds of meters high. It is prone to problems such as gear fatigue, tooth wear, gear fracture and bearing failure due to factors such as long-term operation, variable load conditions and external environment. If these faults are not discovered and addressed in a timely manner, irreversible internal faults may occur in the gear box, leading to shutdown accidents of the wind turbine, which not only increases the cost of unit operation and maintenance, but also causes significant economic losses due to the shutdown of the unit.

[0003] In related technologies, vibration signals are usually used as signal sources to measure the faults of the gear box of the wind turbine. Although this traditional fault diagnosis method can identify faults to some extent, the vibration signals of the fan gear box are often disturbed by noise under actual conditions, especially under complex conditions, making it difficult to extract and identify fault features, and the diagnosis accuracy and robustness are often limited. Therefore, there is currently a lack of a scheme that can accurately diagnose the faults of the gear box of the wind turbine.

[0004] In view of the above, the present application is proposed. SUMMARY

[0005] The present application is proposed in consideration of the above problems. According to one aspect of the present application, an intelligent fault diagnosis method for a gear box of a wind turbine is provided, comprising: obtaining real-time running data of state variables of the gear box of the wind turbine, the state variables including environmental parameters, power grid parameters and component state parameters of the gear box of the wind turbine, and the real-time running data including running data within a preset time length; repeatedly inputting the real-time running data of the state variables into a pre-trained fault diagnosis model to obtain multiple initial fault diagnosis results, wherein the fault diagnosis model is constructed based on a Bayesian graph convolution network, and the initial fault diagnosis results include a probability of each fault type in multiple fault types existing in the gear box of the wind turbine; fusing the multiple initial fault diagnosis results using cooperative game to obtain a final fault diagnosis result.

[0006] Exemplarily, before the real-time running data of the state variables is repeatedly input into the pre-trained fault diagnosis model, the method further comprises: screen the state variables based on the real-time running data of the state variables to determine state variables with higher importance; The repeatedly inputting the real-time running data of the state variables into the pre-trained fault diagnosis model multiple times comprises repeatedly inputting the real-time running data of the state variables with higher importance into the pre-trained fault diagnosis model multiple times.

[0007] Exemplarily, the screening of the state variables based on the real-time running data of the state variables comprises: determining the importance of each state variable by LightGBM's Gini exponent and / or Spearman correlation analysis to screen the state variables.

[0008] Exemplarily, the determining the importance of each state variable by LightGBM's Gini exponent and / or Spearman correlation analysis to screen the state variables comprises: determining, based on LightGBM and the real-time running data of the state variables, the Gini exponent of each state variable relative to the plurality of fault types, wherein the Gini exponent is positively correlated with the importance of the corresponding state variable; selecting the first preset number of state variables with larger Gini exponents as the state variables with higher importance; or, The determining the importance of each state variable by LightGBM's Gini exponent and / or Spearman correlation analysis to screen the state variables comprises: calculating, based on the real-time running data of the state variables, the Spearman correlation between each state variable and other state variables; for any state variable in the state variables, when the Spearman correlation between the state variable and at least part of the other state variables is greater than or equal to a correlation threshold, determining the state variable as a state variable with higher importance.

[0009] Exemplarily, the determining the importance of each state variable by LightGBM's Gini exponent and / or Spearman correlation analysis to screen the state variables comprises: determining, based on LightGBM and the real-time running data of the state variables, the Gini exponent of each state variable relative to the plurality of fault types, wherein the GiniThe index is positively correlated with the importance of the corresponding state variable; choose Gini The first preset number of state variables with larger indexes constitute a first state variable set; Calculating the Spearman correlation between each state variable and other state variables based on the real-time operation data of the state variables; For any state variable among the state variables, when the Spearman correlation between the state variable and at least some of the other state variables is greater than or equal to a correlation threshold, add the state variable to the second state variable set; Taking the intersection of the first state variable set and the second state variable set as a state variable with higher importance; or, The use of LightGBM Gini Index and / or Spearman correlation analysis to determine the importance of each state variable to screen the state variables, including: Calculating the Spearman correlation between each state variable and other state variables based on the real-time operation data of the state variables; For any state variable among the state variables, when the Spearman correlation between the state variable and at least some of the other state variables is greater than or equal to a correlation threshold, add the state variable to the third state variable set; Based on LightGBM and the real-time operation data of each state variable in the third state variable set, determine the relative value of each state variable in the third state variable set to the multiple fault types. Gini Index, where Gini The index is positively correlated with the importance of the corresponding state variable; choose Gini The first preset number of state variables with larger indexes are regarded as state variables with higher importance; or, Based on LightGBM and the real-time operation data of the state variables, determine the relative value of each state variable to the multiple fault types. Gini Index, where Gini The index is positively correlated with the importance of the corresponding state variable; choose Gini The first preset number of state variables with larger exponents constitute a fourth state variable set; calculating, based on the real-time operating data of each state variable in the fourth state variable set, a Spearman correlation between each state variable in the fourth state variable set and other state variables in the fourth state variable set; For any state variable in the fourth set of state variables, when the Spearman correlation between the state variable and at least part of the other state variables in the fourth set of state variables is greater than or equal to a correlation threshold, the state variable is determined to be a state variable with higher importance.

[0010] Exemplarily, before the real-time running data of the state variable is repeatedly input into the pre-trained fault diagnosis model multiple times, the method further comprises: normalizing the real-time running data of the state variable with higher importance.

[0011] Exemplarily, the fusion of the multiple initial fault diagnosis results by using the cooperative game to obtain the final fault diagnosis result comprises: For any one of the multiple initial fault diagnosis results, a consistency correlation coefficient between the initial fault diagnosis result and the remaining fault diagnosis results is calculated, the remaining fault diagnosis results being the sum of the other initial fault diagnosis results except the initial fault diagnosis result in the multiple initial fault diagnosis results; based on the sum of the products of the multiple initial fault diagnosis results and the respective consistency correlation coefficients, the final fault diagnosis result is determined.

[0012] Exemplarily, before the real-time running data of the state variable is repeatedly input into the pre-trained fault diagnosis model multiple times, the method further comprises: processing the real-time running data of the state variable by using a linear interpolation method.

[0013] According to still another aspect of the present application, an electronic device is provided, comprising a processor and a memory, the memory storing a computer program, and the processor being configured to execute the computer program to implement the method as described above.

[0014] According to yet another aspect of the present application, a computer readable storage medium is provided, storing a computer program / instruction, which, when executed by a processor, implements the method as described above.

[0015] In the technical solution, the Bayesian graph convolution network is combined to significantly enhance the ability of nonlinear feature extraction of the wind turbine gearbox fault; the input is repeated multiple times, and the multiple initial fault diagnosis results obtained by the multiple inputs are fused in a cooperative game manner to reduce the influence of accidental noise and avoid accidental errors caused by random model parameters, which helps to improve the robustness and reliability of the method and improve the accuracy of the prediction results; in addition, the scheme does not need to combine engineering practice experience for fault feature extraction and screening, reduces human intervention, can realize the change from the experience-driven artificial feature paradigm to the data-driven representation learning paradigm, helps to improve the result accuracy, and can improve the fault diagnosis speed. In summary, the method can quickly, efficiently and accurately diagnose the fault of the wind turbine gearbox.

[0016] The above description is only a summary of the technical solutions of the present application, in order to enable the technical means of the present application to be more clearly understood, and to be implemented according to the content of the specification, and in order to enable the above and other purposes, features and advantages of the present application to be more apparent and easy to understand, the specific embodiments of the present application are described below. BRIEF DESCRIPTION OF DRAWINGS

[0017] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description of embodiments of the present application taken in conjunction with the accompanying drawings. The drawings provided in the specification and the contents of the specification help to provide further understanding of the embodiments of the present application, and constitute a part of the specification, and are used to explain the present application together with the embodiments of the present application, and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0018] Figure 1 a schematic flowchart of a wind turbine gearbox intelligent fault diagnosis method according to an embodiment of the present application is shown; Figure 2 a schematic block diagram of an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0019] In order to make the purposes, technical solutions and advantages of the present application more apparent, the following will describe the example embodiments according to the present application with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments of the present application, and it should be understood that the present application is not limited by the example embodiments described herein. Based on the embodiments of the present application described in the present application, all other embodiments obtained by those skilled in the art without creative labor should fall within the protection scope of the present application.

[0020] Wind turbine is a complex mechanical and electrical system, usually arranged in remote areas, long-term exposure to harsh environments, working conditions are very poor, repair and maintenance costs are high. Due to the long-term operation of the unit in harsh conditions, the failure rate is high. The gearbox of the unit is often installed in a high and narrow area, and the common gearbox failure types include gear wear, gear breakage, poor lubrication, etc. Bearing failure may also occur, such as pitting, insufficient lubrication, fatigue damage and cracks. Therefore, it is of great practical significance to study the intelligent fault diagnosis method of the gearbox of the wind turbine unit to improve the safe and stable operation of the unit, reduce the operation and maintenance cost of the wind turbine unit, and promote the safe, reliable and efficient operation of the wind turbine unit in the new power system. In view of this, the present application provides an intelligent fault diagnosis method for the gearbox of the wind turbine unit, which can quickly, efficiently and accurately diagnose the fault of the gearbox of the wind turbine unit.

[0021] According to an aspect of an embodiment of the present application, an intelligent fault diagnosis method for the gearbox of a wind turbine unit is provided. Figure 1 A schematic flow chart of the intelligent fault diagnosis method for the gearbox of the wind turbine unit according to an embodiment of the present application is shown. As shown in Figure 1 The method can include the following steps S110, S150 and S160.

[0022] In step S110, real-time operation data of state variables of the gearbox of the wind turbine unit is obtained, the state variables including environmental parameters, power grid parameters and component state parameters of the gearbox of the wind turbine unit, and the real-time operation data including operation data within a preset time length.

[0023] The real-time operation data of each state variable of the gearbox of the wind turbine unit can be collected by a sensor and stored in the data acquisition and monitoring control system of the wind turbine unit. In this embodiment, the state variables include environmental parameters, power grid parameters and component state parameters of the gearbox of the wind turbine unit. Specifically, the environmental parameters can include wind speed, environmental temperature and wind direction. The environmental parameters directly affect the operating state of each component of the wind turbine unit. The power grid parameters can be the active power of the wind turbine unit, the current size generated by the wind turbine unit, etc. The power grid parameters can reflect the dispatching situation of the power grid where the wind turbine unit is located, and are of great significance to the predictive maintenance of the wind turbine unit. The component state parameters of the gearbox of the wind turbine unit can include blade pitch angle, pitch motor temperature (also referred to as blade motor temperature), pitch rate (also referred to as blade pitch speed), bearing temperature of the gearbox, bearing temperature of the generator, oil temperature of the gearbox and the generator, wind wheel speed, etc. In a specific embodiment, the state variables of the wind turbine unit are shown in Table 1.

[0024] Table 1 State variables of wind turbine unit

[0025] In the embodiment, the real-time running data includes running data in a preset time length. That is, for any state variable, the real-time running data of the state variable includes all running data collected in the preset time length with the current time as the end point. For example, if the preset time length is 3s, the current time is 12:13:28, and the state variable is the temperature at the rear end of the high-speed shaft of the gearbox, the real-time running data of the state variable includes multiple temperatures at the rear end of the high-speed shaft of the gearbox collected from 12:13:25 to 12:13:28.

[0026] In step S150, the real-time running data of the state variable is repeatedly input into the pre-trained fault diagnosis model to obtain multiple initial fault diagnosis results, wherein the fault diagnosis model is constructed based on a Bayesian graph convolutional network, and the initial fault diagnosis results include a probability of each fault type in multiple fault types of the gearbox of the wind turbine generator.

[0027] In the embodiment, the multiple fault types include common faults such as root crack, gear fracture, missing tooth, gear pitting, gearbox bearing inner / outer ring fault, and gearbox bearing roller fault.

[0028] In some implementation schemes of the example, the fault diagnosis model is trained in the following manner: historical running data of the state variable of the wind turbine generator is obtained; the historical running data is divided into multiple groups of data in a sliding window manner, and the length of the sliding window is a preset time length; a fault label is added to each group of data according to the input of a user; the multiple groups of data are divided into a training set, a validation set, and a test set in a ratio of 6:2:2, and the training set, the validation set, and the test set are used to complete training of the fault diagnosis model. The fault label can be obtained in combination with expert experience, and the fault label is used to mark the fault type. Those skilled in the art can understand the specific implementation manner of training the fault diagnosis model by using the training set, the validation set, and the test set, and thus no further description is given.

[0029] In the scheme of the example, the fault diagnosis model is constructed based on a Bayesian graph convolutional (Bayesian-GCNN) network. The Bayesian-GCNN network is a graph convolutional neural network that fuses a Bayesian method, and the network implements the parameter family of a random graph observed. Then, inference is performed on the joint posterior of the random graph parameters and the weights and node labels in the GCNN, and the posterior probability of the label is: ; wherein, W is a random variable, and represents a random graph the weight of the Bayesian-GCNN network, and λ is a parameter of the random graph.

[0030] In Bayesian-GCNN, the graph is modeled by a combined membership stochastic block model (a-MMSBM) and the parameters λ = {π, β} are learned by stochastic optimization.

[0031] MMSBM can capture complex relationships in networks and is usually used to model graphs with relatively strong community structure. The core idea is that each node belongs to different communities with a certain probability and the random block model is extended through its classification behavior.

[0032] Since is usually noisy and may not fit the adopted parametric block model well, the maximum a posteriori estimation is used instead of the integral in the parameter π and β is as follows: ; By adding appropriate priors to π and β and using the approximation: ; where W s,i is obtained by sampling from the graph corresponding to the Bayesian-GCNN network using the Monte Carlo method. is sampled from .

[0033] In some embodiments, when training a fault diagnosis model constructed based on a Bayesian-GCNN network, the data of the training set can be input first, the model input X , output Y and observed graph are set, and the GCNN is trained to initialize the inference in MMSBM and the weights in Bayesian-GCNN. Then, the MMSBM is iteratively trained to obtain , and is sampled from . Finally, the weights are obtained by sampling from the Bayesian-GCNN network corresponding to the graph W i using the Monte Carlo method, and the approximation of the posterior probability is calculated. After a certain number of rounds of training on the training set, the model with the highest accuracy on the validation set is obtained as the final fault diagnosis model.

[0034] ​The initial fault probability diagnosis result of the embodiment can be expressed in the form of a two-dimensional matrix, in which the probability of each fault of the current wind turbine is shown.

[0035] In step S160, the plurality of initial fault diagnosis results are fused by using cooperative game to obtain the final fault diagnosis result.

[0036] In step S150, a plurality of initial fault diagnosis results are obtained by repeatedly inputting real-time running data of the state variable. In the scheme of the present example, considering the randomness of Bayesian inference, the weights in the model are randomly sampled each time of training or inference. In order to reduce the influence of contingency, in the embodiment, the input is repeated multiple times and multiple results are obtained. The specific number of repetitions can be determined comprehensively according to the real-time requirement and result accuracy requirement of the user. In a specific embodiment, the number of repetitions can be three. Since the weights in the model are randomly sampled, there are differences between the plurality of initial fault diagnosis results obtained. W W

[0037] After obtaining the plurality of initial fault diagnosis results, the plurality of initial fault diagnosis results are fused by using cooperative game. Cooperative game refers to a game in which cooperation is used, and the result is that the interests of at least one party are increased, so cooperative game can increase the overall interests. In the present example, the plurality of initial fault diagnosis results are fused by using cooperative game, and the plurality of different initial fault probability diagnosis results are regarded as participants of a whole to carry out cooperative game, so as to output the final fault diagnosis result. By fusing the plurality of initial fault diagnosis results by using cooperative game, the accuracy of the final result obtained, as well as the robustness and credibility of the diagnosis, can be improved.

[0038] In the present application, the screened state variables are used as the input feature matrix of the Bayesian graph convolutional neural network X , and the graph structure constructed by the correlation between the state variables G is used to perform inference under the Bayesian inference framework to output the posterior probability of each fault type. Since the Bayesian network has randomness, the inference is repeated three times, and the three results are fused to improve the robustness and credibility of the diagnosis.

[0039] ​​In the technical solution, the Bayesian graph convolution network is combined to significantly enhance the ability of nonlinear feature extraction of the wind turbine gearbox fault; the input is repeated multiple times, and the multiple initial fault diagnosis results obtained by fusing the multiple inputs in a cooperative game manner can reduce the influence of accidental noise and avoid accidental errors caused by random model parameters, which helps to improve the robustness and reliability of the method and the accuracy of the prediction results. In addition, the scheme does not need to combine engineering practice experience for fault feature extraction and screening, reduces human intervention, realizes the change from the experience-driven artificial feature paradigm to the data-driven representation learning paradigm, and helps to improve the result accuracy and fault diagnosis speed. In summary, the method can quickly, efficiently and accurately diagnose the fault of the wind turbine gearbox.

[0040] Exemplarily, before the real-time running data of the state variables is repeatedly input into the pre-trained fault diagnosis model, the method further includes: step S130, screening the state variables based on the real-time running data of the state variables to determine state variables with higher importance; wherein step S150, repeatedly inputting the real-time running data of the state variables into the pre-trained fault diagnosis model includes: repeatedly inputting the real-time running data of the state variables with higher importance into the pre-trained fault diagnosis model.

[0041] In step S110, the real-time running data of each state variable of the wind turbine is obtained. The inventors have found through research that different state variables have different contribution degrees to the target parameter (i.e. the fault type), and too many state variables will increase the redundancy of the feature space, which will increase the calculation amount of the algorithm and also negatively affect the fault diagnosis accuracy of the target, reducing the performance of the model. In view of this, in the scheme of the present example, before the real-time running data of multiple state variables is input into the pre-trained fault diagnosis model, it is considered to screen each state variable and select state variables with higher importance as the input of the model. The specific screening methods include but are not limited to one or a combination of correlation coefficient, random forest, game theory, LightGBM Gini exponent, Spearman correlation analysis, etc., which are not limited in the present example.

[0042] The above scheme can realize dimension reduction of data by screening the real-time running data of state variables with higher importance as the input of the model, reduce the calculation burden of the model training and actual application, and at the same time, can avoid the negative influence of irrelevant state variables on the result and improve the accuracy of the model.

[0043] Exemplarily, step S130, screening the state variables based on the real-time running data of the state variables includes: using the LightGBM GiniExponential and / or Spearman correlation analysis determines the importance of each state variable to screen the state variables.

[0044] In the scheme of the present example, the use of LightGBM is limited to Gini Exponential and / or Spearman correlation analysis. Among them, LightGBM is a high-efficiency machine learning algorithm based on gradient boosting framework, which is widely used in feature screening. The algorithm uses an innovative histogram algorithm to find the optimal split point, and quantifies the feature importance by counting the number of times each feature is selected as a split point, so as to intuitively reflect the contribution of the feature in the model construction. Its core advantages are reflected in three aspects: first, its unique feature splitting strategy combined with efficient parallel computing mechanism significantly improves the model training speed, especially suitable for processing large-scale data sets and high-dimensional feature space; second, based on gradient boosting technology, the model prediction ability is continuously optimized, which can maintain high accuracy in classification and regression tasks; finally, the algorithm supports multiple data format input, has good compatibility, and can flexibly adapt to the needs of different application scenarios. In summary, LightGBM's Gini Exponential and / or Spearman correlation analysis. Among them, LightGBM is a high-efficiency machine learning algorithm based on gradient boosting framework, which is widely used in feature screening. The algorithm uses an innovative histogram algorithm to find the optimal split point, and quantifies the feature importance by counting the number of times each feature is selected as a split point, so as to intuitively reflect the contribution of the feature in the model construction. Its core advantages are reflected in three aspects: first, its unique feature splitting strategy combined with efficient parallel computing mechanism significantly improves the model training speed, especially suitable for processing large-scale data sets and high-dimensional feature space; second, based on gradient boosting technology, the model prediction ability is continuously optimized, which can maintain high accuracy in classification and regression tasks; finally, the algorithm supports multiple data format input, has good compatibility, and can flexibly adapt to the needs of different application scenarios. In summary, LightGBM's Gini Exponential and / or Spearman correlation analysis. Among them, LightGBM is a high-efficiency machine learning algorithm based on gradient boosting framework, which is widely used in feature screening. The algorithm uses an innovative histogram algorithm to find the optimal split point, and quantifies the feature importance by counting the number of times each feature is selected as a split point, so as to intuitively reflect the contribution of the feature in the model construction. Its core advantages are reflected in three aspects: first, its unique feature splitting strategy combined with efficient parallel computing mechanism significantly improves the model training speed, especially suitable for processing large-scale data sets and high-dimensional feature space; second, based on gradient boosting technology, the model prediction ability is continuously optimized, which can maintain high accuracy in classification and regression tasks; finally, the algorithm supports multiple data format input, has good compatibility, and can flexibly adapt to the needs of different application scenarios. In summary, LightGBM's

[0045] In one implementation of this example, LightGBM is used Gini The index and Spearman correlation analysis determine the importance of each state variable to screen the state variables. By combining the two methods, the advantages of the two screening methods can be complemented, thereby further improving the screening efficiency and screening effect.

[0046] For example, using LightGBM Gini Index and / or Spearman correlation analysis to determine the importance of each state variable to screen the state variables, including: based on LightGBM and the real-time operation data of the state variables, determine the importance of each state variable relative to multiple fault types Gini Index, where Gini The index is positively correlated with the importance of the corresponding state variable; Gini The first preset number of state variables with larger indexes are regarded as state variables with higher importance.

[0047] In some embodiments, based on LightGBM and real-time operation data of state variables, the relative performance of each state variable relative to multiple fault types is determined. Gini The index consists of determining each state variable by Gini index: Assume there is n state variables { X 1, X 2…… X n} Input LightGBM, and these state variables have K categories (i.e. K type of failure). Assume V is the importance of the state variable, then j The importance of the state variables is , then the node m of Gini The index is: ; in, p mk Representation node m The selected samples belong to the category k probability.

[0048] When the state variable X j Selected as a node m When the partitioning feature of m is divided into two child nodes (left node and right node), then the node m Before and after division Gini Index change (i.e.Gini Gain) is: ; in, G l and G r Respectively represented by nodes m The two new nodes of the split Gini index.

[0049] State variables X j In the i The importance of a tree is defined as the contribution of the variable to all nodes. Gini Sum of gains: ; State variables X j Importance in the entire forest of LightGBM is defined as: ; in, C Indicates the number of trees in the LightGBM forest.

[0050] To facilitate comparison, the importance of all state variables is normalized. X j The normalized importance of is defined as: ; in, n is the total number of state variables, is a state variable X j The Gini index, for n The sum of the gains of the state variables.

[0051] In this example, after obtaining each state variable Gini After indexing, the state variables can be sorted from largest to smallest, and a preset number of the top-ranked state variables are selected as the state variables of higher importance. This preset number can be selected based on actual needs. For example, the preset number can be n-2, where n is the total number of state variables. Of course, the preset number can also be adjusted based on user input, which will not be explained in detail.

[0052] The above technical solution uses LightGBM GiniThe importance of each state variable is measured by the index, which can efficiently capture the nonlinear relationship and interaction effect between each state variable and the failure of the wind turbine, so as to screen out more important state variables while reducing the dimension, which helps to improve the efficiency of fault diagnosis while ensuring the accuracy of fault diagnosis.

[0053] Exemplarily, the importance of each state variable is determined by the index and / or the Spearman correlation analysis, and the state variables are screened, including: based on the real-time running data of the state variables, the Spearman correlation between each state variable and other state variables is calculated; for any state variable in the state variables, when the Spearman correlation between the state variable and at least part of the other state variables is greater than or equal to a correlation threshold, the state variable is determined to be a state variable with higher importance. Gini

[0054] The correlation threshold can be a preset value or a real-time value input by the user according to the actual situation, which is not limited herein. It can be understood that the Spearman correlation has positive and negative values, and the Spearman correlation in the above scheme herein is the absolute value. In one specific embodiment, the correlation threshold can be 0.1.

[0055] In some implementation schemes, based on the real-time running data of the state variables, the Spearman correlation between each state variable and other state variables is calculated, including: for any state variable, the Spearman correlation between the state variable and any state variable (for the sake of distinction, this state variable is called target variable) in other state variables: In the formula, p XY is the Spearman correlation coefficient between the state variable X and the target variable Y . , are the rankings of the first X , Y data in the sample size i data, h d i is the difference of the ranking of each pair of variables.

[0056] ​​​After obtaining the Spearman correlation between each state variable and other state variables, for any state variable in the state variables, if the Spearman correlation between the state variable and at least part of the other state variables is greater than or equal to the correlation threshold, the state variable is determined to be a state variable with higher importance. The at least part can be one or more, for example, when the total number of state variables is n, the at least part can be n-1, n-2, n-3, etc.

[0057] The above scheme determines the importance of each state variable through or Spearman correlation analysis, which can accurately identify and remove state variables with weak correlation with most other state variables. Such variables may be isolated variables in the system and are difficult to participate in collaborative representation in the modeling process. By excluding them, the dimension compression and noise filtering of the feature space can be achieved, and the collaborative representation ability of each state variable of the input model can be enhanced, which helps to further improve the fault diagnosis accuracy.

[0058] For example, the importance of each state variable is determined by using the LightGBM Gini exponent and / or Spearman correlation analysis to screen the state variables, including: based on the LightGBM and the real-time running data of the state variables, determining the Gini exponent of each state variable relative to a plurality of fault types, wherein Gini the exponent is positively correlated with the importance of the corresponding state variable; selecting the first preset number of state variables with larger Gini exponents to form a first state variable set; based on the real-time running data of the state variables, calculating the Spearman correlation between each state variable and other state variables in the state variables; for any state variable in the state variables, if the Spearman correlation between the state variable and at least part of the other state variables is greater than or equal to the correlation threshold, the state variable is added to a second state variable set; and the intersection of the first state variable set and the second state variable set is taken as the state variable with higher importance.

[0059] The state variables Gini exponent and the specific implementation of the Spearman correlation of the state variables are described in detail above, and will not be repeated.

[0060] In the scheme of the present example, the original obtained plurality of state variables are screened by using the LightGBM Gini exponent and Spearman correlation analysis, and the results of the screening are taken as a union set, so that a combination of state variables with higher correlation with the faults of the wind turbine and stronger collaborative representation ability can be obtained, which helps to further improve the accuracy of fault diagnosis.

[0061] For example, using LightGBM Gini The importance of each state variable is determined by index and / or Spearman correlation analysis to screen the state variables, including: calculating the Spearman correlation between each state variable and other state variables in the state variables based on the real-time operation data of the state variables; for any state variable in the state variables, when the Spearman correlation between the state variable and at least some of the state variables in the other state variables is greater than or equal to the correlation threshold, adding the state variable to the third state variable set; determining the importance of each state variable in the third state variable set relative to multiple fault types based on LightGBM and the real-time operation data of each state variable in the third state variable set; Gini Index, where Gini The index is positively correlated with the importance of the corresponding state variable; Gini The first preset number of state variables with larger indexes are regarded as state variables with higher importance.

[0062] Determine state variables Gini The specific implementation of the Spearman correlation of the index and state variables has been described in detail above and will not be repeated here.

[0063] In the solution of this example, we first consider using Spearman correlation analysis to remove isolated variables from the multiple state variables originally obtained, and then use LightGBM to screen state variables with a high correlation with wind turbine faults. In this way, we can obtain a combination of state variables that has a high correlation with wind turbine faults and a strong collaborative characterization capability, which helps to further improve the accuracy of fault diagnosis.

[0064] For example, based on LightGBM and real-time operation data of state variables, the relative value of each state variable to various fault types is determined. Gini Index, where Gini The index is positively correlated with the importance of the corresponding state variable; Gini A preset number of state variables with larger exponents constitute a fourth state variable set; based on the real-time operation data of each state variable in the fourth state variable set, the Spearman correlation between each state variable in the fourth state variable set and other state variables in the fourth state variable set is calculated; for any state variable in the fourth state variable set, when the Spearman correlation between the state variable and at least some of the other state variables in the fourth state variable set is greater than or equal to a correlation threshold, the state variable is determined to be a state variable with higher importance.

[0065] Determine state variables Gini The specific implementation of the Spearman correlation of the index and state variables has been described in detail above and will not be repeated here.

[0066] In the scheme of the present example, the LightGBM-based Gini exponential is first used for coarse selection, and then the Spearman correlation analysis is used for fine selection. Specifically, first, the feature importance of the original state variable is evaluated based on LightGBM. According to the importance score obtained, a descending order is sorted, so as to identify the state variables that show low feature contribution under multiple fault types. These state variables can be regarded as redundant features with limited contribution to the modeling of the overall system, and can be excluded. Subsequently, the Spearman correlation analysis is performed on the state variable set remaining after the first step of screening, to identify variables with weak correlation with most other variables, i.e. state variables with generally low absolute value of Spearman correlation coefficient. Such state variables, which fail to form a stable association structure with other state variables, may be isolated variables and are difficult to participate in collaborative representation, so they are further excluded. Thus, a combination of state variables with high correlation with the faults of the wind turbine and strong collaborative representation capability can be obtained, which helps to further improve the accuracy of fault diagnosis. The screening scheme can be used as a preferred state variable screening scheme in the present application.

[0067] The screening strategy of "LightGBM first and then Spearman" proposed in the present application has significant advantages and irreplaceability in the precision, stability and performance of the final model. First, LightGBM can identify key variables with the ability to distinguish multiple faults from the model perspective, and has the ability to handle nonlinear relationships and complex interactions between variables. Through the preliminary screening in this stage, the feature space can be effectively compressed, and low-value variables that are not sensitive to multiple faults can be excluded. Subsequently, based on the Spearman correlation analysis of the pairwise relationship between variables, variables with correlation coefficients lower than a predetermined threshold are excluded, thereby excluding "isolated variables" or noise variables that lack collaborative trends with other variables, and retaining a set of variables with structural correlation and potential collaborative information, which helps to improve the stability, structural integrity and fault representation capability of the model. In contrast, if Spearman correlation analysis between variables is performed first and then LightGBM modeling is performed, some features that are weakly correlated with other variables but can form nonlinear interactions with fault variables in the model and have important modeling value may be incorrectly excluded in the early stage, which weakens the model effect. If the union of the screening results of the two methods is used, it is difficult to effectively control the total amount of variables, which may introduce noise features and redundant information, and reduce the model training efficiency and prediction robustness.

[0068] Exemplarily, before repeatedly inputting the real-time running data of the state variables with high importance into the pre-trained fault diagnosis model, the method further comprises: step S140, normalizing the real-time running data of the state variables with high importance.

[0069] In some implementations, any real-time operating data (also referred to as specific operating data) of any state variable (also referred to as a specific state variable) may be normalized using the following formula: ; Where, x i For specific operation data, x min and x max The minimum and maximum values ​​in the real-time running data of a specific state variable.

[0070] In this example, normalizing the real-time data before inputting it into the model can eliminate the interference of feature scale differences on the model, reduce the predictive error of the model, and improve the accuracy and robustness of the model.

[0071] Exemplarily, a cooperative game is used to fuse multiple initial fault diagnosis results to obtain a final fault diagnosis result, including: for any one of the multiple initial fault diagnosis results, calculating the consistency correlation coefficient between the initial fault diagnosis result and the remaining fault diagnosis results, where the remaining fault diagnosis result is the sum of the other initial fault diagnosis results in the multiple initial fault diagnosis results except the initial fault diagnosis result; and determining the final fault diagnosis result based on the sum of the products of the multiple initial fault diagnosis results and their respective corresponding consistency correlation coefficients.

[0072] Optionally, calculating a consistency correlation coefficient between the initial fault diagnosis result and the remaining fault diagnosis results includes calculating the consistency correlation coefficient using the following formula: ; Where, represents the consistency correlation coefficient; Indicates the initial fault diagnosis result; Indicates the first i The value (that is, i probability of a particular failure type); Indicates the remaining fault diagnosis results. i values; express The average value of the probability of each failure in; Represents the average value of the probability of each fault in the remaining fault diagnosis results; m Indicates the total number of fault types.

[0073] Optionally, determining the final fault diagnosis result based on a sum of products of the plurality of initial fault diagnosis results and the respective corresponding consistency correlation coefficients comprises: calculating the sum of products of the plurality of initial fault diagnosis results and the respective corresponding consistency correlation coefficients; and normalizing each element in the sum of products to obtain the final fault diagnosis result.

[0074] In some embodiments, calculating the sum of products of the plurality of initial fault diagnosis results and the respective corresponding consistency correlation coefficients comprises calculating by the following formula: .

[0075] Normalizing each element in the sum of products comprises normalizing by the following formula: .

[0076] The above technical solution can effectively improve the fault diagnosis accuracy by fusing a plurality of fault probability diagnosis matrices through cooperative game.

[0077] Exemplarily, before repeatedly inputting the real-time running data of the state variables into the pre-trained fault diagnosis model, the method further comprises: step S120, processing the real-time running data of the state variables by using a linear interpolation method. Specifically, the linear interpolation method can be used to process the missing values and abnormal values in the real-time running data of each state variable. It can be understood that, due to data transmission and storage anomalies and other reasons, there are missing values and abnormal values in the directly derived system data, and the existence of the missing values and abnormal values will affect the fitting effect of the model. In this embodiment, the missing values and abnormal values can be identified first, and the abnormal values can be directly deleted. Those skilled in the art can understand the identification method of the abnormal values, which will not be described herein. For the deleted abnormal values and missing values, the linear interpolation method can be used to supplement the data to ensure the data integrity.

[0078] In some embodiments, the linear interpolation method can be represented by the following formula: ; wherein, y i is the value of the i th point (i.e. the i th value of the real-time running data of a certain state variable) after being replaced. i i y i-1 , y i-2 is the i th point before the nearest two normal points (i.e. the i-1 th value and the i-2 th value of the real-time running data of a certain state variable). i i i ​​​​​

[0079] The technical solution adopts the linear interpolation method to process the real-time running data of the state variable, which can ensure the continuity of the data and avoid the adverse effect of abnormal data on the result, and ensure the accuracy of the fault diagnosis result.

[0080] In one specific embodiment of the present disclosure, the method comprises steps S110, S120, S130, S140, S150 and S160. The step S150 adopts the preferred state variable screening scheme.

[0081] According to another aspect of the embodiments of the present disclosure, an electronic device is also provided. Figure 2 A schematic block diagram of an electronic device according to one embodiment of the present disclosure is shown. As shown, the electronic device 200 comprises a processor 210 and a memory 220. The memory 220 stores a computer program, and the processor 210 is configured to execute the computer program to implement the method described above. Figure 2 According to another aspect of the embodiments of the present disclosure, an electronic device is also provided.

[0082] According to another aspect of the embodiments of the present disclosure, a computer readable storage medium is also provided. The storage medium stores a computer program / instruction, and the computer program / instruction is executed by a processor to implement the method described above. The storage medium may, for example, include a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a compact disc read-only memory (CD-ROM), a USB memory, or any combination of the above storage media. The computer readable storage medium can be any combination of one or more computer readable storage media.

[0083] Those skilled in the art can easily understand the implementation structure, working principle and beneficial effects of the electronic device and the computer readable storage medium by reading the above method. For brevity, they will not be described here.

[0084] Although the example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the example embodiments described above are merely exemplary and are not intended to limit the scope of the present disclosure. Those skilled in the art can make various changes and modifications without departing from the scope and spirit of the present disclosure.

[0085] Those skilled in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present disclosure.

[0086] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other manners. For example, the division of the units is only a logical function division, and there can be another division manner in actual implementation, for example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the illustrated or described embodiments of the present application are not intended to limit the scope of the present application. Any variations or replacements within the idea of the present application should be construed as falling within the scope of the present application.

[0087] In the specification provided herein, a large number of specific details are illustrated. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure the understanding of the present specification.

[0088] In addition, those skilled in the art can understand that the combination of features of different embodiments means to be within the scope of the present application and form different embodiments, although some embodiments described herein include certain features rather than other features included in other embodiments.

[0089] The various component embodiments of the present application can be implemented in hardware, or implemented in software modules running on one or more processors, or implemented in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some modules in the electronic device according to the embodiments of the present application. The present application can also be implemented as a device program (for example, a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present application can be stored on a computer readable medium, or can have the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0090] The above description is merely a specific implementation or explanation of the present application, and the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, and all such changes or replacements should be encompassed within the protection scope of the present application.

Claims

1. A wind turbine gearbox intelligent fault diagnosis method, characterized in that: include: Acquiring real-time operating data of state variables of a wind turbine gearbox, wherein the state variables include environmental parameters, grid parameters, and component state parameters of the wind turbine gearbox, and the real-time operating data includes operating data within a preset time period; repeatedly inputting the real-time operating data of the state variable into a pre-trained fault diagnosis model to obtain a plurality of initial fault diagnosis results, wherein the fault diagnosis model is constructed based on a Bayesian graph convolutional network, and the initial fault diagnosis results include a probability of each of a plurality of fault types existing in the wind turbine gearbox; The multiple initial fault diagnosis results are fused using cooperative game to obtain a final fault diagnosis result.

2. The method according to claim 1, characterized in that Before repeatedly inputting the real-time operating data of the state variables into the pre-trained fault diagnosis model, the method further includes: Screening the state variables based on the real-time operating data of the state variables to determine state variables with higher importance; The step of repeatedly inputting the real-time operation data of the state variables into the pre-trained fault diagnosis model includes: repeatedly inputting the real-time operation data of the state variables with higher importance into the pre-trained fault diagnosis model.

3. The method according to claim 2, characterized in that The screening of the state variables based on the real-time operation data of the state variables includes: Using LightGBM Gini Index and / or Spearman correlation analysis determines the importance of each state variable to screen the state variables.

4. The method according to claim 3, characterized in that The use of LightGBM Gini Index and / or Spearman correlation analysis to determine the importance of each state variable to screen the state variables, including: Based on LightGBM and the real-time operation data of the state variables, determine the relative value of each state variable to the multiple fault types. Gini Index, where Gini The index is positively correlated with the importance of the corresponding state variable; choose Gini The first preset number of state variables with larger indexes are regarded as state variables with higher importance; or, The use of LightGBM Gini Index and / or Spearman correlation analysis to determine the importance of each state variable to screen the state variables, including: Calculating the Spearman correlation between each state variable and other state variables based on the real-time operation data of the state variables; For any state variable among the state variables, when the Spearman correlation between the state variable and at least some of the other state variables is greater than or equal to a correlation threshold, the state variable is determined to be a state variable with higher importance.

5. The method according to claim 3, characterized in that The use of LightGBM Gini Index and / or Spearman correlation analysis to determine the importance of each state variable to screen the state variables, including: Based on LightGBM and the real-time operation data of the state variables, determine the relative value of each state variable to the multiple fault types. Gini Index, where Gini The index is positively correlated with the importance of the corresponding state variable; choose Gini The first preset number of state variables with larger exponents constitute a first state variable set; Calculating the Spearman correlation between each state variable and other state variables based on the real-time operation data of the state variables; For any state variable among the state variables, when the Spearman correlation between the state variable and at least some of the other state variables is greater than or equal to a correlation threshold, add the state variable to the second state variable set; Taking the intersection of the first state variable set and the second state variable set as a state variable with higher importance; or, The use of LightGBM Gini Index and / or Spearman correlation analysis to determine the importance of each state variable to screen the state variables, including: Calculating the Spearman correlation between each state variable and other state variables based on the real-time operation data of the state variables; For any state variable among the state variables, when the Spearman correlation between the state variable and at least some of the other state variables is greater than or equal to a correlation threshold, add the state variable to the third state variable set; Based on LightGBM and the real-time operation data of each state variable in the third state variable set, determine the relative value of each state variable in the third state variable set to the multiple fault types. Gini Index, where Gini The index is positively correlated with the importance of the corresponding state variable; choose Gini The first preset number of state variables with larger indexes are regarded as state variables with higher importance; or, Based on LightGBM and the real-time operation data of the state variables, determine the relative value of each state variable to the multiple fault types. Gini Index, where Gini The index is positively correlated with the importance of the corresponding state variable; choose Gini The first preset number of state variables with larger exponents constitute a fourth state variable set; calculating, based on the real-time operating data of each state variable in the fourth state variable set, a Spearman correlation between each state variable in the fourth state variable set and other state variables in the fourth state variable set; For any state variable in the fourth state variable set, when the Spearman correlation between the state variable and at least some of the other state variables in the fourth state variable set is greater than or equal to a correlation threshold, the state variable is determined to be a state variable with higher importance.

6. The method according to any one of claims 2 to 5, characterized in that: Before repeatedly inputting the real-time operation data of the state variables with higher importance into the pre-trained fault diagnosis model, the method further includes: The real-time operation data of the state variables with higher importance are normalized.

7. The method according to any one of claims 1 to 5, characterized in that The adopting cooperative game to fuse the multiple initial fault diagnosis results to obtain a final fault diagnosis result includes: For any one of the multiple initial fault diagnosis results, calculating a consistency correlation coefficient between the initial fault diagnosis result and the remaining fault diagnosis results, where the remaining fault diagnosis result is the sum of the other initial fault diagnosis results in the multiple initial fault diagnosis results except the initial fault diagnosis result; The final fault diagnosis result is determined based on the sum of the products of the multiple initial fault diagnosis results and their corresponding consistency correlation coefficients.

8. The method according to any one of claims 1 to 4, characterized in that Before repeatedly inputting the real-time operating data of the state variables into the pre-trained fault diagnosis model, the method further includes: The real-time operation data of the state variables are processed using a linear interpolation method.

9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program, and the processor is configured to execute the computer program to implement the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that A computer program / instruction is stored, and when the computer program / instruction is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Fault diagnosis method and device for wind turbine generator gearbox

    CN114036656A

  • Wind turbine generator gearbox fault diagnosis system and method, computer equipment and storage medium

    CN114417514A

  • Wind turbine generator gearbox fault diagnosis method based on ICEEMDAN-MPE-RF and SVM

    CN116451105A

  • Gearbox oil temperature fault early warning method based on multilayer perception neural network

    CN117034754A

  • Fusion signal hydraulic motor fault diagnosis method and system

    CN117628005A