GMM-L2 fault diagnosis method and device based on DBGSA algorithm, equipment and medium
By adopting the GMM-L2 fault diagnosis method based on DBGSA algorithm in wind power engine fault detection, the problem of insufficient comprehensive and accurate fault detection in the prior art is solved, and more effective detection and diagnosis of wind power engine faults is achieved.
Patent Information
- Application Number
- CN202510177928.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-13
AI Technical Summary
The existing wind power engine fault detection algorithm can only judge a type of fault based on real-time monitored data, and cannot comprehensively and accurately monitor the faulty parts, increasing the risk of missed detection.
The GMM-L2 fault diagnosis method based on the DBGSA algorithm is adopted, and the original data of the wind power engine is obtained for preprocessing and data interpolation, the first and second Gaussian mixed models are generated, the health degree is calculated and the alarm limit is set, and the data is cleaned and dimensionality is reduced. The fuzzy clustering DBGSA algorithm is used to update the cluster center and calculate the membership probability of the fault type.
By improving the DBGSA algorithm, avoiding falling into local optimality and adding fuzzy membership weights can make the data points belong to multiple fault types at the same time, reducing the risk of missed detection, and achieving more comprehensive and accurate detection of wind power engine failures.
Smart Images

Figure CN119989234A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of engine fault detection, and in particular to a GMM-L2 fault diagnosis method, device, equipment and medium based on a DBGSA algorithm. Background Art
[0002] As a core power source, wind turbine engines have extremely high requirements for safety and reliability, but they are relatively expensive in design, production and maintenance costs. Due to long-term operation in harsh environments such as high temperature and high load, engine performance gradually deteriorates, increasing the risk of failure. In existing solutions, algorithms can be used to monitor data obtained from wind turbine engines to determine the location of wind turbine engine failures.
[0003] However, existing algorithms can only determine one type of fault based on real-time monitoring data. Due to this limitation of the algorithm, the faulty part of the wind turbine engine often cannot be fully and accurately monitored, thereby increasing the risk of missed detection. Summary of the invention
[0004] The present application provides a GMM-L2 fault diagnosis method, apparatus, device and medium based on the DBGSA algorithm, which avoids missed detection by improving the algorithm.
[0005] In order to achieve the above purpose, this application adopts the following technical solutions: In a first aspect, the present application provides a GMM-L2 fault diagnosis method based on a DBGSA algorithm, the method comprising: acquiring original data of a wind turbine engine; Preprocessing the original data of the wind turbine engine and removing abnormal data points to obtain first data; Performing data interpolation on the first data to obtain second data; Select k data points in the second data as initial cluster centers, calculate the Euclidean distance of each data point in the second data to all initial cluster centers, and assign each data point to the cluster with the closest Euclidean distance to the initial cluster center to form k clusters, where k is a positive integer; Calculate the cluster center of each cluster during the t-th iteration, where t is the first iteration number and t is a positive integer; Determine whether the distance between the cluster center of each cluster in the t-th iteration and the cluster center of each cluster in the t-1-th iteration is less than a first distance threshold, and obtain a first determination result; If the first judgment result indicates that the distance between the location of the cluster center of each cluster in the t-th iteration and the location of the cluster center of each cluster in the t-1-th iteration is less than the first distance threshold, the iteration is stopped, and the cluster center of each cluster in the t-th iteration is determined as the first cluster center; The expectation maximization algorithm is used to update the first cluster center of each cluster to obtain the second cluster center, and the first Gaussian mixture model is generated according to the second cluster center; The second Gaussian mixture model is generated based on the operating data of the wind turbine engine three weeks after maintenance; Calculating health according to the first Gaussian mixture model and the second Gaussian mixture model; Setting an alarm limit according to the health level, determining whether the second data is less than the alarm limit, and if the second data is less than the alarm limit, determining that the second data is in an abnormal state; Performing data cleaning and dimensionality reduction on the second data in an abnormal state to obtain third data; The DBGSA algorithm of fuzzy clustering is used to update the second cluster center on the third data to obtain the third cluster center; A membership threshold of the fault type is set, and the membership probability of the fault type to which each cluster belongs is calculated to obtain the fault type to which the third data belongs.
[0006] In some possible implementations, performing data interpolation on the first data to obtain the second data includes:
[0007] in, represents the random noise matrix, represents the mask matrix, represents the first data matrix, represents the data matrix at the vacant positions in the first data matrix, is the generator network, is the data interpolation at the vacant positions in the first data matrix, is the second data matrix.
[0008] In some possible implementations, the method further includes: If the first judgment result indicates that the distance between the cluster center of each cluster during the t-th iteration and the cluster center of each cluster during the t-1-th iteration is not less than the first distance threshold, then the distance from each data point to each cluster center during the t-th iteration is calculated, and the clusters to which the data points belong are redistributed.
[0009] In some possible implementations, the method of updating the first cluster center of each cluster by using the expectation maximization algorithm to obtain the second cluster center includes:
[0010] in, Represents data points The posterior probability of belonging to the i-th cluster, J represents the total number of data points in all clusters, , represents the mean of the first cluster center of the i-th cluster, i is a positive integer, and n represents the number of mixed components of the Gaussian mixture model, that is, the number of clusters. represents the covariance matrix of the first cluster center of the i-th cluster, represents the jth data point, represents the weight of the i-th cluster, represents the probability density function of the ith cluster.
[0011] In some possible implementations, the method further includes: If the second judgment result indicates that the second data is greater than or equal to the alarm limit, it is determined that the second data is in a normal state, and the fault type to which the second data belongs is no longer determined.
[0012] In some possible implementations, the updating of the second cluster center by using the DBGSA algorithm of fuzzy clustering on the third data to obtain the third cluster center includes:
[0013]
[0014]
[0015]
[0016]
[0017] in, represents the fuzzy membership of the jth data point in the third data to the third cluster center in the i-th cluster, represents the fuzzy membership weight, represents the third cluster center of the i-th cluster, represents the third cluster center of the kth cluster, m is the fuzzy index used to control the fuzziness of the membership, a is the adjustment parameter, represents the difference between the data point x and the cluster center c, Represents the jth data point To the first cluster center of the i-th cluster The difference between Represents the jth data point To the first cluster center of the kth cluster The difference between represents a convex function of a data point x, represents the convex function of the cluster center c, Representation function The gradient at c, Represents data points The importance weight of .
[0018] In some possible implementations, calculating the health level according to the first Gaussian mixture model and the second Gaussian mixture model includes:
[0019]
[0020]
[0021] in, Indicates health, represents the first Gaussian mixture model, represents the second Gaussian mixture model.
[0022] In a second aspect, the present application provides a GMM-L2 fault diagnosis device based on a DBGSA algorithm, the device comprising: The acquisition module is used to acquire the original data of the wind turbine engine; pre-process the original data of the wind turbine engine, remove abnormal data points, and obtain first data; perform data interpolation on the first data to obtain second data; A judgment module is used to select k data points in the second data as initial cluster centers, calculate the Euclidean distance of each data point in the second data to all initial cluster centers, and assign each data point to the cluster with the closest Euclidean distance to the initial cluster center to form k clusters, where k is a positive integer; calculate the cluster center of each cluster during the t-th iteration, t is the first iteration number, and t is a positive integer; judge whether the distance between the cluster center of each cluster during the t-th iteration and the location of the cluster center of each cluster during the t-1-th iteration is less than a first distance threshold, and obtain a first judgment result; if the first judgment result indicates that the distance between the cluster center of each cluster during the t-th iteration and the location of the cluster center of each cluster during the t-1-th iteration is less than the first distance threshold, then stop the iteration and determine the cluster center of each cluster during the t-th iteration as the first cluster center; A calculation module, used for updating the first cluster center of each cluster by using an expectation maximization algorithm to obtain a second cluster center, and generating a first Gaussian mixture model according to the second cluster center; generating a second Gaussian mixture model based on the operating data of the wind turbine engine three weeks after maintenance; calculating the health degree according to the first Gaussian mixture model and the second Gaussian mixture model; setting an alarm limit according to the health degree, judging whether the second data is less than the alarm limit, and determining that the second data is in an abnormal state if the second data is less than the alarm limit; performing data cleaning and dimensionality reduction on the second data in the abnormal state to obtain third data; The classification module is used to update the second cluster center of the third data using the DBGSA algorithm of fuzzy clustering to obtain the third cluster center; set the membership threshold of the fault type, calculate the membership probability of the fault type to which each cluster belongs, and obtain the fault type to which the third data belongs.
[0023] In a third aspect, the present application provides a computing device, including a memory and a processor; One or more computer programs are stored in the memory, and the one or more computer programs include instructions; when the instructions are executed by the processor, the computing device executes the method as described in any one of the first aspects.
[0024] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method as described in any one of the first aspects.
[0025] In a fifth aspect, the present application provides a computer program product, wherein the computer program product comprises one or more computer instructions, and when the computer instructions are executed by a computer, the computer executes the method as described in any one of the first aspects.
[0026] It can be seen from the above technical solution that the present application has at least the following beneficial effects: In the present application, the original data of the wind turbine engine is obtained; the original data of the wind turbine engine is preprocessed and interpolated to obtain the second data; then the first cluster center in the second data is calculated to generate the first Gaussian mixture model, and then the second Gaussian mixture model is generated based on the data obtained when the wind turbine engine is not faulty, and the health is calculated according to the two models; the alarm limit is set according to the health, and it is determined whether the second data is less than the alarm limit. If the second data is less than the alarm limit, it is determined that the second data is in an abnormal state; the second data in the abnormal state is cleaned and dimensionally reduced to obtain the third data; the DBGSA algorithm of fuzzy clustering is used to update the second cluster center for the third data to obtain the third cluster center; the membership threshold of the fault type is set, and the membership probability of the fault type to which each cluster belongs is calculated to obtain the fault type to which the third data belongs. Therefore, the present application continuously adjusts the position of the cluster center by improving the DBGSA algorithm to avoid falling into the local optimum, and also adds fuzzy membership weights, so that a data point can belong to multiple fault types at the same time, avoiding only classifying a data point as one fault and causing missed detection.
[0027] It should be understood that the description of technical features, technical solutions, beneficial effects or similar language in this application does not imply that all features and advantages can be realized in any single embodiment. On the contrary, it is understood that the description of features or beneficial effects means that specific technical features, technical solutions or beneficial effects are included in at least one embodiment. Therefore, the description of technical features, technical solutions or beneficial effects in this specification does not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions and beneficial effects described in the present embodiment can also be combined in any appropriate manner. Those skilled in the art will understand that the embodiment can be realized without one or more specific technical features, technical solutions or beneficial effects of a specific embodiment. In other embodiments, additional technical features and beneficial effects can also be identified in a specific embodiment that does not embody all embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 A flowchart of a GMM-L2 fault diagnosis method based on a DBGSA algorithm provided in an embodiment of the present application; Figure 2 A schematic diagram of a GMM-L2 fault diagnosis device based on a DBGSA algorithm provided in an embodiment of the present application; Figure 3 A schematic diagram of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] The terms "first", "second", "third", etc. in the specification of this application and the accompanying drawings are used to distinguish different objects rather than to limit a specific order.
[0030] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0031] In order to make the description of the following embodiments clear and concise, a brief introduction to the related technology is first given: The GMM algorithm, or Gaussian Mixture Model algorithm, is a commonly used clustering algorithm. The GMM algorithm assumes that the data set is generated by a mixture of multiple Gaussian distributions (also called "components" or "clusters"). Each Gaussian distribution represents a category or subgroup, and the data points are generated by a mixture of these Gaussian distributions with a certain probability. The goal of the GMM algorithm is to determine which category a data point belongs to by estimating the parameters of each Gaussian distribution (such as the mean vector, covariance matrix, and mixing weights).
[0032] The GMM algorithm usually uses the expectation-maximization (EM) algorithm for parameter estimation. The EM algorithm includes two steps: first, based on the current parameter estimate, calculate the probability that each data point belongs to each category (i.e., the posterior probability); second, use the probability of these data points to update the mean vector, covariance matrix, and mixing weights of each category. By continuously iterating these two steps, the GMM algorithm can gradually optimize these parameters until convergence.
[0033] The DGBSA algorithm, namely, a Batch Job Scheduling Algorithm based on a Genetic Algorithm with regard to the Threshold Detector, is an algorithm for optimizing batch job scheduling.
[0034] The DGBSA algorithm combines the ideas of genetic algorithm (GA) and threshold detector to solve the batch job scheduling problem. Genetic algorithm is an optimization algorithm that simulates natural selection and genetic mechanism. Through operations such as selection, crossover and mutation, it continuously iterates and optimizes individuals in the solution space to find the optimal solution or approximate optimal solution. The threshold detector is used to screen and evaluate individuals according to certain threshold conditions during the execution of the algorithm to further improve the performance of the algorithm.
[0035] At present, the detection method based on mathematically driven models can detect and predict the health status of wind turbines by establishing data models. However, the complex data and high-dimensional characteristics require the results of these models to be displayed through visualization technology. Commonly used visualization methods include time series diagrams, scatter plots, heat maps, etc., to help quickly identify abnormal trends and potential failure risks. Since the health status of wind turbines involves multiple parameters, the amount of data is huge and changes rapidly, traditional methods are difficult to display the complex relationship between data, and real-time data streams also face performance and accuracy challenges.
[0036] However, the GMM algorithm has some limitations, such as sensitivity to initial parameters: different initializations may lead to different results; computational complexity: especially for high-dimensional data or large-scale data, calculating the covariance matrix and probability density may be very complicated; local optimality: the EM algorithm may converge to a local optimal solution, which requires multiple runs to avoid; data distribution assumptions: if the data deviates significantly from the Gaussian distribution assumption, the GMM algorithm may perform poorly.
[0037] The existing DBGSA algorithm can predict the data of wind turbine engines and obtain a fault type, but the data anomaly is usually caused by multiple fault types. Therefore, it is necessary to improve the DBGSA algorithm to predict the data of wind turbine engines in order to predict multiple fault types on the data of wind turbine engines.
[0038] Therefore, this application generates an online monitoring and fault diagnosis model for the health status of wind turbines by integrating the improved DBGSA algorithm with the GMM algorithm, thereby realizing health monitoring of wind turbines and prediction of fault types.
[0039] In view of this, an embodiment of the present application provides a GMM-L2 fault diagnosis method based on the DBGSA algorithm, in which the original data of the wind turbine engine is first obtained; the original data of the wind turbine engine is preprocessed and data interpolated to obtain second data; then the first cluster center in the second data is calculated to generate a first Gaussian mixture model, and then a second Gaussian mixture model is generated based on the data obtained when the wind turbine engine has no fault, and the health is calculated according to the two models; an alarm limit is set according to the health, and it is determined whether the second data is less than the alarm limit. If the second data is less than the alarm limit, it is determined that the second data is in an abnormal state; data cleaning and dimensionality reduction are performed on the second data in the abnormal state to obtain third data; the second cluster center is updated by using the fuzzy clustering DBGSA algorithm for the third data to obtain a third cluster center; a membership threshold of the fault type is set, and the membership probability of the fault type to which each cluster belongs is calculated to obtain the fault type to which the third data belongs. Therefore, the present application improves the DBGSA algorithm, continuously adjusts the position of the cluster center to avoid falling into the local optimum, and also adds fuzzy membership weights, so that a data point can belong to multiple fault types at the same time, avoiding classifying a data point as only one fault and causing missed detection.
[0040] In order to make the technical solution of the present application clearer and easier to understand, the following describes a GMM-L2 fault diagnosis method based on the DBGSA algorithm provided by an embodiment of the present application in conjunction with the accompanying drawings. Figure 1 As shown, this figure is a flow chart of a GMM-L2 fault diagnosis method based on the DBGSA algorithm provided in an embodiment of the present application. The method is applied to a processing device, and the GMM-L2 fault diagnosis method based on the DBGSA algorithm includes: S101. A processing device obtains original data of a wind turbine.
[0041] The processing equipment obtains the real-time data of characteristic parameters of the wind turbine from the SCADA system as the original data of the wind turbine for subsequent processing.
[0042] S102: The processing device pre-processes the original data of the wind turbine engine and obtains first data after removing abnormal data points.
[0043] The processing device uses the DBSCAN algorithm and the local outlier factor LOF algorithm to pre-process the extracted original data of the wind turbine engine, removes abnormal data points, and obtains the first data.
[0044] The local outlier factor LOF algorithm determines the degree of anomaly by comparing the density near the sample object with the density of the neighborhood. The LOF value of a data point is:
[0045] in, represents the kth local outlier factor of each point, represents a data point, represents the kth local reachable density of point o, represents the kth local reachable density of point p, represents the k-th distance neighborhood.
[0046] The DBSCAN algorithm (Density-Based Spatial Clustering of Applications with Noise) is a density-based spatial clustering algorithm that can find clusters of any shape in a noisy data set. The DBSCAN algorithm determines whether a point is a core point by calculating the number of points contained in the ε neighborhood of each point, and then expands the cluster. If the ε neighborhood of a point contains enough points (that is, it reaches or exceeds the preset MinPts threshold), it is regarded as a core point, and the points in its neighborhood are continued to be expanded to form a cluster. The DBSCAN algorithm is robust to noise points, can identify and exclude noise, and is applicable to data distributions of various shapes and densities. It is one of the commonly used clustering algorithms in the field of data mining and machine learning.
[0047] S103: The processing device performs data interpolation on the first data to obtain second data.
[0048] The processing device uses a lightweight generative adversarial network (AT-SGAIN) based on an adaptive attention mechanism to interpolate the missing data of SCADA during the collection process. First, a transformer encoder layer is added to the generator and discriminator of SGAIN, and then the transformer encoder layer is improved to add spatial attention and channel attention mechanisms. The input is divided into a random noise matrix, an initial data matrix, and a mask matrix. The initial data matrix is the first data after data cleaning and outlier removal, called X. Assume , n is the number of samples, d is the feature dimension. The mask matrix is , where 1 refers to non-missing data and 0 refers to missing data. Random noise matrix , we let each element of N be independently sampled from a uniform distribution in the interval [-0.01, 0.01], and .
[0049] Therefore, data interpolation of the first data can be achieved through the following formula:
[0050] in, represents the random noise matrix, represents the mask matrix, represents the first data matrix, represents the data matrix at the vacant positions in the first data matrix, is the generator network, is the data interpolation at the vacant positions in the first data matrix, is the second data matrix.
[0051] The AT-SGAIN model is mainly implemented by combining a lightweight generative adversarial network (SGAIN) and an adaptive Transformer attention mechanism. The model integrates the Transformer encoder in the generator and the discriminator, which can effectively capture long-term time series characteristics and adapt to dynamic changes in data. In addition, a parallel channel and spatial attention mechanism is designed, which can accurately adjust the attention allocation through adaptive weight coefficients, enhance the model's ability to capture local information, and maintain sensitivity to time series features. It can generate more realistic interpolation values that conform to data distribution and improve the accuracy of interpolation. S104. The processing device selects k data points in the second data as initial cluster centers, calculates the Euclidean distance of each data point in the second data to all initial cluster centers, and assigns each data point to the cluster with the shortest Euclidean distance to the initial cluster center to form k clusters.
[0052] The processing device randomly selects k data points in the second data as initial cluster centers, and calculates the Euclidean distance from each data point in the second data to all initial cluster centers. The calculation formula of the Euclidean distance is as follows:
[0053] in, represents the Euclidean distance between the jth data point and the cluster center of the i-th cluster, represents the jth data point, Represents the cluster center of the i-th cluster.
[0054] Then assign each data point to the group with the closest Euclidean distance to the initial cluster center to form k clusters.
[0055] S105: The processing device calculates the cluster center of each cluster in the t-th round of iteration.
[0056] In step S104, the processing device obtains k clusters, each of which contains J data points, so the processing device will recalculate the cluster center in each cluster. During the iteration process, the processing device calculates the cluster center of each cluster in the tth round of iteration, where t is the first iteration number and t is a positive integer.
[0057] S106: The processing device determines whether the distance between the location of the cluster center of each cluster in the t-th iteration and the location of the cluster center of each cluster in the t-1-th iteration is less than a first distance threshold, and obtains a first determination result.
[0058] If the first judgment result indicates that the distance between the location of the cluster center of each cluster in the t-th iteration and the location of the cluster center of each cluster in the t-1-th iteration is less than the first distance threshold, then execute S107; If the first judgment result indicates that the distance between the location of the cluster center of each cluster during the t-th iteration and the location of the cluster center of each cluster during the t-1-th iteration is not less than the first distance threshold, S115 is executed.
[0059] S107 , stopping the iteration, and determining the cluster center of each cluster in the t-th round of iteration as the first cluster center.
[0060] The calculation formula in the first cluster is as follows, and the covariance of the first cluster center is calculated:
[0061]
[0062] in, represents the jth data point, represents the cluster center of the i-th cluster, The number of data points represented by Represents the covariance matrix of the cluster centers of the ith cluster.
[0063] S108: The processing device updates the first cluster center of each cluster using an expectation-maximization algorithm to obtain a second cluster center, and generates a first Gaussian mixture model according to the second cluster center.
[0064] The processing device updates the first cluster center of each cluster using an expectation maximization algorithm to obtain a second cluster center, including:
[0065] in, Represents data points The posterior probability of belonging to the i-th cluster, J represents the total number of data points in all clusters, , represents the mean of the first cluster center of the i-th cluster, i is a positive integer, and n represents the number of mixed components of the Gaussian mixture model, that is, the number of clusters. represents the covariance matrix of the first cluster center of the i-th cluster, represents the jth data point, represents the weight of the i-th cluster, represents the probability density function of the ith cluster.
[0066] The processing device generates a first Gaussian mixture model according to the first cluster center and the AIC criterion.
[0067] Among them, AIC (Akaike Information Criterion) is the Akaike information criterion, which is used to evaluate GMM models with different numbers of components, select an optimal model, avoid overfitting or underfitting, and select the optimal number of mixture components n (the value of n when AIC is the smallest).
[0068]
[0069] Among them, k is the number of parameters of the model. For the Gaussian mixture model (GMM), it is the total number of all parameters in the model, including the mean, covariance matrix and weight of each Gaussian component, which is a penalty for model complexity. The 2k term in AIC penalizes the situation where the model has too many parameters to avoid overly complex models. L is the maximum likelihood function.
[0070] Assuming the data is d-dimensional, then
[0071] Get the first Gaussian mixture model formula
[0072] in, represents the first Gaussian mixture model, represents the weight of the mixed component of the first Gaussian mixture model, n represents the number of mixed components, represents the probability density function of the multivariate Gaussian distribution, represents the mean vector of the mixture component a, Represents the covariance matrix of the mixture component a.
[0073] S109: The processing device generates a second Gaussian mixture model based on the operating data of the wind turbine engine three weeks after maintenance.
[0074] The operating data of the wind turbine engine three weeks after maintenance is used as the baseline state data to generate the second Gaussian mixture model:
[0075] in, represents the second Gaussian mixture model, represents the weight of the mixed component of the first Gaussian mixture model, n represents the number of mixed components, represents the probability density function of the multivariate Gaussian distribution, represents the mean vector of the mixture component b, Represents the covariance matrix of the mixture component b.
[0076] S110: The processing device calculates the health level according to the first Gaussian mixture model and the second Gaussian mixture model.
[0077] The processing equipment uses the GMM-L2 model to calculate the health. Compared with the traditional health evaluation based on a single indicator, this method can combine more dimensional features for comprehensive evaluation and consider the inherent laws of data distribution. The health calculation formula is as follows:
[0078]
[0079]
[0080] in, Indicates health, represents the first Gaussian mixture model, represents the second Gaussian mixture model.
[0081] The smaller the CV health degree is, the more serious the health decline of the wind turbine group is. On the contrary, the larger the CV is, the closer the current state of the wind turbine group is to the state at the reference time, indicating that there is almost no fault in the wind turbine group at this time.
[0082] S111. The processing device sets an alarm limit according to the health level, determines whether the second data is less than the alarm limit, and obtains a second determination result.
[0083] In the embodiment of the present application, the processing device calculates the health of data between multiple wind turbine engine normal operation conditions, arranges them from small to large, and takes the 5% quantile as the CV alarm limit.
[0084] If the second judgment result indicates that the second data is less than the alarm limit, the second data is determined to be in an abnormal state; if the second judgment result indicates that the second data is greater than or equal to the alarm limit, the second data is determined to be in a normal state, and the fault type to which the second data belongs is no longer determined.
[0085] S112: The processing device performs data cleaning and dimension reduction on the second data in an abnormal state to obtain third data.
[0086] The processing device uses DBSCAN and LOF algorithms to clean the second data in an abnormal state, and uses the spearman rank correlation coefficient and PCA (Principal Component Analysis) algorithm to reduce the dimension of the data to obtain the third data.
[0087] Among them, the Spearman rank correlation coefficient, also known as the Spearman correlation coefficient, is a non-parametric correlation coefficient used to measure the monotonic relationship between two variables; the PCA (Principal Component Analysis) algorithm is an unsupervised learning algorithm, mainly used for data dimensionality reduction and feature extraction.
[0088] S113: The processing device updates the second cluster center by using the DBGSA algorithm of fuzzy clustering on the third data to obtain a third cluster center.
[0089] The calculation process of the third cluster center is: First, Bregman divergence is used to measure the difference between point x and cluster center c. Bregman divergence is used to measure the asymmetric distance between two points, minimizing the total divergence from each data point to the cluster center in the entire data, which is defined as:
[0090] Each data point Cluster Center The fuzzy membership of is defined by the fuzzy membership function:
[0091]
[0092]
[0093]
[0094] in, represents the fuzzy membership of the jth data point in the third data to the third cluster center in the i-th cluster, represents the fuzzy membership weight, represents the third cluster center of the i-th cluster, represents the third cluster center of the kth cluster, m is the fuzzy index used to control the fuzziness of the membership, a is the adjustment parameter, represents the difference between the data point x and the cluster center c, Represents the jth data point To the first cluster center of the i-th cluster The difference between Represents the jth data point To the first cluster center of the kth cluster The difference between represents a convex function of a data point x, represents the convex function of the cluster center c, Representation function The gradient at c, Represents data points The importance weight of .
[0095] Using Bregman divergence as the distance metric, the distance metric can be adaptively adjusted according to the data distribution, so as to adapt to different data distributions (such as Gaussian distribution, exponential distribution, etc.). Compared with the traditional Euclidean distance, Bregman divergence performs better when processing non-uniformly distributed data. The dynamically optimized global search capability is combined with the Gravitational Search Algorithm (GSA), which makes the cluster center have stronger global search capabilities in the initial stage. The position of the cluster center is dynamically adjusted through the gravitational mechanism to avoid falling into the local optimum. Low parameter dependence DBGSA reduces the sensitivity to hyperparameters (such as initial cluster center, learning rate, etc.) and reduces the complexity of parameter adjustment through dynamic weight adjustment. At the same time, we added fuzzy membership weights to the DBGSA calculation formula, which allows the data of a point to belong to multiple classes at the same time, avoiding the omission of detection by classifying only one point as a fault.
[0096] S114: The processing device sets a membership threshold of the fault type, calculates the membership probability of the fault type to which each cluster belongs, and obtains the fault type to which the third data belongs.
[0097] The data collected when the wind turbine engine has a fault during operation is obtained and determined as historical data. The improved DBGSA algorithm is used to classify the historical data and determine the fault type.
[0098] According to the occurrence frequency of failures of different components of the wind turbine, the six groups of parameters that have the greatest impact on the health status of the wind turbine are determined to be power, controller temperature, gearbox oil temperature, gearbox bearing temperature, generator bearing 1 temperature and generator bearing 2 temperature. The spearman rank correlation coefficient is used to select the above parameters as the parameters of the feature input as the fault type distribution when the third data is in an abnormal state.
[0099] The processing device calculates the membership probability of the fault type to which each cluster in the third data belongs, and further obtains the fault type to which the third data belongs.
[0100] S115: The processing device calculates the distance from each data point to each cluster center during the t-th iteration, and reallocates the clusters to which the data points belong.
[0101] In the embodiment of the present application, a wind turbine engine fault type diagnosis model can be trained through the above steps, and the fault type prediction of real-time data can be realized through the model.
[0102] Based on the above content, this application improves the DBGSA algorithm, continuously adjusts the position of the cluster center to avoid falling into the local optimum, and also adds fuzzy membership weights, so that a data point can belong to multiple fault types at the same time, avoiding classifying a data point as only one fault and causing missed detection.
[0103] The present application also provides a GMM-L2 fault diagnosis device based on the DBGSA algorithm, such as Figure 2 As shown, the figure is a schematic diagram of a GMM-L2 fault diagnosis device based on the DBGSA algorithm provided in an embodiment of the present application, the device includes: an acquisition module 201, a judgment module 202, a calculation module 203 and a classification module 204; The acquisition module 201 is used to acquire the original data of the wind turbine engine; pre-process the original data of the wind turbine engine, remove abnormal data points, and obtain first data; perform data interpolation on the first data to obtain second data; The judgment module 202 is used to select k data points in the second data as initial cluster centers, calculate the Euclidean distance of each data point in the second data to all initial cluster centers, and assign each data point to the cluster with the closest Euclidean distance to the initial cluster center to form k clusters, where k is a positive integer; calculate the cluster center of each cluster during the t-th iteration, t is the first iteration number, and t is a positive integer; judge whether the distance between the cluster center of each cluster during the t-th iteration and the location of the cluster center of each cluster during the t-1-th iteration is less than a first distance threshold, and obtain a first judgment result; if the first judgment result indicates that the distance between the cluster center of each cluster during the t-th iteration and the location of the cluster center of each cluster during the t-1-th iteration is less than the first distance threshold, then stop the iteration and determine the cluster center of each cluster during the t-th iteration as the first cluster center; The calculation module 203 is used to update the first cluster center of each cluster using the expectation maximization algorithm to obtain the second cluster center, and generate a first Gaussian mixture model based on the second cluster center; generate a second Gaussian mixture model based on the operating data of the wind turbine engine three weeks after maintenance; calculate the health degree according to the first Gaussian mixture model and the second Gaussian mixture model; set an alarm limit according to the health degree, determine whether the second data is less than the alarm limit, and if the second data is less than the alarm limit, determine that the second data is in an abnormal state; perform data cleaning and dimensionality reduction on the second data in the abnormal state to obtain third data; The classification module 204 is used to update the second cluster center of the third data using the fuzzy clustering DBGSA algorithm to obtain the third cluster center; set the membership threshold of the fault type, calculate the membership probability of the fault type to which each cluster belongs, and obtain the fault type to which the third data belongs.
[0104] In some possible implementations, the acquisition module 201 is specifically configured to perform data interpolation on the first data to obtain the second data, including:
[0105] in, represents the random noise matrix, represents the mask matrix, represents the first data matrix, represents the data matrix at the vacant positions in the first data matrix, is the generator network, is the data interpolation at the vacant positions in the first data matrix, is the second data matrix.
[0106] In some possible implementations, the judgment module 202 is also used to calculate the distance from each data point to each cluster center during the tth round of iteration and reallocate the clusters to which the data points belong if the first judgment result indicates that the distance between the cluster center of each cluster during the tth round of iteration and the cluster center of each cluster during the t-1th round of iteration is not less than a first distance threshold.
[0107] In some possible implementations, the calculation module 203 is specifically configured to update the first cluster center of each cluster using an expectation maximization algorithm to obtain a second cluster center, including:
[0108] in, Represents data points The posterior probability of belonging to the i-th cluster, J represents the total number of data points in all clusters, , represents the mean of the first cluster center of the i-th cluster, i is a positive integer, and n represents the number of mixed components of the Gaussian mixture model, that is, the number of clusters. represents the covariance matrix of the first cluster center of the i-th cluster, represents the jth data point, represents the weight of the i-th cluster, represents the probability density function of the ith cluster.
[0109] In some possible implementations, the calculation module 203 is further used to determine that the second data is in a normal state if the second judgment result indicates that the second data is greater than or equal to the alarm limit, and no longer continue to judge the fault type to which the second data belongs.
[0110] In some possible implementations, the classification module 204 is specifically configured to update the second cluster center using the DBGSA algorithm of fuzzy clustering on the third data to obtain the third cluster center, including:
[0111]
[0112]
[0113]
[0114]
[0115] in, represents the fuzzy membership of the jth data point in the third data to the third cluster center in the i-th cluster, represents the fuzzy membership weight, represents the third cluster center of the i-th cluster, represents the third cluster center of the kth cluster, m is the fuzzy index used to control the fuzziness of the membership, a is the adjustment parameter, represents the difference between the data point x and the cluster center c, Represents the jth data point To the first cluster center of the i-th cluster The difference between Represents the jth data point To the first cluster center of the kth cluster The difference between represents a convex function of a data point x, represents the convex function of the cluster center c, Representation function The gradient at c, Represents data points The importance weight of .
[0116] In some possible implementations, the calculation module 203 is specifically configured to calculate the health level according to the first Gaussian mixture model and the second Gaussian mixture model, including:
[0117]
[0118]
[0119] in, Indicates health, represents the first Gaussian mixture model, represents the second Gaussian mixture model.
[0120] The present application also provides a computing device. Figure 3 As shown, this figure is a schematic diagram of a computing device provided in an embodiment of the present application, and the computing device 400 includes a bus 401, a processor 402, a communication interface 403 and a memory 404. The processor 402, the memory 404 and the communication interface 403 communicate with each other through the bus 401.
[0121] The bus 401 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0122] The processor 402 may be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0123] The communication interface 403 is used for communicating with the outside.
[0124] The memory 404 may include a volatile memory, such as a random access memory (RAM). The memory 404 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0125] The memory 404 stores executable codes, and the processor 402 executes the executable codes to perform the aforementioned GMM-L2 fault diagnosis method based on the DBGSA algorithm.
[0126] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-mentioned GMM-L2 fault diagnosis method based on the DBGSA algorithm.
[0127] The embodiment of the present application further provides a computer program product, which includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, the process or function described in the embodiment of the present application is generated in whole or in part.
[0128] The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer or data center to another website, computer or data center via wired (e.g., coaxial cable, optical fiber) or wireless (e.g., infrared, wireless, microwave, etc.) means.
[0129] When the computer program product is executed by a computer, the computer executes any one of the aforementioned GMM-L2 fault diagnosis methods based on the DBGSA algorithm. The computer program product may be a software installation package, and when any one of the aforementioned GMM-L2 fault diagnosis methods based on the DBGSA algorithm needs to be used, the computer program product may be downloaded and executed on a computer.
[0130] The descriptions of the processes or structures corresponding to the above-mentioned figures have different emphases. For parts that are not described in detail in a certain process or structure, please refer to the relevant descriptions of other processes or structures.
[0131] The above description is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be included in the protection scope of the present application.
Claims
1. A GMM-L2 fault diagnosis method based on DBGSA algorithm, characterized in that: The method comprises: Obtain the original data of wind turbines; Preprocessing the original data of the wind turbine engine and removing abnormal data points to obtain first data; Performing data interpolation on the first data to obtain second data; Select k data points in the second data as initial cluster centers, calculate the Euclidean distance of each data point in the second data to all initial cluster centers, and assign each data point to the cluster with the closest Euclidean distance to the initial cluster center to form k clusters, where k is a positive integer; Calculate the cluster center of each cluster during the t-th iteration, where t is the first iteration number and t is a positive integer; Determine whether the distance between the cluster center of each cluster in the t-th iteration and the cluster center of each cluster in the t-1-th iteration is less than a first distance threshold, and obtain a first determination result; If the first judgment result indicates that the distance between the location of the cluster center of each cluster in the t-th iteration and the location of the cluster center of each cluster in the t-1-th iteration is less than the first distance threshold, the iteration is stopped, and the cluster center of each cluster in the t-th iteration is determined as the first cluster center; The expectation maximization algorithm is used to update the first cluster center of each cluster to obtain the second cluster center, and the first Gaussian mixture model is generated according to the second cluster center; The second Gaussian mixture model is generated based on the three-week operation data of the wind turbine engine after maintenance; Calculating health according to the first Gaussian mixture model and the second Gaussian mixture model; Setting an alarm limit according to the health level, determining whether the second data is less than the alarm limit, and obtaining a second determination result; If the second judgment result indicates that the second data is less than the alarm limit, determining that the second data is in an abnormal state; Performing data cleaning and dimensionality reduction on the second data in an abnormal state to obtain third data; The DBGSA algorithm of fuzzy clustering is used to update the second cluster center on the third data to obtain the third cluster center; A membership threshold of the fault type is set, and the membership probability of the fault type to which each cluster belongs is calculated to obtain the fault type to which the third data belongs.
2. The method according to claim 1, characterized in that The step of performing data interpolation on the first data to obtain the second data includes: in, represents the random noise matrix, represents the mask matrix, represents the first data matrix, represents the data matrix at the vacant positions in the first data matrix, is the generator network, is the data interpolation at the vacant positions in the first data matrix, is the second data matrix.
3. The method according to claim 1, characterized in that The method further comprises: If the first judgment result indicates that the distance between the cluster center of each cluster during the t-th iteration and the cluster center of each cluster during the t-1-th iteration is not less than the first distance threshold, then the distance from each data point to each cluster center during the t-th iteration is calculated, and the clusters to which the data points belong are redistributed.
4. The method according to claim 1, characterized in that: The method of using the expectation maximization algorithm to update the first cluster center of each cluster to obtain the second cluster center includes: in, Represents data points The posterior probability of belonging to the i-th cluster, J represents the total number of data points in all clusters, , represents the mean of the first cluster center of the i-th cluster, i is a positive integer, and n represents the number of mixed components of the Gaussian mixture model, that is, the number of clusters. represents the covariance matrix of the first cluster center of the i-th cluster, represents the jth data point, represents the weight of the i-th cluster, represents the probability density function of the ith cluster.
5. The method according to claim 1, characterized in that The method further comprises: If the second judgment result indicates that the second data is greater than or equal to the alarm limit, it is determined that the second data is in a normal state, and the fault type to which the second data belongs is no longer determined.
6. The method according to claim 1, characterized in that The updating of the second cluster center by using the DBGSA algorithm of fuzzy clustering on the third data to obtain the third cluster center includes: in, represents the fuzzy membership of the jth data point in the third data to the third cluster center in the i-th cluster, represents the fuzzy membership weight, represents the third cluster center of the i-th cluster, represents the third cluster center of the kth cluster, m is the fuzzy index used to control the fuzziness of the membership, a is the adjustment parameter, represents the difference between the data point x and the cluster center c, Represents the jth data point To the first cluster center of the i-th cluster The difference between Represents the jth data point To the first cluster center of the kth cluster The difference between represents a convex function of a data point x, represents the convex function of the cluster center c, Representation function The gradient at c, Represents data points The importance weight of .
7. The method according to claim 1, characterized in that Calculating the health degree according to the first Gaussian mixture model and the second Gaussian mixture model includes: in, Indicates health, represents the first Gaussian mixture model, represents the second Gaussian mixture model.
8. A GMM-L2 fault diagnosis device based on DBGSA algorithm, characterized in that: The device comprises: The acquisition module is used to acquire the original data of the wind turbine engine; pre-process the original data of the wind turbine engine, remove abnormal data points, and obtain first data; perform data interpolation on the first data to obtain second data; A judgment module is used to select k data points in the second data as initial cluster centers, calculate the Euclidean distance of each data point in the second data to all initial cluster centers, and assign each data point to the cluster with the closest Euclidean distance to the initial cluster center to form k clusters, where k is a positive integer; calculate the cluster center of each cluster during the t-th iteration, t is the first iteration number, and t is a positive integer; judge whether the distance between the cluster center of each cluster during the t-th iteration and the location of the cluster center of each cluster during the t-1-th iteration is less than a first distance threshold, and obtain a first judgment result; if the first judgment result indicates that the distance between the cluster center of each cluster during the t-th iteration and the location of the cluster center of each cluster during the t-1-th iteration is less than the first distance threshold, then stop the iteration and determine the cluster center of each cluster during the t-th iteration as the first cluster center; A calculation module, used for updating the first cluster center of each cluster by using an expectation maximization algorithm to obtain a second cluster center, and generating a first Gaussian mixture model according to the second cluster center; generating a second Gaussian mixture model based on the operating data of the wind turbine engine three weeks after maintenance; calculating the health degree according to the first Gaussian mixture model and the second Gaussian mixture model; setting an alarm limit according to the health degree, judging whether the second data is less than the alarm limit, and determining that the second data is in an abnormal state if the second data is less than the alarm limit; performing data cleaning and dimensionality reduction on the second data in the abnormal state to obtain third data; The classification module is used to update the second cluster center of the third data using the fuzzy clustering DBGSA algorithm to obtain the third cluster center; set the membership threshold of the fault type, calculate the membership probability of the fault type to which each cluster belongs, and obtain the fault type to which the third data belongs.
9. A computing device, characterized in that including memory and processor; One or more computer programs are stored in the memory, and the one or more computer programs include instructions; when the instructions are executed by the processor, the computing device executes the method as claimed in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method according to any one of claims 1 to 7.
Citation Information
Cited By
Processing data quality monitoring system and method based on machine learning
CN120541730A
A Machine Learning-Based Processing Data Quality Monitoring System and Method
CN120541730B
Bearing fault diagnosis method and device based on unsupervised learning and medium
CN120597131A
Wind energy data quality control method and device based on clustering algorithm
CN121834384A