A smart early warning platform for the entire lifecycle of wind turbines based on big data analytics
The intelligent early warning platform for the entire life cycle of wind turbines, which utilizes big data analysis, solves the problem that static models cannot adapt to dynamic evolution. It achieves highly reliable early warning throughout the entire life cycle of wind turbines, and can identify abnormal changes and reduce false alarm rates.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-04-03
AI Technical Summary
Existing wind turbine full life cycle early warning systems suffer from insufficient early warning reliability due to the inability of static models to adapt to dynamic evolution, making it difficult to meet the fault requirements of wind turbines at different life cycles.
The wind turbine intelligent early warning platform based on big data analysis determines the structured labels of the basic physical model through the physical theory module, generates structured semantic descriptions through the operation status module, calculates the degree of correlation and correlation similarity through the evolution determination module, and corrects the data through the fault early warning module to generate target early warning data to adapt to the dynamic evolution of wind turbines.
It improves the reliability of early warning throughout the entire life cycle of wind turbine units, can distinguish between normal evolution and abnormal changes, reduces the false alarm rate, and ensures that early warning data is consistent with the current state.
Smart Images

Figure CN121280001B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of wind turbine generators, and in particular to an intelligent early warning platform for the entire life cycle of wind turbine generators based on big data analysis. Background Technology
[0002] Currently, the wind power industry widely adopts data-driven modeling methods in the field of predictive maintenance to provide early warnings of faults by analyzing massive amounts of operational data, such as wind turbine data acquisition and monitoring data (SCADA, Supervisory Control and Data Acquisition).
[0003] However, wind turbines undergo dynamic evolution over time, rendering static models inadequate for meeting the early warning needs of dynamically evolving wind turbines. In particular, the lifespan of a wind turbine is quite long, making it difficult for pre-trained models to meet the early warning requirements throughout its entire lifecycle. Specifically, as wind turbines age, gearboxes experience efficiency decline due to wear, and blades undergo aerodynamic performance changes due to increased surface roughness. These factors cause systematic shifts in the normal operating parameters of the wind turbine. Further complicating matters, the dominant failure modes evolve across different lifecycle stages: early failures are primarily due to manufacturing or installation defects, mid-stage failures are dominated by fatigue damage, and later failures are primarily due to material aging and cumulative wear.
[0004] Faced with a wind turbine that changes over time, common practices often fall into a static processing framework. The typical development process is to collect a historical dataset for a specific time period, train it to obtain a model with optimal performance, and then deploy it online for continuous monitoring. This makes the model unable to adapt to early warnings throughout the entire life cycle of the wind turbine.
[0005] Therefore, improving the reliability of early warning for the entire life cycle of wind turbines has become an urgent technical problem to be solved. Summary of the Invention
[0006] The technical problem solved by this invention is the insufficient reliability of early warning for the entire life cycle of wind turbine units.
[0007] To address the aforementioned technical problems, this invention provides the following technical solution: a wind turbine intelligent early warning platform based on big data analysis throughout its entire lifecycle, comprising: a physical theory module for determining structured labels for each basic physical model of the wind turbine; an operation status module for acquiring real-time operation data streams of the wind turbine and generating structured semantic descriptions based on the real-time operation data streams; an evolution determination module for: determining the degree of correlation between each structured semantic description and each structured label to obtain the current correlation matrix; determining the correlation similarity between the data distribution of the current correlation matrix and the historical correlation matrix; and determining the correction coefficient for each basic physical model based on the correlation similarity; and a fault early warning module for: determining the original early warning data for each basic physical model based on the real-time operation data stream; and correcting each original early warning data according to the correction coefficient for each basic physical model to obtain target early warning data.
[0008] Preferably, generating a structured semantic description based on the real-time running data stream includes: marking data streams in the real-time running data stream whose variance is less than a preset variance threshold within a preset time window as quasi-steady-state operating condition data segments; determining the autocorrelation function of the fault data sequence in the quasi-steady-state operating condition data segment; recording the time length required for the autocorrelation function to first decay to the preset variance threshold as the dynamic window length under the quasi-steady-state operating condition data segment; truncating the quasi-steady-state operating condition data segment based on the dynamic window length to obtain the fault data sequence; reconstructing the phase space of the fault data sequence to obtain a high-dimensional phase space point cloud; generating a persistent graph based on the continuous cohomology of the high-dimensional phase space point cloud; extracting topological invariants from the persistent graph to obtain the actual topological feature vector; and determining the structured semantic description based on the actual topological feature vector.
[0009] Preferably, the phase space reconstruction of the fault data sequence to obtain a high-dimensional phase space point cloud includes: determining the mutual information function and permutation entropy function of the fault data sequence under multiple preset delay times; determining the first local minimum point in the permutation entropy function, wherein the first local minimum point is the smallest independent variable value among the multiple local minimum points of the permutation entropy function; determining the second local minimum point in the mutual information function within the domain from the preset initial value to the first local minimum point, wherein the second local minimum point is the smallest independent variable value among the multiple local minimum points of the mutual information function; and then reconstructing the second local minimum point into a higher-dimensional phase space point cloud. The minimum value is determined as the optimal time delay parameter; the signal-to-noise ratio (SNR) of the fault data sequence is determined; the adaptive threshold of the fault data sequence is determined based on the SNR, wherein the adaptive threshold is negatively correlated with the SNR; starting from the preset minimum embedding dimension, the correlation dimension under the current embedding dimension and the relative rate of change of the correlation dimension compared with the previous correlation dimension are determined; the current embedding dimension when the relative rate of change is less than the adaptive threshold is determined as the optimal embedding dimension; the phase space of the fault data sequence is reconstructed based on the optimal time delay parameter and the optimal embedding dimension to obtain a high-dimensional phase space point cloud.
[0010] Preferably, generating a persistent graph based on the continuous cohomology of a high-dimensional phase space point cloud includes: selecting a preset number of landmark point sets from the high-dimensional phase space point cloud using a preset maximum-minimum distance method; determining potential simplexes in the landmark point sets; for each potential simplex: determining witness points corresponding to the potential simplex from the high-dimensional phase space point cloud; determining the set formed by removing potential simplexes from the landmark point set as the landmark point difference set corresponding to the potential simplex; if the distance from the witness point to the vertex of the potential simplex corresponding to the witness point is less than the distance from the witness point to any point in the landmark point difference set corresponding to the potential simplex, then adding the potential simplex to the Witness complex; constructing the Witness complex based on each potential simplex, and determining the continuous cohomology of the Witness complex to obtain a persistent graph.
[0011] Preferably, extracting topological invariants from the persistent graph to obtain the actual topological feature vector includes: setting grid partitioning rules for the persistence dimension and birth time dimension of the persistent graph to obtain multiple grids, wherein the persistent graph is a two-dimensional point set, each point in the two-dimensional point set represents a topological feature, the horizontal axis of the coordinate system corresponding to the two-dimensional point set is the birth time, the vertical axis is the persistence, and the persistence is the difference between the death time and the birth time of the topological feature; selecting topological features corresponding to the first homology group from the persistent graph, wherein the topological features of the first homology group are used to characterize the one-dimensional ring structure in the data; for each topological feature on the first homology group in the persistent graph, determining the grid coordinates to which the topological feature belongs based on the birth value and death value of the topological feature; counting the number of topological features falling into each grid in the persistent graph; and arranging the statistics in the grid to obtain the actual topological feature vector.
[0012] Preferably, the real-time operating data stream includes at least one of a vibration data stream, a temperature data stream, a pressure data stream, and a power data stream. Before marking the data streams in the real-time operating data stream whose variance is less than a preset variance threshold within a preset time window as quasi-steady-state operating condition data segments, the operating status module is further configured to: perform spectral analysis on the vibration data stream in the real-time operating data stream, and determine the first type of fault data sequence of the vibration data stream in the real-time operating data stream based on the energy value of a preset frequency band; perform rate of change analysis on the temperature data stream and pressure data stream in the real-time operating data stream, and determine the second type of fault data sequence of the temperature and pressure data streams in the real-time operating data stream based on the abnormal rate of change; perform efficiency curve analysis on the power data stream in the real-time operating data stream, and determine the third type of fault data sequence of the power data stream in the real-time operating data stream based on the deviation of the power curve; and determine a structured semantic description based on the actual topological feature vector, including: determining the structured semantic description based on the actual topological feature vectors corresponding to the first type of fault data sequence, the second type of fault data sequence, and the third type of fault data sequence.
[0013] Preferably, determining the structured label of each basic physical model of the wind turbine includes: determining the expected fault sequence of the basic physical model based on the fault modes of the basic physical model; determining the expected topological feature vector of the basic physical model based on the expected fault sequence; determining the structured label of the basic physical model based on the components, subsystems, fault modes, and expected topological feature vectors corresponding to the basic physical model; determining the degree of association between each structured semantic description and each structured label to obtain the current association matrix includes: determining the Euclidean distance between the actual topological feature vector of each structured semantic description and the expected topological feature vector of each structured label; determining the degree of association between each structured semantic description and each structured label based on the Euclidean distance to obtain the current association matrix, wherein the Euclidean distance is negatively correlated with the degree of association.
[0014] Preferably, the historical correlation matrix is the correlation matrix corresponding to the real-time operating data stream at the first moment, and the current correlation matrix is the correlation matrix corresponding to the real-time operating data stream at the second moment. Before determining the correction coefficient of each basic physical model based on the correlation similarity, the evolution determination module is further used to: read the maintenance event records of the wind turbine, determine the maintenance events that occurred in the wind turbine from the first moment to the second moment; determine whether the components corresponding to the basic physical model experienced maintenance events during the period from the first moment to the second moment; if the components corresponding to the basic physical model experienced maintenance events during the period from the first moment to the second moment, then record the occurrence time of the maintenance event as the first timestamp; if the components corresponding to the basic physical model did not experience maintenance events during the period from the first moment to the second moment, and the first timestamp is not recorded as the first timestamp. If a maintenance event occurred in a historical moment prior to the current moment, the time of the latest maintenance event in that historical moment is recorded as the first timestamp. If no maintenance event occurred in the components corresponding to the basic physical model during the period from the first moment to the second moment, and no maintenance event occurred in the historical moments prior to the first moment, the commissioning time of the wind turbine is recorded as the first timestamp. The aging factor corresponding to the basic physical model is determined based on the length of the period from the first timestamp to the second moment. The correction coefficient of each basic physical model is determined based on the correlation similarity, including: substituting the correlation similarity into the exponential part of a preset natural exponential function to obtain the distribution stability index, which is positively correlated with the correlation similarity; the correction coefficient of each basic physical model is determined based on the distribution stability index and the aging factor.
[0015] Preferably, the correction coefficients for each basic physical model are determined based on the distribution stability index and the aging factor, including: using the distribution stability index and the aging factor as observational evidence, updating the posterior probability distribution of the preset correction coefficients through Bayes' theorem, and obtaining the updated correction coefficients.
[0016] Preferably, the fault warning module is further configured to: read the target warning data of each basic physical model; determine the comprehensive risk score of the component based on the target warning data of multiple basic physical models related to the same component; determine the comprehensive risk score of the subsystem based on the comprehensive risk scores of all components related to the same subsystem; and output a multi-level risk report based on the comprehensive risk score of the subsystem, the comprehensive risk score of the component, and the target warning data.
[0017] The beneficial effects of this invention are as follows: By determining the degree of association between each structured semantic description and each structured tag, a current association matrix is obtained, which reflects the real-time status of the wind turbine. By determining the association similarity between the data distribution of the current association matrix and the historical association matrix, a correction coefficient for each basic physical model is determined based on the association similarity, so as to reflect the state changes of the wind turbine as the deployment time progresses, and to quantify the specific changes through the correction coefficient. By determining the original early warning data of each basic physical model based on the real-time operation data stream, and correcting each original early warning data according to the correction coefficient of each basic physical model, target early warning data is obtained, so that the target early warning data can fit the current state of the wind turbine for early warning, thereby improving the reliability of early warning for the entire life cycle of the wind turbine. Attached Figure Description
[0018] Figure 1 A schematic diagram of the basic structure of a wind turbine intelligent early warning platform based on big data analysis provided in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating topology data analysis of a wind turbine in a healthy state, according to an embodiment of the present invention. Figure 3 A schematic diagram illustrating topology data analysis of a wind turbine under early fault conditions, provided as an embodiment of the present invention; Figure 4 This is a schematic diagram of topology data analysis of a wind turbine under severe fault conditions, provided as an embodiment of the present invention. Detailed Implementation
[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0020] Example 1, referring to Figure 1 As an embodiment of the present invention, a wind turbine intelligent early warning platform based on big data analysis is provided, comprising: a physical theory module for determining structured labels for each basic physical model of the wind turbine; an operation status module for acquiring real-time operation data streams of the wind turbine and generating structured semantic descriptions based on the real-time operation data streams; an evolution determination module for: determining the degree of correlation between each structured semantic description and each structured label to obtain a current correlation matrix; determining the correlation similarity between the data distribution of the current correlation matrix and the historical correlation matrix; determining the correction coefficient of each basic physical model based on the correlation similarity; and a fault early warning module for: determining the original early warning data of each basic physical model based on the real-time operation data stream; and correcting each original early warning data according to the correction coefficient of each basic physical model to obtain target early warning data.
[0021] Traditional solutions suffer from a mismatch between static physical models and dynamic evolving systems, leading to a decline in the accuracy of early warning platforms throughout the entire lifecycle of wind turbines. Specifically, once established, the fundamental physical models (such as gearbox wear models and bearing fatigue models) have fixed inherent mathematical relationships and parameters. However, as a physical entity, the dynamic characteristics of a wind turbine continuously evolve due to factors such as aging, wear, and maintenance interventions. A model accurate in the initial stages of operation may become inaccurate after several years.
[0022] Based on this, the real-time data can be transformed into structured semantic descriptions through the running status module. The evolution determination module calculates the degree of association between each structured semantic description and each structured label (representing a physical model), forming a current association matrix. Each column of this matrix can be understood as the strength of the explanatory power of each physical model for the system state at the current moment. Furthermore, the evolution of the association relationship can be quantified through the association similarity measure. The current association matrix is compared with the historical association matrix (representing the association relationship at the previous stable moment), and the association similarity is calculated. If the association similarity is high, the relationship between the current model and the data of the wind turbine is stable and predictable compared with the past, that is, the wind turbine is evolving according to a known pattern. If the association similarity is low, the relationship between the current model and the data of the wind turbine has undergone drastic and unpredictable changes compared with the past, that is, the wind turbine has exhibited abnormal behavior that deviates from the historical trajectory.
[0023] A correction coefficient is generated based on the correlation similarity to transform the correlation similarity into a quantitative parameter that can be used to adjust the original warning value. When the correlation similarity is high, the correction coefficient may be close to 1 or grow slowly by an aging-related factor, indicating that the original output of the model is trusted. When the correlation similarity is low, the correction coefficient will be significantly adjusted (amplified or reduced), indicating that the original output of the model may no longer be reliable and needs to be corrected.
[0024] The fault warning module multiplies the original warning data of each model by its corresponding correction coefficient to obtain the target warning data. The final warning value not only includes the model's judgment on the current state, but also integrates the platform's judgment on whether the current state evolution mode of the wind turbine is normal.
[0025] For example, the basic physical model includes a gearbox wear model (label M1) and a generator misalignment model (label M2). At historical time T0, the wind turbine has just completed a routine maintenance and is operating smoothly. The historical correlation matrix shows that the correlation degree of M1 is 0.8, and the correlation degree of M2 is 0.1, indicating that the wind turbine's condition is mainly explained by normal gearbox wear. At this time, the initial warning data output by the gearbox wear model is 0.2 (low risk).
[0026] Assuming the scenario evolves into normal aging, at time T1: one year later, with no maintenance, the gearbox experiences normal wear and vibration increases slowly. The current correlation matrix shows a correlation degree of 0.85 for M1 and 0.05 for M2. The wind turbine's condition is still primarily explained by gearbox wear, and the correlation remains stable. The evolution determination module calculates the correlation similarity between the current and historical correlation matrices, resulting in a high similarity of 0.95. Based on this similarity, the system calculates a correction coefficient for the M1 model, for example, 1.2 (to appropriately amplify the risk considering one year of normal aging). At this point, the original warning data output by the gearbox wear model has risen to 0.4, and the fault warning module outputs a target warning data of 0.4 * 1.2 = 0.48 for low-risk warning, indicating that the risk is increasing as expected.
[0027] Assuming the scenario evolves into a sudden anomaly, at time T2: after T1, a bearing in the gearbox suddenly develops pitting corrosion, leading to a sharp increase in high-frequency vibration. The current correlation matrix shows that the correlation degree of M1 drops sharply to 0.3, while the correlation degree of a bearing pitting model not in the library (or a general high-frequency vibration model) soars to 0.7, indicating a change in the dominant dynamic mode of the wind turbine gearbox. The evolution determination module calculates the correlation similarity between the current correlation matrix and the historical correlation matrix (T0 or T1), with a result of 0.25 (low). This indicates an abnormal system evolution trajectory. Based on the correlation similarity, the platform calculates a correction coefficient for the M1 model, for example, 0.1, because the M1 model can no longer interpret the current data, and its original output has extremely low reliability. At this point, the original warning data output by the gearbox wear model may be falsely reported as 0.9 (high risk) due to high-frequency vibration interference. The fault warning module outputs target warning data = 0.9 * 0.1 = 0.09. The platform successfully suppressed the false alarms of the M1 model. At the same time, based on the current correlation matrix, the platform generates a new, high-risk warning for the bearing pitting model with soaring correlation.
[0028] By monitoring the evolution of the relationship between the model and data, intelligent correction of early warning outputs is achieved, enabling it to distinguish between normal evolution and abnormal mutations, thereby maintaining high accuracy and low false alarm rate throughout the entire life cycle.
[0029] Furthermore, the continuous variation in wind turbine operating parameters (such as wind speed and rotational speed) leads to strong non-stationarity in the data. Directly analyzing long-term data without differentiation will mix the dynamic characteristics under different operating conditions, resulting in distorted analysis results that fail to reflect the true state under any single operating condition. Moreover, the harsh operating environment of wind turbines means that sensor data contains significant background noise. Early, weak fault signals often have energy far lower than the noise energy, making them easily submerged and rendering traditional signal-to-noise ratio analysis methods ineffective. Therefore, traditional Fourier transform-based spectral analysis or statistical moment analysis, while sensitive to signal amplitude and frequency, cannot effectively capture early, progressive, intermittent faults in wind turbine components.
[0030] The core idea of the assumptions based on linear systems and stationary stochastic processes is to decompose complex signals into a superposition of a series of simple basis functions (such as sine waves), and infer the system state by analyzing the coefficients (spectrum, energy) of the basis functions. This method is effective for analyzing stable, periodic phenomena. However, wind turbine data exhibits typical nonlinear and non-stationary dynamic characteristics. Therefore, it is necessary to effectively extract deep dynamic features that can characterize early signs of intermittent faults from the noisy, non-stationary real-time operation data stream of wind turbines. Preferably, generating a structured semantic description based on the real-time running data stream includes: marking data streams in the real-time running data stream whose variance is less than a preset variance threshold within a preset time window as quasi-steady-state operating condition data segments; determining the autocorrelation function of the fault data sequence in the quasi-steady-state operating condition data segment; recording the time length required for the autocorrelation function to first decay to the preset variance threshold as the dynamic window length under the quasi-steady-state operating condition data segment; truncating the quasi-steady-state operating condition data segment based on the dynamic window length to obtain the fault data sequence; reconstructing the phase space of the fault data sequence to obtain a high-dimensional phase space point cloud; generating a persistent graph based on the continuous cohomology of the high-dimensional phase space point cloud; extracting topological invariants from the persistent graph to obtain the actual topological feature vector; and determining the structured semantic description based on the actual topological feature vector.
[0031] The nonlinear and non-stationary dynamic characteristics of wind turbines can be transformed into a description through attractors in phase space. A healthy wind turbine system has a stable attractor, while the occurrence and development of faults can be manifested as changes in the shape, size, or topology of this attractor. Therefore, the analysis can be transformed from the traditional frequency domain or statistical domain of signals to the topological analysis of the system state, so as to get closer to the physical nature of the fault.
[0032] Data streams in the real-time operating data stream whose variance is less than a preset variance threshold within a preset time window are marked as quasi-steady-state operating condition data segments. By monitoring the stability of key operating parameters, relatively pure data segments with consistent dynamic characteristics are cut out from non-stationary data streams, ensuring that subsequent analyses are based on a single operating condition and avoiding mixed contamination from different operating condition characteristics. Furthermore, relatively stable quasi-steady-state segments are first identified through variance thresholds to ensure data consistency. Instead of using a fixed window, the dynamic window length that can capture the intrinsic dynamics of the wind turbine is dynamically calculated based on the decay characteristics of the data's autocorrelation function, adapting to the current operating state.
[0033] Phase space reconstruction is performed on the fault data sequence to obtain a high-dimensional phase space point cloud and a persistent graph generated from the persistent cohomology of the high-dimensional phase space point cloud. Phase space reconstruction restores the one-dimensional time series to the evolution trajectory (attractor) of the system in the high-dimensional state space. Persistent cohomology is a powerful mathematical tool that does not focus on the precise coordinates of the trajectory, but rather analyzes its topological invariants (such as the number and persistence of loops and holes). This makes it extremely sensitive to small changes in the shape of the attractor, which are the deep dynamic characteristics of early faults.
[0034] Topological invariants are extracted from the persistent graph to obtain the actual topological feature vector. Topological features (such as a ring) are robust to random noise. Random noise may add some outliers to the point cloud, but it is difficult to create or completely eliminate a persistent ring structure dominated by system dynamics. Therefore, by analyzing the persistence of the topological structure, we can effectively distinguish between short-term disturbances caused by noise and structural changes caused by faults. The resulting topological feature vector is a quantitative description of this robust, structural change.
[0035] For example, early pitting corrosion on gear teeth is analyzed using vibration acceleration sensor data from a wind turbine gearbox. In a healthy state, the platform operates under stable conditions with a wind speed of 12 m / s. A segment of data is extracted, and the reconstructed phase space point cloud forms a clear, approximately ring-shaped attractor, corresponding to the periodic meshing of the gears. In the H1 dimension (representing a one-dimensional ring structure), the persistence map shows a highly persistent topological feature point, representing the main loop. In the quantized vectors, most of the weights are concentrated within the grid representing the high-persistence loop, such as... Figure 2As shown; if an early pitting failure occurs and the platform operates at a wind speed of 12 m / s, another segment of data is extracted. Because the pitting on the tooth surface generates a weak impact with each engagement, the original smooth annular attractor is slightly disturbed or stretched, and may even generate a tiny secondary ring associated with the main ring. This causes changes in the persistence map of dimension H1. The high-persistence main ring feature points in the original healthy state may still exist, but their birth / death coordinates have shifted. More importantly, a new, lower-persistence topological feature point is very likely to appear, representing the secondary ring caused by the pitting impact, such as... Figure 3 As shown. Furthermore, the quantized vectors differ significantly from the vectors in the healthy state; for example, the grid counts representing high-durability cycles decrease, while the grid counts representing medium / low-durability cycles increase.
[0036] Although the impact caused by pitting may be completely submerged in noise in the original vibration signal, resulting in no significant change in the FFT spectrum, this application successfully extracted early signs of fault by capturing changes in the system dynamic topology of the wind turbine. The platform can identify this difference by comparing the real-time topology feature vector with the baseline vector in the healthy state and output early warning.
[0037] Furthermore, standard phase space reconstruction parameter selection methods (such as the autocorrelation function method) only capture linear correlations and are not applicable to nonlinear systems. While the mutual information method can capture nonlinear dependencies, its function curves generate numerous spurious minima in noisy environments, leading to time delays in incorrect selection and posing risks of overfitting and underfitting in embedding dimension selection. For high-noise data, the correlation dimension may never saturate, resulting in the selection of an excessively large embedding dimension (overfitting, amplifying noise). For data with varying quality, a fixed threshold cannot adapt adaptively, potentially leading to the selection of an excessively small embedding dimension (underfitting, insufficient development of dynamic information). Therefore, standard phase space reconstruction parameter selection methods fail when applied to noisy, non-stationary industrial data (such as wind turbine vibration data), resulting in a reconstructed phase space that cannot accurately represent the true dynamic structure of the system. Therefore, preferably, the phase space reconstruction of the fault data sequence to obtain a high-dimensional phase space point cloud includes: determining the mutual information function and permutation entropy function of the fault data sequence under multiple preset delay times; determining the first local minimum point in the permutation entropy function, wherein the first local minimum point is the smallest independent variable value among the multiple local minimum points of the permutation entropy function; determining the second local minimum point in the mutual information function within the domain from the preset initial value to the first local minimum point, wherein the second local minimum point is the smallest independent variable value among the multiple local minimum points of the mutual information function; and then reconstructing the second local minimum point. The minimum point is determined as the optimal time delay parameter; the signal-to-noise ratio (SNR) of the fault data sequence is determined; an adaptive threshold for the fault data sequence is determined based on the SNR, wherein the adaptive threshold is negatively correlated with the SNR; starting from the preset minimum embedding dimension, the correlation dimension under the current embedding dimension and the relative rate of change of the correlation dimension compared to the previous correlation dimension are determined; the current embedding dimension when the relative rate of change is less than the adaptive threshold is determined as the optimal embedding dimension; the phase space of the fault data sequence is reconstructed based on the optimal time delay parameter and the optimal embedding dimension to obtain a high-dimensional phase space point cloud.
[0038] Permutation entropy is less sensitive to random noise than mutual information, and can more stably identify the interval with the lowest system complexity. By finding the first local minimum of the permutation entropy function, a high-quality search range [0, t_pe] is determined, which can effectively exclude subsequent regions heavily contaminated by noise. Within the robust search range, the second local minimum of the mutual information function is then searched. Since the search range has been purified, this minimum point has a higher probability of being real and determined by system dynamics, rather than a noise artifact. This effectively combines the robustness of permutation entropy and the ability of mutual information to capture nonlinear relationships, avoiding the risk of selecting the wrong time delay parameter τ due to noise interference in the global scope.
[0039] To address the risks of overfitting and underfitting in embedding dimension selection, an adaptive threshold based on signal-to-noise ratio (SNR) can be introduced. The threshold ϵ is inversely proportional to the SNR (ϵ = ϵ0 / SNR). When the SNR is high (data is clean), the SNR value is large, and the calculated adaptive threshold ϵ is small. The algorithm will require the rate of change of the correlation dimension to be very small before stopping the iteration, thus accurately finding the saturation point and avoiding underfitting. When the SNR is low (data is full of noise), the SNR value is small, and the calculated adaptive threshold ϵ is large. The algorithm will stop the iteration earlier, and even if the correlation dimension has changed to some extent, it will be considered to have saturated, thus preventing the algorithm from continuing to fit noise in high-dimensional space, thereby preventing overfitting. This realizes a direct correlation between the selection of embedding dimension and data quality, thereby achieving the optimal trade-off of embedding dimension under different data quality conditions, ensuring the stability and effectiveness of phase space reconstruction.
[0040] To determine the optimal time delay parameter (τ), for the input fault data sequence, the mutual information function MI(t) and the permutation entropy function PE(t) are calculated simultaneously under a series of preset delay times t. On the permutation entropy function PE(t), the first local minimum point is found, and its corresponding delay time is t_pe. In the domain [1, t_pe] of the mutual information function MI(t), the first local minimum point is found, and its corresponding delay time is t_mi. t_mi is then assigned to the optimal time delay parameter τ.
[0041] Permutation entropy measures the degree of order in a time series. For noise, its permutation entropy is high and stable. For deterministic signals, its permutation entropy changes with the delay time. Therefore, the first local minimum point t_pe of PE(t) usually corresponds to a delay range that can maximally represent the determinism of the system. This range is not sensitive to noise. Mutual information measures the generalized statistical correlation between two sequences. Within the robust range [0, t_pe] determined by permutation entropy, the function shape of MI(t) is less affected by noise. Its first local minimum point t_mi can more accurately reflect that the redundancy of the system state itself is minimized, that is, the reconstructed coordinates contain the maximum amount of independent information. This method uses permutation entropy to provide a clean search interval for mutual information, avoiding the pseudo-minimal problem caused by noise in the global search, thereby achieving a robust and accurate selection of τ.
[0042] To determine the optimal embedding dimension, the signal-to-noise ratio (SNR) of the faulty data sequence is calculated. An adaptive threshold ϵ is calculated using the formula ϵ = ϵ0 / SNR, where ϵ0 is a preset constant. Iterative calculations (a~g) are then performed: a. Initialize the current embedding dimension m_c to the preset minimum embedding dimension m_min; b. Calculate the correlation dimension under the current embedding dimension m_c. c. Calculate the correlation dimension in m_c-1 dimensions. d. Calculate the relative rate of change e. Determine if ΔD is less than the adaptive threshold ϵ; f. If If the iteration stops, the current m_c is determined as the optimal embedding dimension m; g. If If m_c = m_c + 1, then return to step b until m_c reaches the preset maximum embedding dimension m_max.
[0043] The core of the optimal embedding dimension calculation method lies in the fact that the threshold ϵ is a function of SNR, which allows the algorithm's stopping condition to be dynamically adjusted according to data quality. When SNR is low, the value of ϵ is large, meaning that even if the correlation dimension D2 still increases to a certain extent (this increase is likely caused by noise being amplified in high-dimensional space), the algorithm will determine that it has saturated and stop early, thus effectively preventing the selection of an excessively large m to fit the noise; when SNR is high, the value of ϵ is small, and the algorithm requires that the rate of change of D2 must be very small before stopping, thus accurately finding the true embedding dimension where D2 tends to stabilize (saturate), avoiding underfitting caused by premature stopping.
[0044] Furthermore, the standard Vietoris-Rips (VR) complex construction algorithm exhibits computational complexity that grows exponentially with the point cloud size N (worst-case complexity O(2^N)). For high-dimensional phase space point clouds containing tens or even hundreds of thousands of points reconstructed from wind turbine data, the computation time is unacceptable on existing engineering hardware. Moreover, constructing the VR complex requires storing all simplexes that grow with the scale parameter, resulting in exponential memory consumption. This leads to rapid exhaustion of computational resources when processing large-scale point clouds. In wind farm predictive maintenance scenarios, the platform needs to perform near real-time analysis of continuously flowing data on edge computing devices. An algorithm that takes hours or even days to complete a single calculation is meaningless for early warning. Therefore, an approximate calculation strategy based on landmark and witness point complexes can address the computational complexity bottleneck. Preferably, generating a persistent graph based on the continuous cohomology of a high-dimensional phase space point cloud includes: selecting a preset number of landmark point sets from the high-dimensional phase space point cloud using a preset maximum-minimum distance method; determining potential simplexes in the landmark point sets; for each potential simplex: determining witness points corresponding to the potential simplex from the high-dimensional phase space point cloud; determining the set formed by removing potential simplexes from the landmark point set as the landmark point difference set corresponding to the potential simplex; if the distance from the witness point to the vertex of the potential simplex corresponding to the witness point is less than the distance from the witness point to any point in the landmark point difference set corresponding to the potential simplex, then adding the potential simplex to the Witness complex; constructing the Witness complex based on each potential simplex, and determining the continuous cohomology of the Witness complex to obtain a persistent graph.
[0045] By selecting a much smaller set of landmarks from the original large-scale point cloud, the subsequent topology calculation object is reduced from N points to k points (k << N), thus achieving dimensionality reduction. Instead of directly constructing a composite on the k landmarks, a Witness composite is constructed. The construction of this composite relies on points in the original point cloud as witnesses to verify the validity of the topological structure between landmarks, thereby completing an approximate construction. The computational complexity is transformed from exponential, depending on the size N of the original point cloud, to polynomial, depending on the number of landmarks k. Although the witnessing process still requires traversing the original point cloud, its computational complexity is controllable (e.g., O(N*k^2)), far lower than exponential, thus achieving complexity transfer.
[0046] For landmark selection, initialize the landmark set Landmarks to empty, randomly select a point p1 from the original high-dimensional phase space point cloud P, and add it to Landmarks. For i ranging from 2 to a preset number N_landmarks, execute a~d: a. Traverse all points in P that do not belong to Landmarks; b. For each point p, calculate the minimum distance d(p, Landmarks) from it to all points in Landmarks; c. Find the point p_i that maximizes d(p, Landmarks); d. Add p_i to Landmarks.
[0047] Among them, the maximum-minimum distance algorithm ensures that the newly added landmarks are farthest from the existing set of landmarks, so that the final set of selected landmarks can cover the geometric distribution of the original point cloud to the greatest extent, avoiding the clustering of landmarks in local areas, thus ensuring the representativeness of subsequent approximate calculations. This step reduces the problem size from N to k and is the basis of the entire computational optimization.
[0048] For Witness complex construction, all potential simplexes (such as all vertices, edges, triangles, etc.) of the landmark set Landmarks can be generated. For each potential simplex σ (e.g., a triangle formed by landmarks l_i, l_j, l_k), perform a~b: a. Determine the witness point: In the original point cloud P, find a point p that satisfies the following condition: the distance from point p to all vertices of simplex σ is less than the distance from point p to any other point in Landmarks, i.e., dist(p, l_i), dist(p, l_j), dist(p, l_k) are all less than dist(p, l_q) for all l_q ∈ Landmarks and l_q ∉ {l_i, l_j, l_k}; b. Construct the complex: If there exists at least one such witness point p, then add simplex σ to the Witness complex.
[0049] Instead of constructing connections between landmarks, the original set of all data points is used as ground reality to verify the validity of these connections. If a simplex formed by a landmark cannot find any supporting points in the original data, it is considered unimportant and will not be added to the complex. This filters out a large number of unimportant simplexes, making the final Witness complex much sparser than the VR complex constructed on the same number of landmarks, thus greatly reducing the computational cost of subsequent continuous cohomology calculations.
[0050] For the computation of persistent homology, standard persistent homology algorithms (such as through matrix reduction) are applied to the constructed, sparse Witness complex to calculate the birth and death times of homology groups in each dimension, and output a persistent graph.
[0051] Since the size of the Witness complex is determined by the number of landmarks k and the sparsification effect of the witnessing process, its size is controllable and much smaller than that of the original problem. Therefore, continuous homology calculations on this basis have a polynomial-level computational complexity and can be completed within an engineering-acceptable timeframe.
[0052] For example, the original high-dimensional phase space point cloud P contains 10,000 points, which roughly form a high-dimensional ring structure, with a preset number of landmarks N_landmarks = 100. The process is as follows: For landmark selection, the maximum-minimum distance method is used to select 100 representative landmarks L1, L2, ..., L100 that are evenly distributed in space from 10,000 points. For Witness complex construction, a potential 2-simplex (triangle) is considered, with its vertices being landmarks L5, L23, and L47. The platform searches for witness points among the original 10,000 points. Assuming a point p_x is found, the distances dist(p_x, L5), dist(p_x, L23), and dist(p_x, L47) are calculated. The distances from p_x to the other 97 landmarks are calculated, and the minimum value min_dist_to_others is found. If dist(p_x, L5), dist(p_x, L23), and dist(p_x, L47) are all less than min_dist_to_others, then the triangle (L5, L23, L47) is considered a valid landmark. L47 is then added to the Witness complex, and this process is repeated for all possible simplexes. For the calculation of continuous homology, assuming that the final constructed Witness complex contains 500 vertices, 1200 edges and 300 triangles, the platform calculates continuous homology on a complex containing 2000 simplexes, which may take from a few seconds to a few minutes.
[0053] The final persistent graph output will show a high-persistence feature point in the H1 dimension, successfully representing the ring structure in the original data.
[0054] Through the above processing, a problem that originally required processing 10,000 points and was computationally infeasible was transformed into an approximate problem that required processing 100 landmark points and constructing a sparse complex under the witness of 10,000 points, and was computationally feasible. The resulting topological features (loops) were successfully preserved, while the computational efficiency was improved exponentially, meeting the requirements of real-time engineering.
[0055] Furthermore, a persistent graph is essentially a two-dimensional set of points, the number and location of which vary with the input data. Persistent graphs contain rich topological information, but this information is geometric and distributed. Directly using point coordinates cannot effectively quantify the differences between different persistent graphs. Therefore, it is necessary to encode the distribution patterns of points in the persistent graph. Preferably, extracting topological invariants from the persistent graph to obtain the actual topological feature vector includes: setting grid partitioning rules for the persistence dimension and birth time dimension of the persistent graph to obtain multiple grids, wherein the persistent graph is a two-dimensional point set, each point in the two-dimensional point set represents a topological feature, the horizontal axis of the coordinate system corresponding to the two-dimensional point set is the birth time, the vertical axis is the persistence, and the persistence is the difference between the death time and the birth time of the topological feature; selecting topological features corresponding to the first homology group from the persistent graph, wherein the topological features of the first homology group are used to characterize the one-dimensional ring structure in the data; for each topological feature on the first homology group in the persistent graph, determining the grid coordinates to which the topological feature belongs based on the birth value and death value of the topological feature; counting the number of topological features falling into each grid in the persistent graph; and arranging the statistics in the grid to obtain the actual topological feature vector.
[0056] A persistent graph is a geometric object whose value lies in the overall distribution, density, and outliers of its points. For example, a persistent graph consisting of many high-persistence points represents very different system dynamics than a persistent graph consisting of a large number of low-persistence points.
[0057] The geometric distribution information of a persistent graph can be encoded into a fixed-length numerical vector, which is the signature of the persistent graph. This allows the similarity between two persistent graphs to be quantified by calculating the Euclidean distance between their signature vectors, thus enabling the encoding of the distribution pattern of points in the persistent graph.
[0058] By defining a grid on the two-dimensional coordinate system of the persistent graph, the continuous and infinitely possible point distribution space is discretized into a finite number of ordered grid cells, providing a structured framework for quantization. By counting the number of topological features falling within each grid cell, the complex point distribution pattern is transformed into a simple histogram, which describes the density distribution of topological features in the two-dimensional space of birth time and persistence. The two-dimensional statistical histogram is flattened into a one-dimensional vector in a fixed order (e.g., row priority). Since the number of grid cells is preset and fixed, the dimension of the final generated vector is also fixed, solving the problem of variable-length input.
[0059] Specifically, a two-dimensional coordinate system is defined for the persistence graph, with the horizontal axis representing birth_time and the vertical axis representing persistence, where persistence = death_time - birth_time. The number of grid cells is set to N_bin, and the ranges for both the horizontal and vertical axes are determined as [0, Birth_max] and [0, Persistence_max]. Each axis is then uniformly divided into N_bin intervals, forming an N_bin × N_bin grid. This step maps the abstract two-dimensional space to a concrete, quantifiable grid structure, transforming a continuous geometric distribution into the basis of discrete numerical statistics, thus resolving the data structure mismatch problem.
[0060] Then, H1 topological features are filtered out. From the topological features (H0, H1, H2...) of all dimensions in the persistent graph, only features belonging to the first homology group (H1) are extracted.
[0061] The H1 feature mathematically corresponds to a one-dimensional ring structure. In wind turbine vibration analysis, this is usually associated with periodic motion or impact response caused by specific faults (such as bearing damage or gear tooth breakage). Filtering the H1 feature can focus on the topological information most relevant to the target fault, filtering out irrelevant information from H0 (connected components), and improving the signal-to-noise ratio of subsequent feature vectors.
[0062] For each selected H1 topological feature point (b, p) (where b is the birth time and p is the persistence), calculate its corresponding grid coordinates (i, j) using the formula: i = floor(b / (Birth_max / N_bin)), j = floor(p / (Persistence_max / N_bin)). This step precisely maps each topological feature point to a specific cell in the grid structure, converting geometric location information into discrete index information.
[0063] To count the number of features within a grid, an N_bin × N_bin zero matrix C is initialized. All H1 topological feature points are iterated through, and for each point, the counter C_ij in its corresponding grid matrix C is incremented. This process generates a two-dimensional histogram that accurately quantifies the distribution density of H1 topological features in the birth time and persistence spaces. For example, high counts in high-persistence regions indicate a robust periodic structure in the system; high counts in low-persistence regions may indicate noise or transient perturbations. This solves the information quantification problem.
[0064] For the actual topological feature vector obtained by permutation, the N_bin × N_bin counting matrix C can be flattened into a one-dimensional vector containing N_bin × N_bin elements in row-major or column-major order to generate the final fixed-length numerical vector. This vector is a compact and structured numerical representation of the original persistent graph and can be directly used as input to any standard machine learning model to build a data link between topology analysis and fault diagnosis applications.
[0065] For example, for the H1 persistent plot of wind turbine gearbox vibration data, N_bin = 3, Birth_max = 30, Persistence_max = 120 can be set. Therefore, each grid cell has a width of 10 on the horizontal axis and a height of 40 on the vertical axis.
[0066] In a healthy state, the H1 persistence graph has only one topological feature point located at (birth=15, persistence=100), which represents a very stable and robust periodic oscillation loop. The processing flow includes: calculating the grid coordinates, i = floor(15 / 10) = 1, j = floor(100 / 40) = 2, the topological feature point falls on grid (1, 2); in the counting matrix C, C_12 = 1, and all other positions are 0; then vectorization (row priority): [0, 0, 0, 0, 0, 1, 0, 0, 0].
[0067] In early tooth surface wear failures, the original point (15, 100) in the H1 persistence graph still exists, but due to the small impact caused by wear, two new points with lower persistence are added, located at (birth=8, persistence=25) and (birth=22, persistence=30). The processing flow includes: calculating the mesh coordinates, (15, 100) -> (1, 2), (8,25) -> (0, 0), (22, 30) -> (2, 0); in the counting matrix C, C_12 = 1, C_00 = 1, C_20 = 1, and the rest are 0; then vectorization (row priority) is performed: [1, 0, 0, 0, 0, 1, 0, 0, 1].
[0068] The feature vector for a healthy state is [0, 0, 0, 0, 0, 1, 0, 0, 0], while the feature vector for an early failure state is [1, 0, 0, 0, 0, 1, 0, 0, 1]. These two vectors differ significantly in value, transforming the previously incalculable topological distribution differences into computable vector differences, thereby enabling automatic detection of early tooth surface wear faults.
[0069] For specific differences, please refer to Figures 2-4 Vibration signal analysis refers to acquiring vibration data from wind turbine sensors and identifying quasi-steady-state data segments, as shown in the legend in the upper left corner. Phase space reconstruction refers to reconstructing a one-dimensional time series into a high-dimensional phase space point cloud using optimal time delay parameters and embedding dimensions, as shown in the legend in the upper right corner. Persistent graph generation refers to generating an H1-dimensional (one-dimensional ring structure) persistent graph based on an approximate calculation strategy using landmark and witness point composites, as shown in the legend in the lower left corner. Feature vector extraction refers to dividing the persistent graph into a grid, counting the number of topological features within each grid, and forming a fixed-length topological feature vector, as shown in the legend in the lower right corner.
[0070] Preferably, the real-time operating data stream includes at least one of a vibration data stream, a temperature data stream, a pressure data stream, and a power data stream. Before marking the data streams in the real-time operating data stream whose variance is less than a preset variance threshold within a preset time window as quasi-steady-state operating condition data segments, the operating status module is further configured to: perform spectral analysis on the vibration data stream in the real-time operating data stream, and determine the first type of fault data sequence of the vibration data stream in the real-time operating data stream based on the energy value of a preset frequency band; perform rate of change analysis on the temperature data stream and pressure data stream in the real-time operating data stream, and determine the second type of fault data sequence of the temperature and pressure data streams in the real-time operating data stream based on the abnormal rate of change; perform efficiency curve analysis on the power data stream in the real-time operating data stream, and determine the third type of fault data sequence of the power data stream in the real-time operating data stream based on the deviation of the power curve; and determine a structured semantic description based on the actual topological feature vector, including: determining the structured semantic description based on the actual topological feature vectors corresponding to the first type of fault data sequence, the second type of fault data sequence, and the third type of fault data sequence.
[0071] For different types of operational data streams, traditional analysis methods best suited to their physical characteristics are employed for parallel monitoring. For example, vibration data is analyzed using spectral analysis to capture dynamic anomalies in mechanical structures. A sliding window Fourier transform (STFT) or wavelet transform is performed on the vibration data stream to calculate the energy values of one or more preset fault characteristic frequency bands (such as the sidebands of gear meshing frequencies, bearing fault characteristic frequencies, etc.). If the energy value exceeds a preset threshold, the data segment within that window is marked as a first-type fault data sequence. Temperature / pressure data is analyzed using rate of change analysis to capture transient anomalies in thermodynamic or fluid systems. The first or higher-order time derivatives (i.e., rates of change) of the temperature and pressure data streams are calculated. If the absolute value of the rate of change exceeds a preset threshold within a short period, the data segment within that time period is marked as... The data sequence is designated as the second type of fault data. Many system-level faults (such as lubrication failure and cooling system blockage) first manifest as a sharp change in temperature or pressure. By capturing transient anomalies, it supplements the diagnostic dimensions that vibration analysis cannot cover. Power data is analyzed for efficiency curves to capture anomalies in energy conversion efficiency. The deviation between the actual power data points and the predicted values of the benchmark model is calculated in real time. If the deviation continues to exceed a preset threshold, the data segment within that time period is marked as the third type of fault data sequence. This monitors the overall health of the wind turbine from the perspective of energy conversion efficiency. Performance degradation is a common manifestation of multiple potential faults, providing a macroscopic, system-level fault detection perspective.
[0072] For any data segment marked as a fault data sequence by the above method, a complete topology analysis pipeline is applied to generate an actual topology feature vector, which is used to characterize the quantified data of the fault event under a specific data mode.
[0073] The topological feature vectors generated from all data modes are combined to form a unified structured semantic description. This description encodes the comprehensive topological features of the fault event across multiple dimensions, including vibration, thermodynamics, and energy, providing rich and unique input for subsequent accurate diagnosis.
[0074] For example, suppose that within a time window, the system simultaneously identifies three types of fault data sequences: type 1, type 2, and type 3. The topology analysis pipeline is executed on each of these three sequences to obtain topology feature vectors V_vib, V_temp, and V_pow. These three vectors are concatenated or weighted and summed to generate the final structured semantic description, such as S = [V_vib, V_temp, V_pow].
[0075] A specific physical fault (such as severe wear of gearbox bearings) may simultaneously generate specific vibration topological features, temperature change topological features, and power reduction topological features, which together form a unique multimodal topological signature with a much higher distinguishability than any single-mode feature vector, thus enabling precise location of the fault root cause.
[0076] For example, when distinguishing between early pitting on gear teeth and poor lubrication of generator bearings.
[0077] Fault A: Early pitting on gear tooth surfaces. For vibration characteristics, a weak sideband is generated near the meshing frequency, with energy exceeding the threshold, marking it as a Type I fault sequence. For temperature characteristics, the gearbox oil temperature rises extremely slowly, but the rate of change does not exceed the threshold, resulting in no Type II sequence. For power characteristics, transmission losses increase slightly, and the power efficiency deviation does not exceed the threshold, resulting in no Type III sequence. Topology analysis and fusion are performed on the vibration sequences, yielding a topology vector V_vib_A, characterized by a low-persistence secondary perturbation loop appearing on the main loop. The final structured semantic description is S_A = [V_vib_A, 0, 0].
[0078] Fault B: Poor lubrication of generator bearings. For vibration characteristics, the high-frequency vibration energy at the generator end increases, marking it as a first-type fault sequence. For temperature characteristics, the temperature of the generator bearing housing rises rapidly, and the rate of change exceeds the threshold, marking it as a second-type fault sequence. For power characteristics, the generator friction loss increases, the power efficiency decreases, and the deviation exceeds the threshold, marking it as a third-type fault sequence. Topological analysis of the vibration sequence yields a topological vector V_vib_B, which may be characterized by multiple unrelated loops in a high-frequency noise background. Topological analysis of the temperature sequence yields a topological vector V_temp_B, which is characterized by a topological structure corresponding to a monotonically increasing trend. Topological analysis of the power sequence yields a topological vector V_pow_B, which is characterized by a topological pattern of efficiency curve deviation. The final structured semantic description is S_B = [V_vib_B, V_temp_B, V_pow_B].
[0079] Therefore, although both faults trigger vibration alarms, the resulting S_A and S_B, obtained through multimodal fusion, are distinctly different high-dimensional vectors. S_A primarily represents mechanical impact, while S_B is a comprehensive signature containing multidimensional information on mechanics, thermodynamics, and energy efficiency. A downstream classification model can easily identify S_A as gear pitting and S_B as poor bearing lubrication, thus achieving high-precision root cause diagnosis.
[0080] Furthermore, there is a mapping gap between the data-driven feature space (the space corresponding to the topological feature vectors) and the physics-based knowledge space (the space of fault mode distribution). Therefore, it is necessary to map the data-driven feature space to the physics-based knowledge space. Preferably, determining the structured label of each basic physical model of the wind turbine includes: determining the expected fault sequence of the basic physical model based on the fault modes of the basic physical model; determining the expected topological feature vector of the basic physical model based on the expected fault sequence; determining the structured label of the basic physical model based on the components, subsystems, fault modes, and expected topological feature vectors corresponding to the basic physical model; determining the correlation degree between each structured semantic description and each structured label to obtain the current correlation matrix, including: determining the Euclidean distance between the actual topological feature vector of each structured semantic description and the expected topological feature vector of each structured label; determining the correlation degree between each structured semantic description and each structured label based on the Euclidean distance to obtain the current correlation matrix, wherein the Euclidean distance is negatively correlated with the correlation degree.
[0081] The data-driven feature space is a high-dimensional Euclidean space consisting of all possible topological feature vectors. In this space, the distance between two vectors reflects the similarity of the dynamic modes they represent, but the space itself does not contain any physical concepts about gears, bearings, or misalignment.
[0082] Physics-based knowledge space is a conceptual space in the engineering field, in which the elements are specific physical components (such as gearboxes and generators), subsystems (such as transmission chains and hydraulic systems), and failure modes (such as wear, fracture, and overheating). It contains discrete, well-defined physical entities and phenomena.
[0083] Therefore, it is necessary to establish a computable and traceable mapping relationship between these two spaces, so that the platform can automatically convert an abstract topological vector into a specific physical diagnostic conclusion.
[0084] Determining the expected fault sequence can be achieved using high-fidelity multiphysics simulation software (such as ANSYS, Romax) or by analyzing rigorously validated historical full-cycle fault data. For specific underlying physical models (such as gear wear models), theoretical time-series data on the occurrence and development of faults can be generated. Determining the expected topological feature vectors involves applying a complete topology analysis pipeline (quasi-steady-state identification, phase space reconstruction, continuous coherence, and feature vectorization) to the expected fault sequence generated in the previous step. Determining structured labels involves creating a data structure (such as a class or a tuple) to bind physical metadata (components, subsystems, fault modes) to mathematically quantified expected topological feature vectors. Determining the Euclidean distance involves calculating the Euclidean distance in high-dimensional space between a real-time generated actual topological feature vector V_real and an expected topological feature vector V_theory from a theoretical library: Distance = sqrt(Σ(V_real[i] - V_theory[i])²), a standard metric for measuring the similarity of high-dimensional vectors; a smaller distance indicates a closer similarity in the topological patterns of the two vectors. It provides a quantifiable and objective similarity criterion. To determine the degree of association, the Euclidean distance is converted into an association score between 0 and 1 using a monotonically decreasing function (e.g., Similarity = 1 / (1 + Distance)).
[0085] For example, a theory library has been built with tag A (gear pitting): {component: gear, subsystem: gearbox, failure mode: pitting, expected topology feature vector: V_theory_pitting}, and tag B (bearing outer ring failure): {component: bearing, subsystem: gearbox, failure mode: outer ring failure, expected topology feature vector: V_theory_bearing}.
[0086] The real-time diagnostic process includes: real-time data input, where the platform acquires abnormal data from the vibration sensor of the wind turbine gearbox, processes it through the TDA pipeline, and generates a real topological feature vector V_real_unknown; matching calculation, calculating the distance to label A, Distance_A = EuclideanDistance(V_real_unknown, V_theory_pitting), assuming the result is Distance_A = 0.85; calculating the distance to label B, Distance_B = EuclideanDistance(V_real_unknown, V_theory_bearing), assuming the result is Distance_B = 0.15; and correlation calculation, calculating the correlation with label A, Similarity_A = 1 / (1 + 0.85) ≈ 0.54; and the correlation with label B, Similarity_B = 1 / (1 + 0.15) ≈ 0.87; Generate the current association matrix. The platform generates a 1x2 association matrix (or a dictionary): Current_Association_Matrix = { "Gear pitting": 0.54, "Bearing outer ring failure": 0.87}.
[0087] Through quantitative comparison, the theoretical fingerprint correlation between the currently observed abnormal state and bearing outer ring failure (0.87) was determined to be significantly higher than that with gear pitting (0.54). Therefore, the most likely diagnostic conclusion output by the system is that the bearing in the gearbox system has experienced an outer ring failure. This process is entirely based on data-driven quantitative calculations, achieving an automatic and accurate mapping from abstract topological features to specific physical faults.
[0088] Furthermore, static early warning models cannot distinguish between expected state evolution caused by deterministic events such as normal equipment aging and planned maintenance, and abnormal state evolution caused by sudden failures. This leads to a decrease in the accuracy and robustness of the early warning system throughout the equipment's entire lifecycle. A single maintenance event (such as replacing a bearing) resets the health age of its corresponding component. A model that does not consider maintenance history will incorrectly assume that newly replaced bearings and old bearings have the same failure risk, resulting in a large number of false alarms after maintenance. Moreover, the failure risk of wind turbines typically increases monotonically with operating time, causing the model to be unable to adapt to this aging trend. This results in the model being overly sensitive (false alarms) in the early stages of wind turbine operation and overly insensitive (missed alarms) in later stages. Therefore, it is necessary to focus on continuous aging timelines (a monotonically increasing, continuous wear process from commissioning or the last maintenance) and discrete maintenance timelines (discontinuous state reset points composed of maintenance events). Preferably, the historical correlation matrix is the correlation matrix corresponding to the real-time operating data stream at the first moment, and the current correlation matrix is the correlation matrix corresponding to the real-time operating data stream at the second moment. Before determining the correction coefficient of each basic physical model based on the correlation similarity, the evolution determination module is further used to: read the maintenance event records of the wind turbine, determine the maintenance events that occurred in the wind turbine from the first moment to the second moment; determine whether the components corresponding to the basic physical model experienced maintenance events during the period from the first moment to the second moment; if the components corresponding to the basic physical model experienced maintenance events during the period from the first moment to the second moment, then record the occurrence time of the maintenance event as the first timestamp; if the components corresponding to the basic physical model did not experience maintenance events during the period from the first moment to the second moment, and the first timestamp is not recorded as the first timestamp. If a maintenance event occurred in a historical moment prior to the current moment, the time of the latest maintenance event in that historical moment is recorded as the first timestamp. If no maintenance event occurred in the components corresponding to the basic physical model during the period from the first moment to the second moment, and no maintenance event occurred in the historical moments prior to the first moment, the commissioning time of the wind turbine is recorded as the first timestamp. The aging factor corresponding to the basic physical model is determined based on the length of the period from the first timestamp to the second moment. The correction coefficient of each basic physical model is determined based on the correlation similarity, including: substituting the correlation similarity into the exponential part of a preset natural exponential function to obtain the distribution stability index, which is positively correlated with the correlation similarity; the correction coefficient of each basic physical model is determined based on the distribution stability index and the aging factor.
[0089] The problem of dynamic evolution can be solved by determining the effective age of the components, quantifying the evolutionary stability based on the effective age, and generating a comprehensive correction coefficient.
[0090] Determining the effective age of components can be achieved by querying maintenance event records. For each basic physical model (corresponding to a specific component), the starting point (first timestamp) of its wear clock can be dynamically determined. This starting point could be the most recent maintenance time or the commissioning time. Specifically, the maintenance event records within the time window from the first time point to the second time point are queried. If a record exists, the occurrence time of that maintenance event is recorded as the first timestamp. If not, the query continues backward through maintenance records before the first time point to find the most recent maintenance event and uses its time as the first timestamp. If no such event exists, the commissioning time of the wind turbine is used as the first timestamp.
[0091] Quantifying evolutionary stability can be achieved by comparing the current association matrix with historical association matrices and calculating the association similarity. This metric measures whether the current model and data relationship of the system follow historical evolutionary patterns; high similarity indicates stable evolution, while low similarity indicates anomalies. Specifically, the duration Δt between the second time point and the first timestamp is calculated. Then, Δt is substituted into a predefined function f(Δt) to obtain the aging factor. This function can be linear (e.g., 1 + k*Δt) or nonlinear (e.g., an exponential function or a function describing a bathtub curve).
[0092] The comprehensive generation of correction coefficients involves combining the aging factor, representing long-term aging trends, with the distribution stability index, representing short-term evolutionary stability, to generate a correction coefficient. This coefficient dynamically adjusts the original warning output of the basic physical model, enabling it to simultaneously reflect information on both how much the components have aged and whether their current state is normal. Specifically, the calculated correlation similarity (a value between 0 and 1) is substituted into a natural exponential function, for example, distribution stability index = exp(correlation similarity).
[0093] When the association similarity is high (evolutionary stability), the index value is large, indicating that the original output of the model can be trusted; when the association similarity is low (evolutionary anomaly), the index value is small (close to 1), indicating that the original output of the model needs to be suppressed because it may no longer be able to explain the current state.
[0094] Determining the correction coefficient involves combining the distribution stability index and the aging factor to integrate long-term aging trends with short-term stability assessments. For example, this can be achieved through multiplication: Correction Coefficient = Distribution Stability Index × Aging Factor. Under normal aging conditions, the distribution stability index is high, and as the aging factor increases over time, the correction coefficient will gradually increase, improving early warning sensitivity. In cases of sudden anomalies, the distribution stability index drops sharply, and even if the aging factor is high, the correction coefficient will be lowered, thus suppressing false alarms based on aging trends and indicating unmodeled anomalies on the platform.
[0095] Preferably, the correction coefficients for each basic physical model are determined based on the distribution stability index and the aging factor, including: using the distribution stability index and the aging factor as observational evidence, updating the posterior probability distribution of the preset correction coefficients through Bayes' theorem, and obtaining the updated correction coefficients.
[0096] By introducing a Bayesian framework, the process of determining the correction coefficients is transformed from an isolated, deterministic calculation into a continuous probabilistic reasoning process that integrates prior knowledge and new evidence.
[0097] A prior probability distribution P(C) is defined for the correction coefficient C. This distribution represents our belief in the possible values of C before obtaining the current observational evidence, based on historical data, engineering standards, or expert experience. Using the distribution stability index and aging factor as the observational evidence E, a likelihood function P(E|C) is defined to characterize the probability of observing the current set of evidence E if the true value of the correction coefficient is C. Applying Bayes' theorem P(C|E) ∝ P(E|C) * P(C), the prior belief P(C) is combined with the likelihood P(E|C) of the new evidence E to obtain the posterior probability distribution P(C|E). This distribution represents the updated and more accurate belief in the correction coefficient C after considering all information. A representative value (such as the expected value or the maximum a posteriori probability estimate) is extracted from the posterior probability distribution as the updated correction coefficient.
[0098] Defining a prior probability distribution can be achieved by choosing a suitable family of probability distributions to characterize the uncertainty of the correction coefficient C, such as a Gaussian distribution (normal distribution) or a gamma distribution, and setting the parameters of that distribution. For example, if a Gaussian distribution N(μ, σ²) is chosen, then the mean μ and variance σ² need to be determined. μ can be set as the theoretical design value or the historical average, while σ² reflects the degree of uncertainty we have about this prior belief.
[0099] Define the likelihood function P(E|C), where the observed evidence E is a vector containing a distribution stability index and an aging factor. A model needs to be built to describe the probability of observing evidence E given the true correction coefficient C. A simplified implementation assumes the existence of a deterministic function g(C) that maps C to the expected value of the evidence E_expected, and the observed evidence E fluctuates around E_expected with some measurement noise distribution (such as a Gaussian distribution N(0, σ_e²)). Therefore, P(E|C) = N(E; g(C), σ_e²).
[0100] The posterior probability distribution P(C|E) can be calculated using Bayes' theorem: P(C|E) = [P(E|C) * P(C)] / P(E), where P(E) is a normalization constant that ensures the integral (or summation) of the posterior distribution is 1. In practical calculations, if the prior and likelihood are chosen as conjugate distributions (such as Gaussian-Gaussian), the posterior distribution has an analytical solution, making computation efficient. For non-conjugate cases, numerical methods, such as Markov chain Monte Carlo (MCMC) or variational inference, can be used to approximate the posterior distribution.
[0101] Extract the updated correction coefficients and calculate a point estimate from the calculated posterior probability distribution P(C|E) as the final correction coefficients. Common methods include: maximum a posteriori estimation, which finds the C value that maximizes P(C|E); and posterior expectation estimation, which calculates the expected value E[C|E] of C relative to P(C|E).
[0102] For example, performing a Bayesian update on the correction coefficient C of a wind turbine bearing model.
[0103] Based on the design specifications, we believe that the ideal value of C is 1.0, but there is uncertainty. We can set the prior distribution as a Gaussian distribution: P(C) = N(μ=1.0, σ²=0.1²).
[0104] Obtaining observational evidence E: At time T2, the distribution stability index is calculated to be 2.2, and the aging factor is 1.8. A simple deterministic function g(C) = 1.5 * C can be defined, and the observation noise variance is assumed to be σ_e² = 0.2². Therefore, the expected value of the evidence should be E_expected = 1.5 * C. The likelihood function is P(E|C) = N(E_observed; 1.5*C, 0.2²).
[0105] Calculate the posterior P(C|E): Since the prior and likelihood are both Gaussian distributions, the posterior is also a Gaussian distribution, and its mean μ_posterior and variance σ_posterior² have analytical solutions.
[0106] μ_posterior = (σ_e² * μ_prior + σ_prior² * (E_observed / 1.5)) / (σ_e² + σ_prior²); therefore, μ_posterior = (0.04 * 1.0 + 0.01 * (2.0 / 1.5)) / (0.04 + 0.01) = (0.04 + 0.0133) / 0.05 = 1.066. For simplicity, we use (distribution stability index + aging factor) / 2 as E_observed, i.e., (2.2 + 1.8) / 2 = 2.0.
[0107] The posterior distribution is P(C|E) = N(μ=1.066, σ²≈0.008²).
[0108] The posterior expectation is taken as the updated correction coefficient, i.e., C_updated = 1.066.
[0109] The initial belief was C=1.0, and new observational evidence (E=2.0) suggested that C might be higher (since 2.0 / 1.5 ≈ 1.33). The Bayesian update, C=1.066, is a weighted average between the prior belief (1.0) and the evidence-suggested value (1.33). Due to the relatively large observational noise (σ_e=0.2 vs σ_prior=0.1), the platform gave a higher weight to the prior belief, and this result is much more robust than simply using 2.2 * 1.8 = 3.96 as a correction factor, as it fully accounts for the uncertainty of both the evidence and the prior.
[0110] Preferably, the fault warning module is further configured to: read the target warning data of each basic physical model; determine the comprehensive risk score of the component based on the target warning data of multiple basic physical models related to the same component; determine the comprehensive risk score of the subsystem based on the comprehensive risk scores of all components related to the same subsystem; and output a multi-level risk report based on the comprehensive risk score of the subsystem, the comprehensive risk score of the component, and the target warning data.
[0111] The unprocessed target early warning data set is a one-dimensional, unstructured list where each data point is logically equal and has no hierarchical or aggregation relationship.
[0112] Therefore, target early warning data from multiple basic physical models belonging to the same physical component are merged to generate a comprehensive risk score for the component; comprehensive risk scores from all components belonging to the same subsystem are merged to generate a comprehensive risk score for the subsystem; and the scoring data from the three levels of subsystem, component, and model are organized according to the physical hierarchy to output a structured multi-level risk report, providing maintenance personnel with a risk situation map that can be drilled down layer by layer from macro to micro.
[0113] By determining the correlation degree between each structured semantic description and each structured label, a current correlation matrix is obtained to reflect the real-time status of the wind turbine. The correlation similarity between the data distribution of the current correlation matrix and the historical correlation matrix is determined, and a correction coefficient for each basic physical model is determined based on this correlation similarity. This reflects the state changes of the wind turbine as deployment progresses, and the correction coefficient quantifies these changes. The original early warning data for each basic physical model is determined based on the real-time operational data stream. Each original early warning data is then corrected based on the correction coefficient of each basic physical model to obtain target early warning data. This ensures that the target early warning data accurately reflects the current state of the wind turbine, thereby improving the reliability of early warnings throughout the entire lifecycle of the wind turbine.
[0114] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0115] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A wind turbine intelligent early warning platform for the entire lifecycle based on big data analysis, characterized in that, include: The physical theory module is used to determine the structured labels for each basic physical model of the wind turbine. The operation status module is used to acquire the real-time operation data stream of the wind turbine and generate a structured semantic description based on the real-time operation data stream; the evolution determination module is used to: determine the degree of association between each structured semantic description and each structured label to obtain the current association matrix; determine the association similarity between the data distribution of the current association matrix and the historical association matrix; and determine the correction coefficient of each basic physical model based on the association similarity. The fault early warning module is used to: determine the original early warning data for each basic physical model based on the real-time running data stream; Each original early warning data point is corrected based on the correction coefficient of each basic physical model to obtain target early warning data. The correlation similarity is used to determine the correction coefficient of each basic physical model, including: substituting the correlation similarity into the exponential part of a preset natural index function to obtain a distribution stability index, where the distribution stability index is positively correlated with the correlation similarity; determining the correction coefficient of each basic physical model based on the distribution stability index and the aging factor; the determination of the correction coefficient of each basic physical model based on the distribution stability index and the aging factor includes: using the distribution stability index and the aging factor as observational evidence, updating the preset correction coefficient using Bayes' theorem. The posterior probability distribution is used to obtain the updated correction coefficients. The step of updating the posterior probability distribution of the preset correction coefficients using Bayes' theorem includes: setting a prior probability distribution P(C) for the correction coefficient C; defining a likelihood function P(E|C) using the distribution stability index and aging factor as observed evidence E, to characterize the probability of observing the current set of evidence E if the true value of the correction coefficient is C; applying Bayes' theorem P(C|E)∝P(E|C)*P(C) to combine the prior belief P(C) with the likelihood P(E|C) of the new evidence E to obtain the posterior probability distribution P(C|E); and extracting a representative value from the posterior probability distribution as the updated correction coefficients.
2. The platform as described in claim 1, characterized in that, The step of generating a structured semantic description based on the real-time running data stream includes: marking data streams in the real-time running data stream whose variance is less than a preset variance threshold within a preset time window as quasi-steady-state operating condition data segments; determining the autocorrelation function of the fault data sequence in the quasi-steady-state operating condition data segment; recording the time length required for the autocorrelation function to first decay to the preset variance threshold as the dynamic window length under the quasi-steady-state operating condition data segment; truncating the quasi-steady-state operating condition data segment based on the dynamic window length to obtain a fault data sequence; reconstructing the phase space of the fault data sequence to obtain a high-dimensional phase space point cloud; generating a persistent graph based on the continuous cohomology of the high-dimensional phase space point cloud; extracting topological invariants from the persistent graph to obtain an actual topological feature vector; and determining the structured semantic description based on the actual topological feature vector.
3. The platform as described in claim 2, characterized in that, The step of reconstructing the phase space of the fault data sequence to obtain a high-dimensional phase space point cloud includes: determining the mutual information function and permutation entropy function of the fault data sequence under multiple preset delay times; determining a first local minimum point in the permutation entropy function, wherein the first local minimum point is the smallest independent variable value among the multiple local minimum points of the permutation entropy function; determining a second local minimum point in the mutual information function within the domain from the preset initial value to the first local minimum point, wherein the second local minimum point is the smallest independent variable value among the multiple local minimum points of the mutual information function; and setting the second local minimum point as the smallest independent variable value among the multiple local minimum points of the mutual information function. The optimal time delay parameter is determined; the signal-to-noise ratio (SNR) of the fault data sequence is determined; an adaptive threshold for the fault data sequence is determined based on the SNR, wherein the adaptive threshold is negatively correlated with the SNR; starting from a preset minimum embedding dimension, the correlation dimension under the current embedding dimension and the relative rate of change of the correlation dimension compared to the previous correlation dimension are determined; the current embedding dimension when the relative rate of change is less than the adaptive threshold is determined as the optimal embedding dimension; the phase space of the fault data sequence is reconstructed based on the optimal time delay parameter and the optimal embedding dimension to obtain a high-dimensional phase space point cloud.
4. The platform as described in claim 3, characterized in that, The step of generating a persistent graph based on the continuous cohomology of a high-dimensional phase space point cloud includes: selecting a preset number of landmark points from the high-dimensional phase space point cloud using a preset maximum-minimum distance method; determining potential simplexes in the landmark point set; for each potential simplex: determining witness points corresponding to the potential simplex from the high-dimensional phase space point cloud; determining the set formed by removing the potential simplex from the landmark point set as the landmark point difference set corresponding to the potential simplex; if the distance from the witness point to the vertex of the potential simplex corresponding to the witness point is less than the distance from the witness point to any point in the landmark point difference set corresponding to the potential simplex, then adding the potential simplex to the Witness complex; constructing the Witness complex based on each potential simplex, and determining the continuous cohomology of the Witness complex to obtain the persistent graph.
5. The platform as described in claim 4, characterized in that, The step of extracting topological invariants from the persistent graph to obtain the actual topological feature vector includes: setting grid partitioning rules for the persistence dimension and birth time dimension of the persistent graph to obtain multiple grids, wherein the persistent graph is a two-dimensional point set, each point in the two-dimensional point set represents a topological feature, the horizontal axis of the coordinate system corresponding to the two-dimensional point set is the birth time, the vertical axis is the persistence, and the persistence is the difference between the death time and the birth time of the topological feature; filtering out topological features corresponding to a first homology group from the persistent graph, wherein the topological features of the first homology group are used to characterize one-dimensional ring structures in the data; for each topological feature on the first homology group in the persistent graph, determining the grid coordinates to which the topological feature belongs based on the birth value and death value of the topological feature; counting the number of topological features falling into each grid in the persistent graph; and arranging the statistics in the grids to obtain the actual topological feature vector.
6. The platform as described in claim 5, characterized in that, The real-time operating data stream includes at least one of a vibration data stream, a temperature data stream, a pressure data stream, and a power data stream. Before marking the data stream in the real-time operating data stream whose variance is less than a preset variance threshold within a preset time window as a quasi-steady-state operating condition data segment, the operating status module is further configured to: perform spectral analysis on the vibration data stream in the real-time operating data stream, and determine a first type of fault data sequence of the vibration data stream in the real-time operating data stream based on the energy value of a preset frequency band; perform rate of change analysis on the temperature data stream and pressure data stream in the real-time operating data stream, and determine a second type of fault data sequence of the temperature and pressure data streams in the real-time operating data stream based on the abnormal rate of change; perform efficiency curve analysis on the power data stream in the real-time operating data stream, and determine a third type of fault data sequence of the power data stream in the real-time operating data stream based on the deviation of the power curve; the step of determining the structured semantic description based on the actual topology feature vector includes: determining the structured semantic description based on the actual topology feature vector corresponding to the first type of fault data sequence, the actual topology feature vector corresponding to the second type of fault data sequence, and the actual topology feature vector corresponding to the third type of fault data sequence.
7. The platform as described in claim 6, characterized in that, The process of determining the structured labels for each basic physical model of a wind turbine includes: determining the expected fault sequence of the basic physical model based on its fault modes; determining the expected topological feature vector of the basic physical model based on the expected fault sequence; and determining the structured labels for the basic physical model based on its corresponding components, subsystems, fault modes, and expected topological feature vectors. The process of determining the correlation between each structured semantic description and each structured label to obtain a current correlation matrix includes: determining the Euclidean distance between the actual topological feature vector of each structured semantic description and the expected topological feature vector of each structured label; and determining the correlation between each structured semantic description and each structured label based on the Euclidean distance to obtain a current correlation matrix, wherein the Euclidean distance is negatively correlated with the correlation degree.
8. The platform as described in claim 7, characterized in that, The historical correlation matrix is the correlation matrix corresponding to the real-time operating data stream at the first moment, and the current correlation matrix is the correlation matrix corresponding to the real-time operating data stream at the second moment. Before determining the correction coefficient of each basic physical model based on the correlation similarity, the evolution determination module is further configured to: read the maintenance event records of the wind turbine, determine the maintenance events that occurred in the wind turbine during the period from the first moment to the second moment; and determine whether the maintenance events occurred in the components corresponding to the basic physical model during the period from the first moment to the second moment. If the maintenance event occurs in the component corresponding to the basic physical model during the period from the first time to the second time, the time of occurrence of the maintenance event is recorded as the first timestamp; if the maintenance event does not occur in the component corresponding to the basic physical model during the period from the first time to the second time, but a maintenance event occurred in a historical time before the first time, the time node of the latest maintenance event in the historical time is recorded as the first timestamp; if the maintenance event does not occur in the component corresponding to the basic physical model during the period from the first time to the second time, and no maintenance event occurred in a historical time before the first time, the commissioning time of the wind turbine is recorded as the first timestamp. The aging factor corresponding to the basic physical model is determined based on the time period from the first timestamp to the second time.
9. The platform as described in claim 8, characterized in that, The fault warning module is also used to: read the target warning data of each basic physical model; and determine the comprehensive risk score of the component based on the target warning data of multiple basic physical models related to the same component. The overall risk score of a subsystem is determined based on the overall risk score of all components related to the same subsystem. A multi-level risk report is output based on the comprehensive risk score of the subsystem, the comprehensive risk score of the component, and the target early warning data.
Citation Information
Patent Citations
Wind turbine generator fault early warning and abnormal parameter inspection method and system
CN116644343A
Transformer potential fault mode identification method and device based on reverse derivation
CN120387017A
Power grid optimization method and device based on two-stage continuous coherence
CN120764764A