Methods and systems for early warning of governor failure in hydro-generator units
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-27
- Publication Date
- 2026-08-14
AI Technical Summary
[0003]当前,针对调速器的故障预警技术主要存在以下局限性:一是多停留在状态监测与异常报警阶段,预警后的决策与调控环节严重依赖人工经验,响应滞后;二是现有方法多为“开环”系统,即故障诊断、风险评估与控制执行相互割裂,缺乏基于实时预警信息的自适应反馈调控能力
[0028]本申请通过构建水轮发电机组调速器故障的感知、诊断、预测、调控的完整技术闭环,实现了水轮发电机组调速器运维效率的提升。
Smart Images

Figure CN122565635A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent operation and maintenance and fault early warning technology for water conservancy and hydropower equipment, and in particular to a fault early warning method and system for a hydro turbine generator set governor. Background Technology
[0002] The governor of a hydro-turbine generator set is a key control device in a hydropower system, and its operating status directly affects the stability of the unit and the safety of the power grid.
[0003] Currently, fault early warning technologies for speed governors suffer from the following limitations: First, they mostly remain at the stage of condition monitoring and abnormal alarms, with decision-making and control after the warning heavily reliant on human experience, resulting in a delayed response. Second, existing methods are mostly "open-loop" systems, meaning that fault diagnosis, risk assessment, and control execution are isolated from each other, lacking adaptive feedback control capabilities based on real-time early warning information. This makes it difficult for traditional early warning systems to proactively intervene in the operating status at the nascent or early stages of a fault, failing to effectively curb performance degradation trends, and even more so, failing to form a closed-loop management system of "perception-diagnosis-decision-control".
[0004] Therefore, developing an integrated method that can link high-precision fault early warning with adaptive control strategies in real time has become an urgent need to improve the intelligent operation and maintenance level of speed governors. Summary of the Invention
[0005] To address the shortcomings of existing technologies, the purpose of this application is to provide a fault early warning method and system for hydro-generator speed governors, which can realize closed-loop intelligent operation and maintenance of hydro-generator speed governors from fault early warning to early warning and control, thereby improving their operational reliability and control autonomy.
[0006] To achieve the above objectives, this application adopts the following technical solution:
[0007] This application provides a method for early warning of faults in the governor of a hydro-generator set, the method comprising the following steps:
[0008] S101 collects real-time data from vibration sensors, pressure sensors, temperature sensors, and displacement sensors deployed on the speed governor, and uses a time synchronization protocol to timestamp the data from each sensor. At the same time, it acquires the control signals of the speed governor and the operating parameters of the unit to form an initial dataset with spatiotemporal alignment.
[0009] S102, Perform multi-level noise reduction processing on the initial dataset to obtain a pre-processed clean dataset;
[0010] S103, Based on the clean dataset, extract time-domain features, frequency-domain features, time-frequency-domain features, and nonlinear features to construct a multi-scale feature vector for characterizing the governor's operating state;
[0011] S104. Based on the multi-scale feature vector, vibration kurtosis, temperature change rate and pressure standard deviation are fused by weighting method to construct governor health status index, and the health status index is compared with the benchmark health status curve established based on historical normal data, so as to calculate the deviation between the current operating status and the benchmark health status in real time.
[0012] S105, when the deviation exceeds the preset range, a deep learning model is used to analyze the multi-scale feature vector, diagnose the fault type and occurrence probability, and combine it with the historical fault case library to perform working condition similarity matching, so as to calculate the correlation and confidence score between the current abnormal mode and the historical fault, thereby outputting a diagnostic conclusion including fault type, occurrence probability and confidence score.
[0013] S106. Based on the diagnostic conclusion, an autoregressive integral moving average model is used to predict the short-term trends of key parameters such as vibration amplitude, oil temperature and pressure fluctuations. Combined with the health status index and health status degradation model, the time function of the failure probability is calculated, and then a risk-time matrix is established to quantify the comprehensive risk value by comprehensively considering the severity of the failure, the probability of occurrence and the scope of impact.
[0014] S107 When the comprehensive risk value exceeds the preset fault warning threshold, the guide vane opening strategy and response speed parameters are adjusted according to the current head, load and fault type in the diagnostic conclusion, an adaptive control strategy is generated and executed, the control effect is monitored in real time and the control parameters are optimized through the feedback mechanism, and an adaptive control record is obtained.
[0015] As a preferred technical solution, step S101 further includes: the vibration sensor includes a triaxial vibration sensor deployed on the relay piston rod, the main pressure regulating valve core, and the electro-hydraulic converter assembly, respectively, for collecting wide-band vibration signals; the time synchronization protocol adopts a network-based precision time synchronization protocol, and integrates a hardware timestamp unit in the programmable logic device of the speed governor controller to achieve high-precision synchronization between the data of each sensor and the event sequence inside the controller; the unit operating condition parameters include at least head, active power, frequency deviation, and opening setpoint, and are obtained in real time through industrial communication protocol interaction with the monitoring system.
[0016] As a preferred technical solution, the multi-stage noise reduction process in step S102 includes: first-stage noise reduction: filtering the original signals of each sensor using a wavelet adaptive thresholding method to eliminate pulse interference; second-stage noise reduction: applying an improved adaptive noise complete set empirical mode decomposition algorithm to the vibration signal to decompose the signal into multiple intrinsic mode functions, and reconstructing the effective components based on the correlation coefficient and kurtosis criterion to remove background noise; third-stage noise reduction: performing cross-validation using causality tests and partial autocorrelation functions on the reconstructed multi-sensor data to remove random fluctuation components unrelated to the governor state, thus obtaining the clean dataset.
[0017] As a preferred technical solution, in step S103, the extracted time-frequency domain features include the time-frequency ridge energy distribution obtained by synchronous compressed wavelet transform; the extracted nonlinear features include permutation entropy, recursive graph quantitative features, and multi-scale fuzzy entropy; the construction of the multi-scale feature vector further includes: using the maximum correlation minimum redundancy algorithm to screen all the extracted initial features, and weighting them with the feature importance evaluated by random forest to form a dimension-reduced optimized multi-scale feature vector.
[0018] As a preferred technical solution, in step S104, the weighting method adopts a game theory-based combined weighting method, which optimizes the combination of subjective weights determined by the order relation analysis method and objective weights determined by the entropy weight method to determine the final fusion weights of vibration kurtosis, temperature change rate, and pressure standard deviation; the benchmark health state curve is a set of multiple operating condition benchmark curves established based on historical normal data under different head and load combinations; the real-time deviation calculation is to calculate the dynamic time warping distance between the current health state index and the benchmark curve corresponding to the current operating condition.
[0019] As a preferred technical solution, in step S105, the deep learning model is a hybrid diagnostic model consisting of a bidirectional gated recurrent neural network based on an attention mechanism and a one-dimensional convolutional neural network in parallel; the condition similarity matching based on the historical fault case library specifically involves using a dynamic time warping algorithm to calculate the morphological similarity between the current abnormal feature sequence and the fault sequences in the case library, and combining it with Euclidean distance to measure amplitude similarity, and then weighting and fusing the two as the correlation degree; the confidence score is jointly determined by the fault probability output by the hybrid diagnostic model and the correlation degree.
[0020] As a preferred technical solution, in step S106, the health status degradation model is a stochastic degradation model based on the Wiener process, and its drift coefficient and diffusion coefficient are obtained by maximum likelihood estimation using historical degradation data; the risk-time matrix is a three-dimensional matrix, whose three dimensions are the severity level of the fault, the probability interval of the fault occurrence within a specific future time window, and the subsystem range affected by the fault; the comprehensive risk value is obtained by weighted fuzzy comprehensive evaluation of the values of each unit of the three-dimensional matrix.
[0021] As a preferred technical solution, in step S107, the generation and execution of the adaptive control strategy specifically includes: based on the fault type in the diagnostic conclusion, calling a preset fuzzy rule base, the input of which is the current head, load, comprehensive risk value, and fault type, and the output is the correction coefficient of the guide vane opening change rate and the adjustment amount of the proportional-integral-derivative control parameters; the real-time monitoring and control effect is evaluated by comparing the sample entropy changes of key parameters before and after the adjustment; the feedback mechanism adopts a proximal strategy optimization algorithm based on reinforcement learning, using the reduction in the comprehensive risk value as a reward signal to optimize the conclusion part parameters of the fuzzy rule base online.
[0022] This application also provides a fault early warning system for a hydro-generator speed governor, the system comprising:
[0023] The data acquisition and synchronization module is used to acquire real-time operating sensor data through vibration sensors, pressure sensors, temperature sensors and displacement sensors deployed on the speed governor, and to align the timestamps of each sensor data using a time synchronization protocol. At the same time, it acquires the control signals of the speed governor and the operating parameters of the unit to form an initial dataset with spatiotemporal alignment.
[0024] The data processing and feature extraction module is used to perform multi-level noise reduction processing on the initial dataset to obtain a clean dataset, and extract time-domain features, frequency-domain features, time-frequency-domain features and nonlinear features based on the clean dataset to construct a multi-scale feature vector for characterizing the operating state of the speed governor.
[0025] The health assessment and fault diagnosis module is used to construct a governor health status index based on the multi-scale feature vector by fusing vibration kurtosis, temperature change rate and pressure standard deviation through a weighted method, and calculate the deviation from the benchmark health status. When the deviation exceeds a preset range, a deep learning model is used to diagnose the fault type and probability of occurrence, and the operating condition similarity is matched with the historical fault case library to output a diagnostic conclusion including fault type, probability of occurrence and confidence score.
[0026] The risk prediction and adaptive control module is used to predict the short-term trend of key parameters based on the diagnostic conclusions using an autoregressive integral moving average model, and to calculate the time function of the probability of failure by combining the health status indicators and the health status degradation model, thereby establishing a risk-time matrix to quantify the comprehensive risk value. When the comprehensive risk value exceeds a preset threshold, an adaptive control strategy is generated and executed based on the current head, load, and failure type, while the control parameters are optimized through a feedback mechanism.
[0027] Compared with the prior art, the beneficial effects of this application are as follows:
[0028] This application improves the operation and maintenance efficiency of hydro-generator governors by constructing a complete technical closed loop for the perception, diagnosis, prediction, and control of governor faults.
[0029] First, by employing multi-scale feature fusion and deep learning diagnostics, early and accurate fault identification and quantitative risk assessment are achieved, resolving the question of when to intervene. Second, this application does not merely provide alarm information after a warning; instead, based on a risk-time matrix and real-time operating conditions, it automatically generates and executes an adaptive guide vane opening control strategy, addressing the question of how to intervene and achieving a shift from passive alarm to proactive protection. Finally, by monitoring the control effect in real time and providing feedback to optimize control parameters, a closed-loop optimization mechanism is formed, significantly enhancing the system's autonomous adaptability and reliability. Ultimately, the core objective of integrating early warning and control is achieved, effectively ensuring the long-term safe and stable operation of the unit. Attached Figure Description
[0030] Figure 1 A flowchart illustrating the steps of a fault early warning method for a hydro-generator speed governor. Detailed Implementation
[0031] To enable those skilled in the art to better understand the present application, the technical solutions in specific embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0032] like Figure 1 As shown, this application provides a method for early warning of faults in the governor of a hydro-generator set, which includes the following steps:
[0033] S101 collects real-time operating sensor data through vibration sensors, pressure sensors, temperature sensors, and displacement sensors deployed on the speed governor, and uses a time synchronization protocol to timestamp the data of each sensor. At the same time, it obtains the control signals of the speed governor and the operating parameters of the unit to form an initial dataset with spatiotemporal alignment.
[0034] The vibration sensors are triaxial IEPE accelerometers, installed on the servo piston rod (measuring axial and radial vibration), the main pressure regulating valve core housing (measuring radial vibration), and the electro-hydraulic converter assembly housing (measuring three-dimensional vibration). Their range is ±50g, frequency response range is 0.5Hz to 5kHz, and sampling frequency is uniformly set to 10kHz. The pressure sensors are high-frequency response variable pressure transmitters, deployed in chambers A and B of the main pressure regulating valve and the pilot valve outlet, with a measurement range of 0-20MPa and a sampling frequency of 1kHz. The temperature sensors are PT100 platinum resistance thermometers, installed on the electro-hydraulic converter coil, main pressure regulating valve body, and bottom of the pressure tank, with a sampling frequency of 10Hz. The displacement sensors are magnetostrictive linear displacement sensors, measuring the servo travel with an accuracy of ±0.1%FS and a sampling frequency of 100Hz.
[0035] Time synchronization is achieved through a Precision Time Protocol (PTP) network based on the IEEE 1588 standard. Within the governor controller, an FPGA with hardware timestamp functionality (such as the Xilinx Zynq-7000 series) is used as the data acquisition and synchronization unit. Each sensor signal, after being conditioned by a circuit, is connected to a dedicated I / O pin on the FPGA, and is precisely timestamped (accuracy up to 100 nanoseconds) by the FPGA's PTP hardware clock at the moment of acquisition. Control signals (such as guide vane opening setpoint and PID output) and unit operating parameters (head, active power, frequency deviation, and opening setpoint) obtained from the monitoring system via Modbus / TCP protocol are also received synchronously and timestamped with the same time base. All timestamped data is packaged into fixed time windows (e.g., 1 second) to form a spatiotemporally strictly aligned initial dataset buffer.
[0036] This approach ensures strict temporal alignment of multi-source heterogeneous data, laying a solid foundation for subsequent joint feature extraction and state analysis, and avoiding analytical errors caused by time asynchrony.
[0037] S102, perform multi-level noise reduction on the initial dataset to obtain a preprocessed clean dataset.
[0038] Multi-stage noise reduction processing includes three stages of noise reduction.
[0039] First-stage noise reduction: The original signals from each sensor are filtered using a wavelet-based adaptive thresholding method to eliminate impulse interference.
[0040] Specifically, the algorithm steps of the first-level noise reduction (wavelet adaptive thresholding method) are as follows: 1. Perform N-level wavelet decomposition on the original signal x(t) (using the 'db4' wavelet basis, N=3) to obtain the detail coefficients of each level. and approximation coefficients 2. For each level of detail coefficient Calculate the adaptive threshold Where σj is the median absolute deviation (MAD) estimate of the detail coefficients at level j. N is the signal length. 3. Use a soft thresholding function to handle detail coefficients: 4. Use the processed coefficients. and approximation coefficients Wavelet reconstruction is performed to obtain the first-level denoised signal.
[0041] Second-stage noise reduction: For vibration signals, an improved adaptive noise complete set empirical mode decomposition (ICEEMDAN) algorithm is applied to decompose the signal into multiple intrinsic mode functions (IMFs), and the effective components are reconstructed based on the correlation coefficient and kurtosis criterion to remove background noise.
[0042] Specifically, the second-stage noise reduction (improved ICEEMDAN) targets vibration signals. The algorithm steps are as follows: 1. Define operators This represents the k-th intrinsic mode function (IMF) obtained after EMD decomposition. 2. Calculate the residual after adding specific noise: Let... To satisfy the standard normal distribution of white noise, a noise amplitude coefficient is added. The signal after the first addition of noise is: 3. For each Perform EMD decomposition to obtain the first IMF population average: Where I represents the number of noise additions (usually set to 100). 4. Calculate the first-order residual: 5. In subsequent stages, higher-order IMFs are iteratively calculated by adding noise to the residuals until the residuals become a monotonic function. This ultimately yields a series of IMF components. 6. Reconstructing effective components: Calculate the relationship between each IMF component and the original signal. correlation coefficient and its kurtosis value Set threshold (Original signal kurtosis). Preserve the satisfaction... and The IMF components are summed to obtain the secondary noise-reduced signal.
[0043] The third stage of noise reduction involves cross-validating the reconstructed multi-sensor data using Granger causality test and partial autocorrelation function (PACF) to eliminate random fluctuation components unrelated to the governor's state.
[0044] Specifically, the algorithm steps for the third-level noise reduction are as follows: 1. Process multi-sensor data (such as vibration data) 1. Standardize the pressure p(t) and displacement d(t). 2. Calculate the partial autocorrelation function (PACF) for the pressure signal p(t). If the PACF truncates rapidly after lag orders 1 and 2 (falling within the confidence interval), it indicates that the current pressure fluctuation is mainly affected by its historical values. 3. Perform a Granger causality test. Using pressure p(t) as the dependent variable, the oscillation... Using displacement d(t) as independent variables, a vector autoregression (VAR) model is constructed. An F-test is performed. If the p-value is greater than the significance level (e.g., 0.05), the null hypothesis is accepted, and it is considered that the past values of vibration and displacement cannot significantly explain the current changes in pressure, i.e., they have no Granger causal relationship. 4. Combining the PACF and Granger test results, if the pressure fluctuation is mainly determined by its own history and has no significant causal relationship with other signals, its autoregressive model (e.g., AR(2)) is used for one-step prediction, and the residual between the actual value and the predicted value is regarded as random fluctuation and smoothed (e.g., using moving average) to obtain the final clean pressure signal. Similar processing was performed on other signals to ultimately obtain a clean dataset.
[0045] This multi-stage progressive noise reduction strategy can selectively filter out different types of noise, significantly improve the signal-to-noise ratio, and thus obtain clean data that better reflects the true state of the equipment, creating conditions for accurate feature extraction.
[0046] S103, based on the clean dataset, extracts time-domain features, frequency-domain features, time-frequency-domain features, and nonlinear features to construct a multi-scale feature vector for characterizing the governor's operating state.
[0047] The extracted time-domain features include at least one of the following: mean, variance, standard deviation, peak-to-peak value, root square amplitude, average amplitude, kurtosis, skewness, waveform factor, peak factor, impulse factor, and margin factor.
[0048] Frequency domain features include the centroid frequency, mean square frequency, and frequency variance obtained through Fast Fourier Transform (FFT). Frequency domain features are calculated after performing an FFT on the signal to obtain the spectrum X(f).
[0049] The time-frequency domain features include the time-frequency ridge energy distribution obtained by synchronous compressed wavelet transform (SST), specifically by calculating the energy integral of the wavelet coefficient modulus along the ridge in the time-frequency plane.
[0050] Time-frequency domain features (based on Synchronous Compressed Wavelet Transform (SST): Perform a continuous wavelet transform (CWT) on the signal to obtain wavelet coefficients. Where 'a' represents the scale and 'b' represents the translation. Calculate the instantaneous frequency at each point (a, b): Synchronous compression is achieved by mapping wavelet coefficients from the scale-time plane (a, b) to the frequency-time plane (ω, b). Where ωl is the frequency center after discretization. Extracting the time-frequency ridge: in the time-frequency plane... Above, for each time point b, find the frequency point with the highest energy. Calculate the energy distribution characteristics of the time-frequency ridge: Using the ridge frequency as the center, take a certain bandwidth (e.g., Calculate the proportion of energy in the total energy within the banded region, and statistically analyze the mean and variance of this proportion over time.
[0051] The extracted nonlinear features include permutation entropy (PE, embedding dimension m=3, delay time τ=1), quantitative features of the recursive graph (such as determinism and laminarity), and multiscale fuzzy entropy (MSE, scale factor of 5).
[0052] Permutation entropy (PE): Given a time series Reconstructing the phase space yields a vector group, where each vector... Where the embedding dimension m=3, and the delay time Arrange the m values in each vector in ascending order to obtain a symbol sequence (permutation pattern) π. Calculate the frequency p(π) of each permutation pattern π, then the permutation entropy is obtained. After normalization, we get
[0053] Quantitative characteristics of recursion graphs: Calculate determinism (DET), laminar flow (LAM), recursion rate (RR), etc., based on the recursion graph.
[0054] Multiscale fuzzy entropy (MSE): This coarsens the original sequence to obtain a scale factor of... The new sequence is then calculated, and the fuzzy entropy of each coarse-grained sequence is calculated. In the fuzzy entropy calculation, the similarity tolerance r is taken as 0.2 times the standard deviation, and the embedding dimension m=2.
[0055] Constructing multi-scale feature vectors also includes: using the maximum correlation minimum redundancy (mRMR) algorithm to filter all the extracted initial features, and weighting them with the feature importance evaluated by random forest to form a dimension-reduced optimized multi-scale feature vector.
[0056] Suppose that the initial feature set F has M features, and the goal is to select m features from it (m=20).
[0057] mRMR screening: The maximum correlation and minimum redundancy criterion is adopted. First, the feature with the highest mutual information with the Health Status Index (HI) is identified. As the first feature. Then iterative selection, at step k, from the remaining feature set. Select features Make it satisfy: in Let S represent mutual information, and S be the set of selected features. Repeat this process until m features are selected, forming a feature subset.
[0058] Random Forest Weighting: A regression model (predicting HI values) is trained on a historical normal dataset using a random forest model containing 100 decision trees. After training, the average reduction in impurity for each feature across all trees (Gini importance) is calculated and normalized to obtain the feature importance weight vector.
[0059] Final feature vector: Features selected by mRMR According to their corresponding random forest importance weights Weighting is performed to form the final optimized multi-scale feature vector. Where ⊙ represents element-wise multiplication.
[0060] This method comprehensively characterizes the operational status at multiple scales and optimizes the feature space through feature selection and weighting, which not only preserves high-value information but also reduces data dimensionality, thereby improving the computational efficiency and generalization ability of subsequent models.
[0061] S104, based on multi-scale feature vectors, uses a weighted method to fuse vibration kurtosis, temperature change rate and pressure standard deviation to construct governor health status indicators. It then compares the health status indicators with a benchmark health status curve established based on historical normal data, thereby calculating the deviation between the current operating status and the benchmark health status in real time.
[0062] The weighting method adopts a game theory-based combination weighting method, which optimizes the subjective weights determined by the order relation analysis method (G1 method) and the objective weights determined by the entropy weight method to determine the final fusion weights of vibration kurtosis, temperature change rate and pressure standard deviation.
[0063] First, experts, based on experience, use the G1 method to assign a subjective weight vector Ws = [0.5, 0.3, 0.2] (corresponding to vibration kurtosis, temperature change rate, and pressure standard deviation, respectively). Second, based on historical normal data, the information entropy of each indicator is calculated, resulting in an objective weight vector Wo = [0.6, 0.25, 0.15]. Finally, with the objective of minimizing the difference between the two weight allocation results, an optimization model is solved to obtain the combined weight W = αWs + (1 - α)Wo. After calculation, α = 0.6, and the final W = [0.54, 0.28, 0.18].
[0064] Health Status Index (HI) Calculation: for vibration kurtosis Rate of temperature change Pressure standard deviation Normalize the values so that they fall within the range [0, 1]. The normalization method uses minimum-maximum scaling. in and These are the minimum and maximum values of this indicator in historical normal data. The combined weights W=[0.54,0.28,0.18] are used to calculate HI:
[0065] The baseline health state curves are multiple baseline curves established based on historical normal data under different head and load combinations.
[0066] Specifically, the historical normal operation data is divided into 3×3=9 operating condition cells according to the head interval (H1, H2, H3) and the load interval (P1, P2, P3). For the historical HI time series within each operating condition cell, an exponentially weighted moving average (EWMA) is used for smoothing to obtain the baseline curve for that operating condition. Where i and j represent the head and load interval indices, respectively. EWMA formula: The smoothing factor λ is set to 0.1.
[0067] Real-time deviation calculation involves determining the dynamic time warping (DTW) distance between the current health status indicator and the baseline curve corresponding to the current operating condition. The DTW distance effectively aligns and measures the shape differences between two time series, even in the presence of nonlinear deformation.
[0068] Deviation (DTW distance) calculation: Let the HI sequence of the current time window (e.g., the most recent 5 minutes) be... The reference curve sequence corresponding to the current operating condition is Construct an m×n cumulative cost matrix D. Initialization: Recursive matrix filling: The final DTW distance (deviation) is the square root of D(m, n). For easier comparison, the distance can be standardized by dividing by the sequence length. A deviation threshold can be set. (e.g., 3 times the standard deviation of the DTW distance calculated from historical normal data). When the deviation exceeds this threshold, the fault diagnosis process is triggered.
[0069] This method constructs a comprehensive health "barometer" by integrating multiple sensitive indicators, and achieves early and sensitive detection of state deviations by dynamically comparing it with historical health benchmarks under the same operating conditions, providing a quantitative basis for fault early warning.
[0070] S105. When the deviation exceeds the preset range, a deep learning model is used to analyze the multi-scale feature vectors, diagnose the fault type and probability of occurrence, and combine the historical fault case library to perform working condition similarity matching to calculate the correlation and confidence score between the current abnormal mode and the historical faults, thereby outputting a diagnostic conclusion that includes the fault type, probability of occurrence and confidence score.
[0071] The deep learning model is a hybrid diagnostic model that combines a bidirectional gated recurrent neural network (BiGRU) based on the attention mechanism with a one-dimensional convolutional neural network (1D-CNN).
[0072] Specifically, the hybrid deep learning model structure is as follows: 1. Input: A normalized sequence of multi-scale feature vectors (time step L, feature dimension D=20). 2. 1D-CNN branch: Convolutional layer 1: 32 filters, 3 kernels, 1 stride, ReLU activation, output shape (L-2, 32). Convolutional layer 2: 64 filters, 5 kernels, 1 stride, ReLU activation, output shape (L-6, 64). Global max pooling layer: outputs a 64-dimensional vector. 3. BiGRU branch: BiGRU layer: 128 hidden units, outputs the forward and backward hidden states at each time step, concatenates the forward hidden state of the last time step with the backward hidden state of the first time step to obtain a 256-dimensional vector. 4. Feature concatenation and attention mechanism: Concatenate the 64-dimensional vector from the CNN branch with the 256-dimensional vector from the BiGRU branch to obtain a 320-dimensional feature vector. The attention weights are calculated using an attention layer: Let the concatenated feature vector be... The weights are calculated using a fully connected layer (output dimension 1) and a Softmax function. The weighted context vector is 5. Output Layer: The context vector c is passed through a fully connected layer (output dimension equal to the number of fault categories K) and a Softmax activation function to output the probability distribution of each fault category (such as main brake jamming, electro-hydraulic converter coil failure, oil circuit blockage, etc.).
[0073] The system combines historical failure case databases to perform operating condition similarity matching.
[0074] Specifically, the historical fault case database stores multi-scale feature vector sequences (labeled with fault types) within a time window when historical faults occurred. Morphological similarity (DTW): Let the current abnormal feature sequence be ( A certain fault sequence in the case library is ( Calculate its DTW distance Morphological similarity is defined as: Where γ is a scaling parameter that controls the rate of similarity decay. Amplitude similarity (Euclidean distance): Calculates the mean Euclidean distance between corresponding feature vectors of two sequences. Amplitude similarity is defined as: Where β is the scaling parameter. Overall correlation: Find the top N cases (e.g., N=3) in the case library that have the highest correlation with the current sequence.
[0075] Let the fault type diagnosed by the hybrid deep learning model be k, and its probability be... The highest correlation among similar failure cases matched from the case library is The confidence score for fault type k is then: The final diagnostic conclusion output is: fault type k, probability of occurrence. Confidence score
[0076] This hybrid diagnostic model combines the local feature extraction capabilities of CNNs with the temporal modeling capabilities of BiGRUs, and improves diagnostic accuracy by focusing on key information through an attention mechanism. Cross-validation through case matching further enhances the interpretability and reliability of diagnostic conclusions, effectively reducing the false positive rate.
[0077] S106. Based on the diagnostic conclusions, an autoregressive integral moving average model is used to predict the short-term trends of key parameters such as vibration amplitude, oil temperature and pressure fluctuations. Combined with health status indicators and health status degradation models, the time function of the failure probability is calculated, and then a risk-time matrix is established to quantify the comprehensive risk value by comprehensively considering the severity of the failure, the probability of occurrence and the scope of impact.
[0078] For short-term trend prediction of key parameters, an autoregressive integral moving average (ARIMA) model is used. Taking the amplitude of vibration as an example, it is first stabilized by differencing, and then the model order (p, d, q) is determined based on the autocorrelation plot (ACF) and partial autocorrelation plot (PACF). For example, an ARIMA(2,1,1) model is established to predict the values of the next 10 sampling points (e.g., the next second).
[0079] Specifically, ARIMA(p, d, q) models are established for key parameters (such as vibration amplitude V, oil temperature T, and pressure fluctuation P'). Taking vibration amplitude V as an example: First, an ADF test is performed. If the sequence is non-stationary, differencing (order d) is performed. The orders of p and q are determined by observing the autocorrelation function (ACF) and partial autocorrelation function (PACF) plots of the differencing sequence. For example, if the ACF tails and the PACF truncates after order p, the AR(p) model is considered; if the ACF truncates after order q and the PACF tails, the MA(q) model is considered. The optimal (p, d, q) combination is determined using the AIC / BIC criteria. Model form: Where B is the lag operator. The noise is white. The model is used to predict the next h steps (e.g., the next 10 sampling points), yielding the predicted value. and its prediction range.
[0080] The health status degradation model is a stochastic degradation model based on the Wiener process, and its drift coefficient λ and diffusion coefficient σ are obtained by maximum likelihood estimation using historical degradation data.
[0081] Specifically, health status degradation modeling (Wiener process): This assumes a degradation process in the health indicator HI. Obeying the Wiener process: Where λ is the drift coefficient, Let B(t) be the diffusion coefficient, and B(t) be the standard Brownian motion. Based on historical HI degradation data from normal to fault (assuming n samples, each sample having an observation value of α °j at time point [i]), the parameters are estimated using maximum likelihood estimation (MLE): Where mi is the number of observations for the i-th sample. Let the failure threshold be D (determined based on engineering experience or statistics, such as H). Given the current health status The probability of reaching the failure threshold for the first time within a future time interval Δt (i.e., the probability of failure) is: Where φ(·) is the standard normal distribution function. For simplification, the first-arrival time approximation is often used: .
[0082] The risk-time matrix is a three-dimensional matrix. Its three dimensions are the severity level of the fault (e.g., divided into 5 levels, S1 minor, S2 moderate, S3 severe, S4 severe, S5 catastrophic), the probability interval of the fault occurrence within a specific future time window (divided into 3 intervals, e.g., T1=[0,1h), T2=[1h,6h), T3=[6h,24h]) (e.g., [0,0.1), [0.1,0.5), [0.5,1.0]), and the subsystem scope affected by the fault (e.g., divided into 3 levels, I1 (local components), I2 (governor subsystem), I3 (hydro turbine generator set)).
[0083] The comprehensive risk value is obtained by weighted fuzzy comprehensive evaluation of the values of each unit in the three-dimensional matrix (predefined by expert experience), representing the basic risk value under severity s, time window t, and impact range i, with a range of, for example, 0-10.
[0084] Fuzzy Comprehensive Assessment: 1. Factor Set: U = {Severity S, Time Probability T, Scope of Influence I}. 2. Evaluation Set: V = {Low Risk, Medium Risk, High Risk}. 3. Weight Vector: (Subjective setting, adjustable). 4. Membership Matrix R: For the currently diagnosed fault, calculate its membership degree to each factor. For severity S, directly map it to a preset level based on the fault type, using triangular fuzzy numbers (e.g., S3 corresponds to (0, 0.5, 1.0)). For time probability T, use a degradation model to calculate the probability of the fault occurring within the next three time windows. It is then categorized into "high," "medium," and "low" probability intervals, and its membership degree is calculated. For influence range I, its influence range level is preset according to the fault type, and its membership degree is calculated. 5. Fuzzy synthesis: Calculate the fuzzy comprehensive evaluation vector. Where “.” represents a fuzzy composition operator (such as a weighted average type). 6. Defuzzification: Using the weighted average method, assign values to the comment set V (e.g., low=30, medium=60, high=90), and calculate the comprehensive risk value. Obtain a value between 0 and 100. Set the warning threshold. (e.g., 70).
[0085] This method combines short-term parameter prediction, medium- and long-term health degradation trend prediction, and multidimensional risk factors to achieve dynamic and forward-looking risk assessment from qualitative to quantitative, providing precise time windows and risk level basis for early warning decisions.
[0086] S107 When the comprehensive risk value exceeds the preset fault warning threshold, the guide vane opening strategy and response speed parameters are adjusted according to the current head, load and fault type in the diagnostic conclusion. An adaptive control strategy is generated and executed. At the same time, the control effect is monitored in real time and the control parameters are optimized through the feedback mechanism, and an adaptive control record is obtained.
[0087] The specific steps for generating and executing the adaptive control strategy include: Based on the fault type in the diagnostic conclusion, calling a preset fuzzy rule base. The inputs to the rule base are the current head, load, comprehensive risk value, and fault type. The outputs are the correction coefficient K of the guide vane opening change rate and the adjustment amounts (ΔKp, ΔKi, ΔKd) of the proportional-integral-derivative (PID) control parameters (Kp, Ki, Kd).
[0088] For example, a rule could be: "IF the fault type is slight jamming of the main and auxiliary components AND high risk value AND high load, then reduce the opening change rate (K=0.8) and fine-tune the PID parameters (ΔKp=+0.1, ΔKi=-0.05, ΔKd=+0.05) to reduce the response speed."
[0089] Real-time monitoring of the control effect is evaluated by comparing the changes in sample entropy of key parameters (such as relay stroke fluctuation and frequency deviation) before and after adjustment. A decrease in sample entropy indicates a reduction in sequence complexity, a more stable control, and a good effect.
[0090] Calculate Sample Entropy: 1. For a time series {u(i)} of length N, given the embedding dimension m (usually 2) and the similarity tolerance r (usually 0.2 times the standard deviation of the series). 2. Construct an m-dimensional vector: X_m(i) = [u(i), u(i+1), ..., u(i+m-1)], 1 ≤ i ≤ Nm. 3. Define the distance between X_m(i) and X_m(j) as the maximum absolute value of the difference between their corresponding elements. 4. Count the number of vector pairs that satisfy the distance less than r, denoted as B_i. Calculate its average value B^m(r). 5. Similarly, calculate B_i^{m+1}(r) and its average value B^{m+1}(r) for the m+1 dimensional vector. 6. The sample entropy is estimated as: SampEn = -ln[B^{m+1}(r) / B^m(r)].
[0091] The feedback mechanism adopts the Proximal Policy Optimization (PPO) algorithm based on reinforcement learning, using the reduction in the comprehensive risk value R as the reward signal to optimize the parameters (such as K, ΔKp, etc.) of the conclusion part of the fuzzy rule base online.
[0092] The feedback optimization mechanism based on PPO consists of: State: Current head, load, fault type, comprehensive risk value, and sample entropy of key parameters. Action: Fine-tuning some parameters of the fuzzy rule base conclusion (such as the center value of the output fuzzy set). For example, adjusting the fuzzy set center value corresponding to "correction coefficient K = small" in the rule conclusion. Reward: Primarily based on the reduction in the comprehensive risk value R. The reward function is defined as: in and These are the risk values calculated before and after optimization. η is the action taken (parameter adjustment), and η is the penalty coefficient to prevent the action from being too large. The PPO algorithm uses the PPO-Clip algorithm based on the Actor-Critic framework. The Actor network (policy network πθ) outputs the probability distribution of actions. The Critic network (value network Vφ) evaluates the value of the state. The objective function is: in It is the probability ratio of the new strategy to the old strategy. The advantage function is estimated (calculated using the generalized advantage estimation method GAE), and ε is a hyperparameter (e.g., 0.2). By continuously interacting with the environment (the speed controller system), the PPO algorithm learns and optimizes the policy network parameters θ, thereby dynamically adjusting the fuzzy rule base to maximize long-term cumulative rewards, i.e., achieving optimal suppression of overall risk.
[0093] This adaptive control strategy realizes a closed loop from fault diagnosis and risk assessment to control execution. It can proactively adjust the control logic in the early stage of faults to suppress state deterioration, and continuously optimize the control effect through self-learning to improve the robustness and operational stability of the speed governor under abnormal operating conditions.
[0094] This application also provides a fault early warning system for a hydro-generator speed governor, the system comprising:
[0095] The data acquisition and synchronization module is used to acquire real-time operating sensor data through vibration sensors, pressure sensors, temperature sensors and displacement sensors deployed on the speed governor, and to align the timestamps of each sensor data using a time synchronization protocol. At the same time, it acquires the control signals of the speed governor and the operating parameters of the unit to form an initial dataset with spatiotemporal alignment.
[0096] The data processing and feature extraction module is used to perform multi-level noise reduction on the initial dataset to obtain a clean dataset, and extract time-domain features, frequency-domain features, time-frequency-domain features, and nonlinear features based on the clean dataset to construct a multi-scale feature vector to characterize the operating state of the speed governor.
[0097] The health assessment and fault diagnosis module is used to construct governor health status indicators based on multi-scale feature vectors, and to calculate the deviation from the baseline health status by fusing vibration kurtosis, temperature change rate, and pressure standard deviation using a weighted method. When the deviation exceeds a preset range, a deep learning model is used to diagnose the fault type and probability of occurrence, and the system performs operating condition similarity matching by combining a historical fault case library, outputting a diagnostic conclusion that includes the fault type, probability of occurrence, and confidence score.
[0098] The risk prediction and adaptive control module is used to predict the short-term trends of key parameters based on diagnostic conclusions using an autoregressive integral moving average model. It also combines health status indicators and a health status degradation model to calculate the time function of the failure probability, establishing a risk-time matrix to quantify the comprehensive risk value. When the comprehensive risk value exceeds a preset threshold, an adaptive control strategy is generated and executed based on the current head, load, and failure type, while simultaneously optimizing control parameters through a feedback mechanism.
[0099] It should be noted that the terms "first," "second," and similar terms used in this application specification and claims do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, "an" or "a" and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. "A plurality" or "several" indicates at least two. Unless otherwise stated, terms such as "front," "back," "left," "right," "lower," and / or "upper" are for illustrative purposes only and are not limited to a location or spatial orientation. Terms such as "comprising" or "including" indicate that the elements or objects preceding "comprising" encompass the elements or objects listed following "comprising" or "including" and their equivalents, and do not exclude other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect.
[0100] The singular forms “a,” “the,” and “the” used in this application specification and appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0101] It should be understood that those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A method for early warning of governor faults in a hydro-generator set, characterized in that, Includes the following steps: S101 collects real-time data from vibration sensors, pressure sensors, temperature sensors, and displacement sensors deployed on the speed governor, and uses a time synchronization protocol to timestamp the data from each sensor. At the same time, it acquires the control signals of the speed governor and the operating parameters of the unit to form an initial dataset with spatiotemporal alignment. S102, Perform multi-level noise reduction processing on the initial dataset to obtain a pre-processed clean dataset; S103, Based on the clean dataset, extract time-domain features, frequency-domain features, time-frequency-domain features, and nonlinear features to construct a multi-scale feature vector for characterizing the governor's operating state; S104. Based on the multi-scale feature vector, vibration kurtosis, temperature change rate and pressure standard deviation are fused by weighting method to construct governor health status index, and the health status index is compared with the benchmark health status curve established based on historical normal data, so as to calculate the deviation between the current operating status and the benchmark health status in real time. S105, when the deviation exceeds the preset range, a deep learning model is used to analyze the multi-scale feature vector, diagnose the fault type and occurrence probability, and combine it with the historical fault case library to perform working condition similarity matching, so as to calculate the correlation and confidence score between the current abnormal mode and the historical fault, thereby outputting a diagnostic conclusion including fault type, occurrence probability and confidence score. S106. Based on the diagnostic conclusion, the autoregressive integral moving average model is used to predict the short-term trend of key parameters such as vibration amplitude, oil temperature and pressure fluctuation. Combined with the health status index and health status degradation model, the time function of the failure probability is calculated, and then a risk-time matrix is established to quantify the comprehensive risk value by comprehensively considering the severity of the failure, the probability of occurrence and the scope of impact. S107 When the comprehensive risk value exceeds the preset fault warning threshold, the guide vane opening strategy and response speed parameters are adjusted according to the current head, load and fault type in the diagnostic conclusion, an adaptive control strategy is generated and executed, the control effect is monitored in real time and the control parameters are optimized through the feedback mechanism, and an adaptive control record is obtained.
2. The method for early warning of governor faults in a hydro-generator unit according to claim 1, characterized in that, Step S101 further includes: the vibration sensor includes a triaxial vibration sensor deployed on the relay piston rod, the main pressure regulating valve core, and the electro-hydraulic converter assembly, respectively, for collecting wide-band vibration signals; the time synchronization protocol adopts a network-based precision time synchronization protocol, and integrates a hardware timestamp unit in the programmable logic device of the governor controller to achieve high-precision synchronization between the data of each sensor and the event sequence inside the controller; the unit operating condition parameters include at least head, active power, frequency deviation, and opening setpoint, and are obtained in real time through industrial communication protocol interaction with the monitoring system.
3. The method for early warning of governor faults in a hydro-generator unit according to claim 1 or 2, characterized in that, The multi-stage noise reduction process in step S102 includes: First-stage noise reduction: filtering the original signals of each sensor using a wavelet adaptive thresholding method to eliminate impulse interference; Second-stage noise reduction: applying an improved adaptive noise complete set empirical mode decomposition algorithm to the vibration signal to decompose the signal into multiple intrinsic mode functions, and reconstructing the effective components based on the correlation coefficient and kurtosis criterion to remove background noise; Third-stage noise reduction: performing cross-validation using causality tests and partial autocorrelation functions on the reconstructed multi-sensor data to remove random fluctuation components unrelated to the governor state, thus obtaining the clean dataset.
4. The method for early warning of governor faults in a hydro-generator unit according to claim 1, characterized in that, In step S103, the extracted time-frequency domain features include the time-frequency ridge energy distribution obtained by synchronous compressed wavelet transform; the extracted nonlinear features include permutation entropy, recursive graph quantitative features, and multi-scale fuzzy entropy; the construction of the multi-scale feature vector also includes: using the maximum correlation minimum redundancy algorithm to screen all the extracted initial features, and weighting them with the feature importance evaluated by random forest to form a dimension-reduced optimized multi-scale feature vector.
5. The method for early warning of governor faults in a hydro-generator unit according to claim 1 or 4, characterized in that, In step S104, the weighting method adopts a game theory-based combination weighting method, which optimizes the subjective weights determined by the order relation analysis method and the objective weights determined by the entropy weight method to determine the final fusion weights of vibration kurtosis, temperature change rate and pressure standard deviation. The baseline health status curves are multiple operating condition baseline curves established based on historical normal data under different head and load combinations. The real-time deviation calculation is the dynamic time-normalized distance between the current health status index and the baseline curve corresponding to the current working condition.
6. The method for early warning of governor faults in a hydro-generator unit according to claim 5, characterized in that, In step S105, the deep learning model is a hybrid diagnostic model consisting of a bidirectional gated recurrent neural network based on an attention mechanism and a one-dimensional convolutional neural network in parallel. The step of performing condition similarity matching by combining the historical fault case library specifically involves using a dynamic time warping algorithm to calculate the morphological similarity between the current abnormal feature sequence and the fault sequences in the case library, and combining it with Euclidean distance to measure amplitude similarity, and then weighting and fusing the two as the correlation degree. The confidence score is jointly determined by the fault probability output by the hybrid diagnostic model and the correlation degree.
7. The method for early warning of governor faults in a hydro-generator unit according to claim 1, characterized in that, In step S106, the health status degradation model is a stochastic degradation model based on the Wiener process, and its drift coefficient and diffusion coefficient are obtained by maximum likelihood estimation using historical degradation data; the risk-time matrix is a three-dimensional matrix, whose three dimensions are the severity level of the fault, the probability interval of the fault occurrence within a specific future time window, and the subsystem range affected by the fault; the comprehensive risk value is obtained by weighted fuzzy comprehensive evaluation of the values of each unit of the three-dimensional matrix.
8. The method for early warning of governor faults in a hydro-generator unit according to claim 1, characterized in that, In step S107, the generation and execution of the adaptive control strategy specifically includes: based on the fault type in the diagnostic conclusion, calling a preset fuzzy rule base, the input of which is the current head, load, comprehensive risk value, and fault type, and the output is the correction coefficient of the guide vane opening change rate and the adjustment amount of the proportional-integral-derivative control parameters; the real-time monitoring and control effect is evaluated by comparing the sample entropy changes of key parameters before and after the adjustment; the feedback mechanism adopts a proximal policy optimization algorithm based on reinforcement learning, using the reduction in the comprehensive risk value as a reward signal to optimize the conclusion part parameters of the fuzzy rule base online.
9. A fault early warning system for a hydro-generator speed governor, characterized in that, include: The data acquisition and synchronization module is used to acquire real-time operating sensor data through vibration sensors, pressure sensors, temperature sensors and displacement sensors deployed on the speed governor, and to align the timestamps of each sensor data using a time synchronization protocol. At the same time, it acquires the control signals of the speed governor and the operating parameters of the unit to form an initial dataset with spatiotemporal alignment. The data processing and feature extraction module is used to perform multi-level noise reduction processing on the initial dataset to obtain a clean dataset, and extract time-domain features, frequency-domain features, time-frequency-domain features and nonlinear features based on the clean dataset to construct a multi-scale feature vector for characterizing the operating state of the speed governor. The health assessment and fault diagnosis module is used to construct a governor health status index based on the multi-scale feature vector by fusing vibration kurtosis, temperature change rate and pressure standard deviation through a weighted method, and calculate the deviation from the benchmark health status. When the deviation exceeds a preset range, a deep learning model is used to diagnose the fault type and probability of occurrence, and the operating condition similarity is matched with the historical fault case library to output a diagnostic conclusion including fault type, probability of occurrence and confidence score. The risk prediction and adaptive control module is used to predict the short-term trend of key parameters based on the diagnostic conclusions using an autoregressive integral moving average model, and to calculate the time function of the failure probability by combining the health status indicators and the health status degradation model, and to establish a risk-time matrix to quantify the comprehensive risk value. When the overall risk value exceeds the preset threshold, an adaptive control strategy is generated and executed based on the current head, load and fault type, while the control parameters are optimized through a feedback mechanism.