An Early Fault Location Method for Complex Data Structures in Large Hydropower Units

By performing outlier detection, standardization processing, and sliding window correlation analysis on multi-source data from large hydro turbine units, and dynamically adjusting thresholds to construct fault feature vectors, the problems of data dimension differences and ambiguous location were solved, enabling accurate fault identification and rapid location, and ensuring the safe and stable operation of the units.

CN120354311BActive Publication Date: 2026-01-06CHINA YANGTZE POWER
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510636025.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2026-01-06
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

The multi-source data of large hydro turbine units have differences in dimensions, inconsistent data quality, and ambiguous fault location, which leads to misjudgment and inaccurate fault location in fault monitoring, affecting the safe and stable operation of the unit.

Method used

The system eliminates differences in data dimensions by outlier detection, linear interpolation imputation, and standardization. It uses a sliding window method to calculate parameter correlation, dynamically adjusts the fault detection threshold, constructs a multi-dimensional fault feature vector, and locates the fault location based on weighted scores.

Benefits of technology

It significantly improved the accuracy of fault early warning and location, reduced the false alarm rate, and ensured the safe and stable operation of the unit and maintenance efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354311B_ABST
    Figure CN120354311B_ABST
Patent Text Reader

Abstract

The application discloses a large-scale water turbine set complex data structure-oriented early fault positioning method, belongs to the technical field of water turbine set fault detection, and comprises the following steps: collecting 52 kinds of key operation parameters, covering multiple types of data such as rotating speed, power and vibration; classifying the complex data structure according to collection positions; and performing preprocessing such as abnormal value processing, missing value filling and normalization, so as to eliminate dimension differences and improve data quality. A sliding window is used to calculate a Pearson correlation coefficient, real-time monitoring of parameter correlation change is performed, a dynamic threshold is set based on historical data standard deviation, and unit abnormalities are accurately judged. After detecting abnormalities, a fault characteristic vector containing a correlation change amount and an abnormal time parameter value is constructed, a fault position score is dynamically adjusted according to the pointing weight of multiple characteristic vectors, and accurate fault positioning is realized. The method systematically solves the problem of multi-source data processing, significantly improves the accuracy of unit fault early warning and positioning, and effectively guarantees the safe and stable operation of large-scale water turbine sets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of turbine unit fault detection technology, and specifically relates to an early fault location method for the complex data structure of large turbine units. Background Technology

[0002] When large hydroelectric turbine units are in operation, it is necessary to monitor a variety of key operating parameters, such as rotational speed, active power, and various vibration and sway data. These data come from a wide range of sources and are of complex types. The presence of numerous problems with this multi-source data seriously affects the accurate assessment of the unit's operating status and fault analysis.

[0003] On the one hand, the dimensions and magnitudes of data for different parameters vary greatly. For example, the dimensions of rotational speed, excitation current, and top cover vibration are different, making it difficult for traditional analysis methods to effectively compare and analyze these data under a unified standard, resulting in low efficiency in the comprehensive utilization of data.

[0004] On the other hand, data quality varies greatly. During sensor data acquisition, communication failures or sensor malfunctions frequently occur, resulting in missing values. Furthermore, interference from complex operating environments and other factors can also generate outliers. This uncertainty significantly reduces the reliability and usability of the data, easily leading to biased analysis results and failing to accurately reflect the true operating status of the unit.

[0005] Furthermore, existing technologies have significant shortcomings in fault monitoring. Traditional methods often use static thresholds to determine whether the unit's operating status is abnormal. However, the operating conditions of hydroelectric units are complex and variable, significantly affected by factors such as season and operating load. Static thresholds are difficult to adapt to these dynamic changes, which can easily lead to false faults. This can result in either issuing a large number of false alarms, interfering with normal operation and maintenance, or missing real potential faults, posing a great risk to the safe and stable operation of the unit.

[0006] Furthermore, in the fault location phase, relying on single parameters or simple parameter combinations for analysis has been insufficient to comprehensively and accurately determine the fault location. Moreover, the lack of a reasonable dynamic weight allocation mechanism in multi-parameter correlation analysis makes fault location ambiguous, making it difficult for maintenance personnel to quickly and accurately troubleshoot faults. This not only delays maintenance opportunities but may also increase maintenance costs due to repeated troubleshooting.

[0007] Therefore, there is an urgent need to develop a technical method that can effectively manage the uncertainty of multi-source data from large-scale hydro-turbine units and improve the accuracy of fault monitoring and location, so as to ensure the safe and efficient operation of hydro-turbine units. Summary of the Invention

[0008] The technical problem to be solved by the present invention is to provide an early fault location method for the complex data structure of large hydro-turbine units, solve the problem of multi-source data processing, significantly improve the accuracy of unit fault early warning and location, and effectively ensure the safe and stable operation of large hydro-turbine units.

[0009] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:

[0010] An early fault location method for complex data structures in large hydro turbine units, comprising the following steps:

[0011] Step 1. Collect key operating parameters of the turbine unit, including speed, power, and vibration;

[0012] Step 2. Divide the key operating parameters into multiple groups according to the data collection location;

[0013] Step 3. Eliminate differences in data units by outlier detection, linear interpolation to fill missing values, and standardization.

[0014] Step 4. Use the sliding window method to dynamically calculate the Pearson correlation coefficient between key operating parameters and monitor short-term correlation changes in real time;

[0015] Step 5. Dynamically adjust the fault detection threshold based on the standard deviation of historical data to adapt to different operating conditions and reduce false alarms;

[0016] Step 6. Quantify the dynamic change magnitude of the correlation between parameters by the change in correlation between adjacent windows Δr;

[0017] Step 7. When the correlation change Δr exceeds the dynamic threshold, an alarm is triggered to identify deviations in the operating status in real time;

[0018] Step 8. Integrate the parameter values ​​and related changes at the time of the anomaly to construct a multi-dimensional fault feature vector;

[0019] Step 9. Based on the pointing weights of multiple feature vectors, dynamically adjust the fault location score to improve positioning accuracy.

[0020] Preferably, step 1 includes the following:

[0021] Collect the following data for the turbine unit under normal operating conditions: speed, active power, head, excitation current, excitation voltage, reactive power, generator terminal current Ia, generator terminal frequency, generator terminal phase voltage Ua, generator terminal current Ib, generator terminal phase voltage Ub, generator terminal current Ic, generator terminal phase voltage Uc, generator terminal line voltage Uab, generator terminal line voltage Ubc, generator terminal line voltage Uca, peak-to-peak value of top cover vibration (X-bandpass peak-to-peak value), peak-to-peak value of top cover vibration (X-high frequency peak-to-peak value), peak-to-peak value of top cover vibration (Y-bandpass peak-to-peak value), peak-to-peak value of top cover vibration (Y-high frequency peak-to-peak value), peak-to-peak value of top cover vibration (Z-bandpass peak-to-peak value), peak-to-peak value of top cover vibration (Z-high frequency peak-to-peak value), water guide X-axis sway, water guide Y-axis sway, and peak-to-peak value of upper frame vibration (X-axis peak-to-peak value). The system includes 52 key operating parameters, such as the peak-to-peak value of the X-axis of upper frame vibration, the peak-to-peak value of the X-axis of upper frame vibration, the peak-to-peak value of the Y-axis of upper frame vibration, the peak-to-peak value of the Y-axis of upper frame vibration, the peak-to-peak value of the Y-axis of upper frame vibration, the peak-to-peak value of the Y-axis of upper frame vibration, the peak-to-peak value of the Z-axis of upper frame vibration, the peak-to-peak value of the Z-axis of upper frame vibration, the peak-to-peak value of the Z-axis of upper frame vibration, the peak-to-peak value of the Z-axis of lower frame vibration, the peak-to-peak value of the X ... Y-axis of lower frame vibration, the peak-to-peak value of the Y-axis of lower frame vibration, the peak-to-peak value of the Y-axis of lower frame vibration, the peak-to-peak value of the Y-axis of lower frame vibration, the peak-to-peak value of the Y-axis of lower frame vibration, the peak-to-peak value of the Z-axis of lower frame vibration, the peak-to-peak value of the Z-axis of lower frame vibration, the peak-to-peak value of the Z-axis of lower frame vibration, the peak-to-peak value of the Z-axis of lower frame vibration, the peak-to-peak value of the Z-axis of lower frame vibration, the peak-to-peak

[0022] Preferably, step 2 includes the following:

[0023] Based on the different collection locations, the key operational parameters were categorized and processed into six datasets, including:

[0024] Group 1: Rotation speed, active power, head, excitation current, excitation voltage, and unit reactive power;

[0025] Group 2: Generator terminal current Ia, generator terminal frequency, generator terminal phase voltage Ua, generator terminal current Ib, generator terminal phase voltage Ub, generator terminal current Ic, generator terminal phase voltage Uc, generator terminal line voltage Uab, generator terminal line voltage Ubc, generator terminal line voltage Uca;

[0026] Group 3: Top cover vibration X peak-to-peak value, top cover vibration X bandpass peak-to-peak value, top cover vibration X high frequency peak-to-peak value, top cover vibration Y peak-to-peak value, top cover vibration Y bandpass peak-to-peak value, top cover vibration Y high frequency peak-to-peak value, top cover vibration Z peak-to-peak value, top cover vibration Z bandpass peak-to-peak value, top cover vibration Z high frequency peak-to-peak value, water guide X-axis swing, water guide Y-axis swing;

[0027] Group 4: Upper rack vibration X peak-to-peak value, upper rack vibration X bandpass peak-to-peak value, upper rack vibration X high frequency peak-to-peak value, upper rack vibration Y peak-to-peak value, upper rack vibration Y bandpass peak-to-peak value, upper rack vibration Y high frequency peak-to-peak value, upper rack vibration Z peak-to-peak value, upper rack vibration Z bandpass peak-to-peak value, upper rack vibration Z high frequency peak-to-peak value;

[0028] Group 5: Lower frame vibration X peak-to-peak value, lower frame vibration X bandpass peak-to-peak value, lower frame vibration X high frequency peak-to-peak value, lower frame vibration Y peak-to-peak value, lower frame vibration Y bandpass peak-to-peak value, lower frame vibration Y high frequency peak-to-peak value, lower frame vibration Z peak-to-peak value, lower frame vibration Z bandpass peak-to-peak value, lower frame vibration Z high frequency peak-to-peak value;

[0029] Group 6: Upper guide X-axis swing, upper guide Y-axis swing, peak-to-peak value of axial displacement, peak-to-peak value of axial displacement bandpass, peak-to-peak value of axial displacement high frequency, lower guide X-axis swing, lower guide Y-axis swing.

[0030] Preferably, the sub-step of step 3 is as follows:

[0031] Step 3-1: The standard anomaly detection method was used to calculate the mean of the entire dataset. and standard deviation The mean is calculated by summing all the data points and dividing by the total number of data points N:

[0032] ;

[0033] ;

[0034] in, Let N be the i-th data point, and N be the total number of data points.

[0035] By dividing the difference between each data point and the mean by the standard deviation, we ensured that the data were evaluated under the same standard; finally, by setting a threshold, we determined the standardized value. Does it exceed the threshold? If it does, the data point is identified as a global outlier.

[0036] Step 3-2: Fill in the missing values ​​in the complex structure data of the turbine unit:

[0037] A) Locating the location of missing values: First, in the data sequence In this process, the entire data sequence is scanned to determine the locations of missing data, that is, to find all points in the data sequence where all values ​​are missing, denoted as . This indicates that the data at that location is missing; all missing value locations are aggregated into a set. , where i is the index of the missing value;

[0038] B) Interpolation Imputation: Missing data is imputed using linear interpolation. Linear interpolation estimates missing values ​​based on the known linear relationships between adjacent data points, ensuring a smooth transition of data over time. For missing values ​​in the data series... The interpolation and filling steps are as follows:

[0039] ① Determine the interpolation boundaries: for each missing value location First, find the nearest previous non-null value at that position. and the next non-null value ,in ,and , ;

[0040] ② Linear interpolation calculation: missing values The calculation is based on the linear relationship between two consecutive non-missing values ​​and is obtained through the following formula;

[0041] ;

[0042] in, The known value is for position a; Let be the known value at position b; a and b are the indices of the nearest non-missing points before and after it; i is the index of the current missing value.

[0043] ③ Handling consecutive missing values: When consecutive missing values ​​exist... At that time, among them Similarly, imputation is performed using linear interpolation; for the k-th missing value in a series of consecutive missing values, the following condition is met: The calculation formula is as follows:

[0044] ;

[0045] Ensure that the data filling within the entire missing interval conforms to the linear trend between the known data before and after it, ensuring a smooth transition and consistency of the data;

[0046] Step 3-3: Normalization of complex structure data of the turbine unit. By applying a standardization formula, the original value of each parameter is converted into a standardized value Z. This value is calculated by subtracting the mean from the original data and then dividing by the standard deviation. The formula for calculating the standardized value Z is:

[0047] ;

[0048] Where X is the original data, μ is the mean, and σ is the standard deviation.

[0049] Preferably, step 4 employs a sliding window method to calculate the Pearson correlation coefficient: by dynamically moving the window in a continuous data stream, updating the dataset within the window for each new data point, the correlation coefficient r between the two parameters within the window is calculated. xy The size of the sliding window is set according to actual monitoring needs and data characteristics to balance computational efficiency and monitoring sensitivity. At each window position, the dataset within the window is analyzed using the Pearson correlation coefficient formula, where the correlation coefficient r... xy Calculate using the following formula;

[0050] ;

[0051] Where n is the number of data points in the window, and x and y are the datasets of the two parameters.

[0052] Preferably, step 5 includes:

[0053] A dynamic threshold for fault location is set. This dynamic threshold is used to monitor and determine abnormal conditions in the operation of the turbine unit. The dynamic threshold is dynamically adjusted based on key parameters, including historical operating data, seasonal variations, and operating load, to reflect the boundary between normal and abnormal states. When the monitored correlation change, i.e., the correlation between turbine unit operating parameters, exceeds the dynamic threshold for fault location, an abnormality is considered to have occurred in the turbine unit. The dynamic threshold T is based on the average of the correlation changes of historical operating parameter data. and standard deviation The calculation yielded the result.

[0054] ;

[0055] in, and These are the mean and standard deviation of the historical correlation changes, respectively. It is an adjustment factor.

[0056] Preferably, in step 6, the change in correlation Δr within two consecutive sliding windows is calculated: the sliding window method is applied to the continuous data stream, the parameter correlation coefficient within each window is calculated respectively, and then the change in correlation is determined by comparing the correlation coefficient values ​​of two adjacent windows; the formula for calculating the change in correlation Δr within two consecutive sliding windows is as follows;

[0057] ;

[0058] in, r xy t1 and r xy t2These are the correlation coefficients within two adjacent sliding windows.

[0059] Preferably, the sub-steps of step 7 are as follows:

[0060] Step 7.1: Construct a baseline model based on the historical operating data of the turbine unit to reflect the normal state under different operating conditions. This model covers multiple key operating parameters.

[0061] Step 7.2: Calculate the standard thresholds for the correlation between these parameters dynamically using statistical analysis methods. These standard thresholds are adjusted in real time according to the real-time operating conditions to match the current operating environment. The anomaly monitoring process continuously monitors the rotational speed, swing, vibration parameters and their correlations, and compares the current data with the dynamic thresholds in real time to identify changes that deviate from the normal pattern. When the change in the correlation of parameters exceeds the dynamic threshold, the system considers this a potential anomaly and automatically triggers an alarm.

[0062] Step 7.3: Real-time monitoring and accurate anomaly detection of the turbine unit's operating status. The anomaly monitoring strategy is as follows;

[0063] ;

[0064] Where Δr is the change in correlation and T is the dynamic threshold.

[0065] Preferably, in step 8, a fault feature vector F is constructed, which includes the following:

[0066] (1) The correlation change of each operating parameter when the abnormality occurs, that is, the correlation change Δr calculated above within two consecutive sliding windows;

[0067] (2) Specific parameter values ​​at the time of the abnormality, including speed, power and vibration level. Accordingly, fault characteristic Zx is the speed parameter value, Zy is the power parameter value and Zz is the vibration level parameter value.

[0068] The formula for constructing the fault feature vector F is shown below;

[0069] ;

[0070] Where Δr is the correlation change, and Zx, Zy, and Zz are the corresponding fault characteristics.

[0071] In step 9, the fault location of the turbine unit is determined by analyzing the indicated parameter changes and correlation changes. If, during this process, multiple fault feature vectors are found to point to the same part of the turbine unit, the fault probability of that part will be re-evaluated, increasing the fault probability score S at that location. Each element in the fault feature vector (including the correlation change Δr, and the specific parameter values ​​Zx, Zy, and Zz) is used as the basis for judgment. The fault probability score is calculated based on the pre-set weight wi corresponding to each fault feature vector (this weight can be determined based on the importance of different fault features and statistical analysis of historical fault data) and the number n of feature vectors pointing to the same fault location.

[0072] The formula for scoring the probability of failure is as follows;

[0073] ;

[0074] Among them, w i is the weight corresponding to the i-th fault feature vector, and n is the number of feature vectors pointing to the same fault location.

[0075] The present invention can achieve the following beneficial effects:

[0076] 1. Through outlier detection, missing value imputation, and normalization, the system effectively solves the problems of dimensional differences, missing data, and anomalies in multi-source data, improves data quality and consistency, provides a reliable foundation for subsequent analysis, ensures that data can be accurately analyzed under the same standard, and greatly improves the efficiency of comprehensive data utilization.

[0077] 2. By using a sliding window Pearson correlation coefficient calculation combined with dynamic threshold setting, it can track the dynamic changes in the correlation between parameters in real time, adapt to the complex and ever-changing operating conditions of the turbine unit, overcome the limitations of traditional static thresholds, greatly reduce the false alarm and false alarm rates, accurately identify abnormal operating conditions of the unit, and buy valuable time for timely countermeasures.

[0078] 3. The constructed fault feature vector encompasses multi-dimensional information, integrating parameter values ​​and correlation changes at abnormal times. Combined with a fault location scoring mechanism based on the weights of multiple feature vectors, this significantly improves the accuracy and timeliness of fault location. The maintenance team can quickly and effectively troubleshoot faults, shortening maintenance time, reducing maintenance costs, ensuring stable unit operation, and minimizing economic losses caused by downtime due to faults. Attached Figure Description

[0079] The present invention will be further described below with reference to the accompanying drawings and embodiments:

[0080] Figure 1 This is a flowchart of the present invention;

[0081] Figure 2 Heatmap of Pearson correlation coefficient under healthy operating conditions;

[0082] Figure 3 A heatmap of Pearson correlation coefficients under fault operation conditions;

[0083] Figure 4 A comparison of the active power influence coefficient and its rate of change between the healthy state and the fault state of the unit;

[0084] Figure 5 A comparison of the head influence coefficient and its rate of change under healthy and fault conditions;

[0085] Figure 6 A comparison of the speed influence coefficient and its rate of change under healthy and fault conditions;

[0086] Figure 7 A summary chart of parameters that have changed significantly before and after the fault;

[0087] Figure 8 A schematic diagram for fault location based on abnormal parameter identification.

[0088] In the diagram: A: Water conductance swing measurement point;

[0089] B: Vibration measuring point of the top cover;

[0090] C: Vibration measuring point on the lower frame;

[0091] D: Vibration measuring point on the upper frame. Detailed Implementation

[0092] Preferred solutions include Figures 1 to 8 As shown, an early fault location method for complex data structures in large hydroelectric units is described, with the following specific implementation steps:

[0093] A large hydroelectric turbine unit is equipped with multiple sensors to monitor 52 parameters, including rotational speed, active power, and head. This invention is used as an example to illustrate its specific implementation method.

[0094] Step 1: Collect complex data structure of the turbine unit. First, collect the following data under normal operating conditions: turbine speed, active power, head, excitation current, excitation voltage, unit reactive power, generator terminal current Ia, generator terminal frequency, generator terminal phase voltage Ua, generator terminal current Ib, generator terminal phase voltage Ub, generator terminal current Ic, generator terminal phase voltage Uc, generator terminal line voltage Uab, generator terminal line voltage Ubc, generator terminal line voltage Uca, top cover vibration X peak-to-peak value, top cover vibration X bandpass peak-to-peak value, top cover vibration X high-frequency peak-to-peak value, top cover vibration Y peak-to-peak value, top cover vibration Y bandpass peak-to-peak value, top cover vibration Y high-frequency peak-to-peak value, top cover vibration Z peak-to-peak value, top cover vibration Z bandpass peak-to-peak value, top cover vibration Z high-frequency peak-to-peak value, water guide X-axis swing, water guide Y-axis swing, and upper frame vibration. The data includes 52 key operating parameters such as X-peak value, X-bandpass peak value of upper frame vibration, X-high frequency peak value of upper frame vibration, Y-peak value of upper frame vibration, Y-bandpass peak value of upper frame vibration, Y-high frequency peak value of upper frame vibration, Z-peak value of upper frame vibration, Z-bandpass peak value of upper frame vibration, Z-high frequency peak value of upper frame vibration, X-peak value of lower frame vibration, X-bandpass peak value of lower frame vibration, X-high frequency peak value of lower frame vibration, Y-peak value of lower frame vibration, Y-bandpass peak value of lower frame vibration, Y-high frequency peak value of lower frame vibration, Z-peak value of lower frame vibration, Z-bandpass peak value of lower frame vibration, Z-high frequency peak value of lower frame vibration, X-axis swing of upper guide, Y-axis swing of upper guide, peak value of axial displacement, bandpass peak value of axial displacement, high frequency peak value of axial displacement, X-axis swing of lower guide, and Y-axis swing of lower guide. These data will serve as the basis for subsequent analysis.

[0095] [1] Step 2: Classification of complex data structure of turbine unit: Classify the data in Step 1 according to the different collection locations. Specifically: Group 1: Speed, active power, head, excitation current, excitation voltage, unit reactive power; Group 2: Generator terminal current Ia, generator terminal frequency, generator terminal phase voltage Ua, generator terminal current Ib, generator terminal phase voltage Ub, generator terminal current Ic, generator terminal phase voltage Uc, generator terminal line voltage Uab, generator terminal line voltage Ubc, generator terminal line voltage Uca; Group 3: Top cover vibration X peak-to-peak value, top cover vibration X bandpass peak-to-peak value, top cover vibration X high frequency peak-to-peak value, top cover vibration Y peak-to-peak value, top cover vibration Y bandpass peak-to-peak value, top cover vibration Y high frequency peak-to-peak value, top cover vibration Z peak-to-peak value, top cover vibration Z bandpass peak-to-peak value, top cover vibration Z high frequency peak-to-peak value, water guide X-axis swing, water guide Y-axis swing; Group 4: Upper frame vibration The fifth group consists of: X-peak value of upper frame vibration, X-band-pass peak value of upper frame vibration, X-high frequency peak value of upper frame vibration, Y-peak value of upper frame vibration, Y-band-pass peak value of upper frame vibration, Y-high frequency peak value of upper frame vibration, Z-peak value of upper frame vibration, Z-band-pass peak value of upper frame vibration, and Z-high frequency peak value of upper frame vibration; the sixth group consists of: X-axis swing of upper guide, Y-axis swing of upper guide, peak value of axial displacement, peak value of axial displacement with band-pass, peak value of axial displacement with high frequency, X-axis swing of lower guide, and Y-axis swing of lower guide.

[0096] Step 3: Data preprocessing for the complex structure of the turbine unit. The mathematical formulas mentioned above are used to standardize the data from each sensor, ensuring the data are on the same order of magnitude for easier subsequent analysis. This invention employs a standard difference anomaly detection method, first averaging the entire dataset. and standard deviation calculate,

[0097] ;

[0098] ;

[0099] Linear interpolation calculation: missing values The calculation is based on the linear relationship between two consecutive non-missing values, and is obtained through the following formula.

[0100] ;

[0101] in, The known value is for position a; Let a be the known value at position b; a and b be the indices of the nearest non-missing points before and after it; and i be the index of the current missing value.

[0102] The standardized formula is as follows.

[0103] ;

[0104] Where X is the original data, μ is the mean, and σ is the standard deviation.

[0105] (2) Sliding window calculation and correlation analysis

[0106] Step 4: Sliding Window Calculation. When performing sliding window calculation and correlation analysis, a suitable sliding window size needs to be set to capture the dynamic changes in the data. In this embodiment, we choose to cover 10 seconds of data per window. This time length is considered sufficient to capture rapidly changing fault signals while also containing enough data points for reliable statistical analysis. The sliding window moves forward at fixed time intervals (every 10 seconds), ensuring data continuity and coverage while reducing information omissions and duplication. Choosing an appropriate window size and sliding step is crucial for improving the accuracy and efficiency of fault detection and needs to be adjusted and optimized based on actual operating data and fault characteristics. Within each sliding window, the correlation between different sensor data can be effectively assessed by calculating the Pearson correlation coefficient. The Pearson correlation coefficient is a statistic that measures the degree of linear correlation between two variables. Due to the complex structure of the turbine unit, fault location only requires focusing on the impact of the parameters themselves on the system; therefore, the correlation coefficient is processed as an absolute value. The correlation calculation results under healthy conditions are as follows: Figure 2 As shown, the correlation results under fault conditions are as follows: Figure 3 As shown, 1 indicates perfect correlation, and 0 indicates no linear correlation. By calculating the correlation coefficient, we can identify which sensors have significantly correlated data trends, which is crucial for understanding the unit's operating status and detecting early faults. If the data from a certain sensor changes suddenly and its correlation with other sensor data also changes significantly, this may be an early signal of a fault. Therefore, correlation analysis not only helps monitor the real-time status of the unit but also provides strong data support for fault early warning. The formula for calculating the Pearson correlation coefficient is as follows.

[0107] ;

[0108] Where n is the number of data points in the window, and x and y are the datasets of the two parameters.

[0109] Step 5: Dynamic Threshold Setting. To accurately identify abnormal states and trigger the fault detection process, it is first necessary to analyze the normal variation range of the Pearson correlation coefficient between sensor data based on a large amount of historical operating data. By statistically analyzing historical data, the average value μ and standard deviation σ of the correlation coefficient between each pair of sensors can be determined. Based on these statistical characteristics, the threshold for anomaly detection can be dynamically set. The method is to define the threshold using the average value plus or minus a certain number of times the standard deviation, i.e., setting the threshold range to [μ-kσ, μ+kσ], where k is a coefficient determined based on actual needs and experience to control the sensitivity of anomaly detection. The advantage of this method is that it can adaptively adjust the threshold according to the changing characteristics of historical data, thereby improving the accuracy and reliability of anomaly detection.

[0110] During real-time monitoring, the system compares the new Pearson correlation coefficient calculated for each sliding window with a dynamically set threshold. If a correlation coefficient exceeds the preset dynamic threshold range, the condition is met.

[0111] Corr(X,Y)<μ-kσ or Corr(X,Y)>μ+kσ;

[0112] Here, Corr(X,Y) represents the Pearson correlation coefficient between the two sensor data.

[0113] The system then marks this state as abnormal. Once an anomaly is detected, a fault detection process is immediately triggered, including but not limited to further data analysis, fault diagnosis, and notification of maintenance personnel for inspection. This process is not only based on statistical principles but also takes into account the flexibility of actual operation and the needs of emergency response, ensuring the efficient and stable operation of the turbine unit. According to calculations, for the algorithm's real-time data, since this example is in a relatively static state, the data analysis uses a setpoint parameter with a change rate exceeding 40% as the anomaly point. The relevant threshold lines are as follows: Figure 4-6 As shown.

[0114] Step 6: Calculate the correlation change of the turbine unit operation. Calculate the correlation change Δr within two consecutive sliding windows from the signals that may indicate changes in operating status or potential faults captured in real time.

[0115] ;

[0116] in, r xy t1 and r xy t2 These are the correlation coefficients within two adjacent sliding windows, respectively. Based on the example, the calculated rate of change and amount of change before and after the two windows are as follows: Figure 4-6As shown in the bar chart.

[0117] Step 7: Detect abnormal points of turbine unit faults, such as... Figure 4-6 Outliers are circled and marked. Parameters exceeding the threshold set in step 5 are also marked, and so on. Figure 7 As shown, statistics are compiled on parameters that have a significant impact on the system, such as head, active power of the generator unit, and rotational speed.

[0118] Step 8: Constructing the turbine unit fault feature vector. For the abnormal states detected by the system, effective fault analysis requires constructing a feature vector containing comprehensive information. This feature vector includes not only the magnitude of the correlation changes but also key information such as the speed and direction of the changes. Specifically, the feature vector can be represented as:

[0119] T1=[Δr1,Z x1 Z y1 Z z1 ];

[0120] Where Δr1 is the change in correlation, calculated as the difference between the current correlation coefficient and the historical normal value; Z x1 Z y1 Z z1 These are the specific values ​​of speed, power, and vibration under abnormal conditions. A comprehensive consideration of these characteristics provides strong data support for accurately identifying fault modes. The correlation change can be constructed using the following formula.

[0121] ;

[0122] Where v1 represents the magnitude of the correlation change, and ΔCorr(X,Y) represents the amount of correlation change between two adjacent sliding windows.

[0123] Step 9: Turbine Unit Fault Location. Using correlation analysis and known turbine unit structure and sensor layout information, the specific location of the fault is determined. In large turbine units, sensors are typically carefully placed in critical areas to collect various important parameters in real time, such as head, active power, and speed—parameters that significantly impact unit operation. When an anomaly occurs, correlation analysis, combined with known turbine unit structure and sensor layout information, can be used to pinpoint the exact location of the fault. Based on... Figure 8 As shown, by marking the abnormal parameters, we can detect abnormal swing parameters near the water guide bearing. Therefore, we can infer that the fault may be near the water guide bearing. Thus, by comparing and marking these abnormal parameters with the monitoring chart, we can accurately locate the fault point.

[0124] Feedback mechanism: The results of fault detection and location are fed back to the operation and maintenance personnel, and this information is used to optimize the parameter settings of the algorithm. By analyzing the false alarms and false alarms in the fault detection process, the dynamic threshold setting parameters in the algorithm (such as adjusting the k value) can be adjusted, as well as the feature vector construction method and the selection of machine learning model can be optimized to improve the detection accuracy in the future.

[0125] The above embodiments are merely preferred technical solutions of the present invention and should not be considered as limitations on the present invention. The scope of protection of the present invention should be defined as the technical solutions described in the claims, including equivalent substitutions of the technical features described in the claims. That is, equivalent substitutions and improvements within this scope are also within the scope of protection of the present invention.

Claims

1. A method for early fault location oriented to complex data structure of large hydro-turbine unit, characterized in that Comprising the following steps: Step 1. Collect the key operating parameters of the hydraulic turbine unit, including the rotating speed, power and vibration; Step 2. Divide the key operating parameters into multiple groups according to the collection location; Step 3. Eliminate the dimensional difference of the data through outlier detection, linear interpolation to fill in missing values and standardization processing; Step 4. Adopt the sliding window method to dynamically calculate the Pearson correlation coefficient between the key operating parameters, and real-time monitor the short-term correlation change; Step 5. Dynamically adjust the fault detection threshold based on the historical data standard deviation, adapt to different operating conditions to reduce false positives; Step 6. Quantify the dynamic change range of the correlation between parameters through the correlation change quantity Ar of adjacent windows; Step 7. Trigger an alarm when the correlation change quantity Ar exceeds the dynamic threshold, and real-time identify the operating state deviation; Step 8. Integrate the parameter values and correlation change quantity at the abnormal moment, and construct a multi-dimensional fault feature vector; Step 9. Position the fault of the hydraulic turbine unit, analyze according to the indicated parameter change and correlation change quantity; if multiple fault feature vectors are found to point to the same part of the hydraulic turbine unit in this process, the fault possibility of the part will be re-scored, and the fault possibility score S of the position will be increased; Wherein, the fault possibility score is calculated based on each element in the fault feature vector, and the number n of feature vectors pointing to the same fault position; The fault possibility score formula is as follows; ; where w i is the weight corresponding to the ith fault feature vector, and n is the number of feature vectors pointing to the same fault location.

2. The method of claim 1, wherein the method is applied to a large hydroelectric generating unit. The content of step 1 is as follows: Collect the rotating speed, active power, water head, excitation current, excitation voltage, unit reactive power, generator terminal current Ia, generator terminal frequency, generator terminal phase voltage Ua, generator terminal current Ib, generator terminal phase voltage Ub, generator terminal current Ic, generator terminal phase voltage Uc, generator terminal line voltage Uab, generator terminal line voltage Ubc, generator terminal line voltage Uca, top cover vibration X peak-to-peak value, top cover vibration X band pass peak-to-peak value, top cover vibration X high frequency peak-to-peak value, top cover vibration Y peak-to-peak value, top cover vibration Y band pass peak-to-peak value, top cover vibration Y high frequency peak-to-peak value, top cover vibration Z peak-to-peak value, top cover vibration Z band pass peak-to-peak value, top cover vibration Z high frequency peak-to-peak value, water guide X swing, water guide Y swing, upper bracket vibration X peak-to-peak value, upper bracket vibration X band pass peak-to-peak value, upper bracket vibration X high frequency peak-to-peak value, upper bracket vibration Y peak-to-peak value, upper bracket vibration Y band pass peak-to-peak value, upper bracket vibration Y high frequency peak-to-peak value, upper bracket vibration Z peak-to-peak value, upper bracket vibration Z band pass peak-to-peak value, upper bracket vibration Z high frequency peak-to-peak value, lower bracket vibration X peak-to-peak value, lower bracket vibration X band pass peak-to-peak value, lower bracket vibration X high frequency peak-to-peak value, lower bracket vibration Y peak-to-peak value, lower bracket vibration Y band pass peak-to-peak value, lower bracket vibration Y high frequency peak-to-peak value, lower bracket vibration Z peak-to-peak value, lower bracket vibration Z band pass peak-to-peak value, lower bracket vibration Z high frequency peak-to-peak value, upper guide X swing, upper guide Y swing, axial displacement peak-to-peak value, axial displacement band pass peak-to-peak value, axial displacement high frequency peak-to-peak value, lower guide X swing and lower guide Y swing, etc. 52 kinds of key operating parameters of the hydraulic turbine unit in normal operating state.

3. The method of claim 2, wherein the method is applied to a large hydroelectric generating unit. The content of step 2 is as follows: According to the different collection positions, the key operating parameters are classified and processed into six data sets; including: The first group: speed, active power, water head, excitation current, excitation voltage, and unit reactive power; The second group: generator terminal current Ia, generator terminal frequency, generator terminal phase voltage Ua, generator terminal current Ib, generator terminal phase voltage Ub, generator terminal current Ic, generator terminal phase voltage Uc, generator terminal line voltage Uab, generator terminal line voltage Ubc, and generator terminal line voltage Uca; The third group: top cover vibration X peak-to-peak value, top cover vibration X band-pass peak-to-peak value, top cover vibration X high-frequency peak-to-peak value, top cover vibration Y peak-to-peak value, top cover vibration Y band-pass peak-to-peak value, top cover vibration Y high-frequency peak-to-peak value, top cover vibration Z peak-to-peak value, top cover vibration Z band-pass peak-to-peak value, top cover vibration Z high-frequency peak-to-peak value, water guide X swing, and water guide Y swing; The fourth group: upper bracket vibration X peak-to-peak value, upper bracket vibration X band-pass peak-to-peak value, upper bracket vibration X high-frequency peak-to-peak value, upper bracket vibration Y peak-to-peak value, upper bracket vibration Y band-pass peak-to-peak value, upper bracket vibration Y high-frequency peak-to-peak value, upper bracket vibration Z peak-to-peak value, upper bracket vibration Z band-pass peak-to-peak value, and upper bracket vibration Z high-frequency peak-to-peak value; The fifth group: lower bracket vibration X peak-to-peak value, lower bracket vibration X band-pass peak-to-peak value, lower bracket vibration X high-frequency peak-to-peak value, lower bracket vibration Y peak-to-peak value, lower bracket vibration Y band-pass peak-to-peak value, lower bracket vibration Y high-frequency peak-to-peak value, lower bracket vibration Z peak-to-peak value, lower bracket vibration Z band-pass peak-to-peak value, and lower bracket vibration Z high-frequency peak-to-peak value; The sixth group: upper guide X swing, upper guide Y swing, axial displacement peak-to-peak value, axial displacement band-pass peak-to-peak value, axial displacement high-frequency peak-to-peak value, lower guide X swing, and lower guide Y swing.

4. The method of claim 1, wherein the method is applied to a large hydroelectric generating unit. In step 4, the sliding window method is used to calculate the Pearson correlation coefficient: by dynamically moving the window in the continuous data stream, updating the data set within the window for each new data point, the correlation coefficient r between the two parameters within the window is calculated xy The size of the sliding window is set according to the actual monitoring requirements and the characteristics of the data to balance the calculation efficiency and monitoring sensitivity; at each window position, the data set within the window is analyzed by the Pearson correlation coefficient formula, and the correlation coefficient r xy is calculated by the following formula; ; Wherein, n is the number of data points in the window, and x and y are the data sets of the two parameters, respectively.

5. The method of claim 4, wherein the method is applied to a large hydroelectric generating unit. The content of step 5 is: The fault positioning dynamic threshold is set, and the dynamic threshold is used for monitoring and judging abnormal conditions in the operation state of the hydro-turbine set. The fault positioning dynamic threshold is dynamically adjusted according to the key parameters including historical operation data, seasonal changes and operation load, and is used for reflecting the boundary between the normal operation state and the abnormal state. When the monitored correlation change amount, i.e. the correlation between the operation parameters of the hydro-turbine set, exceeds the fault positioning dynamic threshold, it is considered that the hydro-turbine set may have an abnormality. The dynamic threshold T is calculated based on the mean value and the standard deviation of the correlation change amount of the historical operation parameter data. and the standard deviation The calculation is as follows: ; wherein, and are the average and standard deviation of the historical correlation change amount, respectively, is an adjustment factor.

6. The method of claim 1, wherein: In step 6, the correlation change Δr between the two continuous sliding windows is calculated: the sliding window method is applied in the continuous data stream, the parameter correlation coefficients in each window are calculated respectively, and then the correlation change is determined by comparing the correlation coefficient values of the adjacent two windows; The formula for calculating the correlation change Δr between the two continuous sliding windows is as follows: ; wherein, r xy t1 and r xy t2 are the correlation coefficients in the front and rear two adjacent sliding windows, respectively; t1 and t2 represent the time sequence numbers of the front and rear two adjacent sliding windows, respectively.

7. The method of claim 1, wherein the method is applied to large hydroelectric generating units. The sub-steps of step 7 are as follows: Step 7.1, based on the historical operating data of the hydraulic turbine unit, a baseline model reflecting the normal state under different operating conditions is constructed, which covers multiple key operating parameters; Step 7.2, the standard threshold of the correlation between these parameters is dynamically calculated by statistical analysis method, which is adjusted in real time according to the real-time operating conditions to match the current operating environment; The abnormality monitoring process monitors the speed, swing, vibration parameters and their correlation continuously, compares the current data with the dynamic threshold in real time, so as to identify the changes deviating from the normal mode; when the correlation change of the parameters exceeds the dynamic threshold, the system regards it as a potential abnormality and automatically triggers an alarm; Step 7.3, real-time monitoring of the operating state of the hydraulic turbine unit and accurate abnormality detection, the abnormality monitoring strategy is as follows: ; Wherein, Δr is the correlation change, and T is the dynamic threshold.

8. The method of claim 6, wherein the method is applied to a large hydroelectric generating unit. In step 8, a fault feature vector F is constructed, and the fault feature vector F includes the following contents: (1) the correlation change amount of each operating parameter when the anomaly occurs, that is, the correlation change amount Δr calculated in the above two sliding windows; (2) the specific parameter value at the anomaly time, and the specific parameter value includes the rotating speed, the power, the vibration level, and correspondingly, the fault feature Zx is the rotating speed parameter value, the fault feature Zy is the power parameter value, and the fault feature Zz is the vibration level parameter value; The construction formula of the fault feature vector F is as follows: ; Wherein, Δr is the correlation change amount, and Zx, Zy and Zz are the corresponding fault features.

Citation Information

Patent Citations

  • Monitoring data anomaly detection method based on space-time correlation coefficient

    CN118312902A

  • Power consumption abnormity diagnosis method based on multi-element composite feature collaborative machine learning

    CN118535853A