Hydroelectric generating set state online monitoring system and method based on data analysis

By building a hydropower unit status monitoring system based on data analysis, using variational autoencoders and mixed density network models, combined with a gated attention mechanism, accurate monitoring of the hydropower unit status and fault identification are achieved, solving the problem of insufficient recognition accuracy in existing technologies and improving the system's adaptability and intelligence.

CN120744720APending Publication Date: 2025-10-03HUNAN WULING POWER TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510932802.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

The existing hydropower unit condition monitoring system has insufficient recognition accuracy in dynamic environments, making it difficult to extract subtle characteristic changes in vibration anomalies. It also lacks the ability to automatically model comprehensive features such as load conditions, frequency domain disturbances, and axis trajectory, resulting in slow response and low recognition accuracy.

Method used

A data analysis-based method is adopted to synchronously collect vibration and electrical energy parameters through sensors, perform spectral energy distribution analysis and feature extraction, construct a variational autoencoder model and a hybrid density network model, and combine the gated attention mechanism for feature fusion and anomaly judgment to achieve accurate extraction and automatic classification of different vibration modes.

Benefits of technology

It has improved the recognition accuracy and fault prevention and control capabilities of the hydropower unit status monitoring system, enhanced the system's adaptability and intelligence in complex environments, and reduced fault downtime and false detection and repair rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744720A_ABST
    Figure CN120744720A_ABST
Patent Text Reader

Abstract

The invention discloses a hydroelectric generating set state online monitoring system and method based on data analysis, and relates to the technical field of data analysis. The method comprises the following steps: setting a training period, extracting a time domain feature vector and a frequency domain feature vector of each data packet, and training to generate a gated attention mechanism model; training and generating a mixed density network model according to the data packet containing the fault state label; according to the real-time fault risk probability, the latent variable representation vector and the real-time center frequency, whether abnormal features are analyzed or not is judged, if the abnormal features are analyzed, the contribution degree of each input feature to the real-time fault risk probability is calculated through a gradient approximation method, and an influence degree score is calculated; and taking the corresponding input features exceeding the influence degree score threshold as abnormal features, and outputting prompt information for operation and maintenance personnel to check, thereby improving dynamic environment response, identification precision degree and reliability of a monitoring system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and in particular to a system and method for online monitoring of the status of a hydropower unit based on data analysis. Background Art

[0002] In hydropower equipment condition monitoring systems, the accuracy of vibration signal analysis directly determines the accuracy and speed of mechanical fault identification. As unit structures become increasingly complex and operating conditions become increasingly diverse, traditional identification mechanisms based on fixed-frequency analysis and manual comparison are gradually exposing problems such as low recognition accuracy and weak response capabilities. Current vibration anomaly identification relies primarily on spectrum interpretation using fixed segments and empirical rules. It lacks the ability to automatically model comprehensive features such as load conditions, frequency domain disturbances, and axis trajectory. Furthermore, it is difficult to extract subtle feature changes in the early stages of an anomaly. Some algorithms are not adaptable to dynamically changing conditions, making it impossible to effectively classify anomalies. Therefore, it is necessary to construct an identification mechanism based on vibration feature interval division and frequency domain dynamic matching to support the accurate extraction of different vibration modes and automatic classification of abnormal types, so as to solve the problems of slow response and insufficient recognition accuracy of existing identification methods in dynamic environments, and improve the reliability and fault prevention and control capabilities of the monitoring system. Summary of the Invention

[0003] The purpose of the present invention is to provide a system and method for online monitoring of the status of a hydropower unit based on data analysis, so as to solve the problems raised in the prior art.

[0004] In order to solve the above technical problems, the present invention provides the following technical solutions: a method for online monitoring of the status of a hydropower unit based on data analysis, the method comprising: Step S100: In the hydropower equipment monitoring system, the operating parameters and power parameters of the hydropower unit are collected and packaged into a data packet with a time stamp, the vibration signal in the data packet is obtained, the spectrum energy distribution curve is drawn, and the energy proportion and energy entropy characteristic value are calculated; Step S200: Obtaining the power parameters of the data packet, calculating the operating condition score, determining the operating condition level of each data packet, constructing a spectrum structure reference template library, and calculating the center frequency interval of each frequency band; Step S300: Obtain a vibration signal within a center frequency interval, calculate the envelope kurtosis, spectrum centroid offset value, and ellipticity eigenvalue, train and generate a variational autoencoder model, construct a sensitive eigenvector based on the variational autoencoder model, and generate a discriminant boundary interval for each dimensional latent variable; Step S400: Setting a training cycle, extracting the time domain feature vector and frequency domain feature vector of each data packet, training and generating a gated attention mechanism model, and training and generating a mixture density network model based on the data packet containing the fault status label; Step S500: constructing the operating condition characteristics of the historical data packet, generating a historical operating condition library, obtaining the real-time operating condition characteristics of the real-time data packet, performing similarity matching with the historical operating condition library, collecting the maximum similarity index, calculating the dynamic classification threshold, and determining the matching status of the real-time data packet; Step S600: Based on the matching status, the real-time data packet is divided into frequency bands and template matched, the real-time center frequency within each frequency band is extracted, and a real-time feature input vector is obtained. The mixed weights of several components are calculated through a gated attention mechanism model and a mixture density network model, and the real-time fault risk probability is calculated. The sensitive feature vector of the real-time data packet is input into a variational autoencoder model to obtain a latent variable representation vector. Step S700: Determine whether to analyze abnormal features based on the real-time fault risk probability, latent variable representation vector, and real-time center frequency. If abnormal features are to be analyzed, the gradient approximation method is used to calculate the contribution of each input feature to the real-time fault risk probability, and the impact score is calculated. The input features corresponding to those exceeding the impact score threshold are regarded as abnormal features, and a prompt message is output for verification by the operation and maintenance personnel.

[0005] Furthermore, step S100 includes: Step S101: Several types of sensors are placed at key locations of the hydropower unit to monitor the operating parameters of the unit and simultaneously collect the power parameters output by the unit. The operating parameters and power parameters are synchronized and processed via an edge acquisition terminal, packaged into a data packet with a timestamp, and sampled according to a preset sampling period. The data packet is then uploaded to the hydropower equipment monitoring system via a communication interface. Step S102: Obtain a data packet from a hydropower equipment monitoring system, collect vibration signals of operating parameters in the data packet, decompose the vibration signal into several intrinsic mode functions by using an empirical mode decomposition method, remove redundant modes of high-frequency noise and background disturbances by analyzing spectral characteristics, retain the dominant modal components of the operating state of the hydropower unit, perform sliding window slicing on the obtained dominant modal component signals according to a preset time window, preset an overlap rate threshold, frame the signal sequence that meets the overlap rate threshold, generate a signal slice sequence within a continuous time period, extract spectral characteristics of each frame slice signal by short-time Fourier transform, further calculate the Mel-frequency cepstrum coefficient, construct a frame-level time-frequency feature matrix, and calculate the power spectral density of each frame slice signal by using the Welch method to obtain the energy distribution of the signal in the frequency domain, and draw a spectral energy distribution curve; Step S103: Slice the vibration signal after frame processing and perform multi-scale frequency domain decomposition using the wavelet packet decomposition method. Select the db6 wavelet basis and set the number of decomposition layers. Decompose each frame signal into several equal-width sub-bands. Calculate the energy value of each sub-band signal and perform normalization processing. Calculate the energy proportion according to the following formula: ; Among them, A a Expressed as the energy ratio of the ath sub-band signal, B a Expressed as the energy value of the ath sub-band signal, B b It is expressed as the energy value of the b-th sub-band signal, c is the total number of sub-band signals, and the energy entropy eigenvalue is calculated according to the following formula: ; Among them, D represents the energy entropy eigenvalue, A d Expressed as the energy proportion of the d-th sub-band signal; By deploying multiple types of sensors at key locations on hydropower units, we can simultaneously collect operating parameters such as vibration, temperature, and pressure, as well as electrical energy parameters such as voltage, current, and active power. Edge computing terminals are introduced for data synchronization and packaging, ensuring data consistency, timeliness, and integrity. Vibration signals can be adaptively decomposed into several intrinsic mode functions through empirical mode decomposition. Combined with spectrum analysis, high-frequency noise and background interference are removed. By retaining the dominant mode components, the characteristic signals most closely related to the unit's operating status are effectively highlighted, avoiding misjudgment caused by irrelevant interference. The spectrum information of each frame signal is extracted through short-time Fourier transform, and the Mel-frequency cepstral coefficients are further extracted to achieve time-frequency joint modeling. At the same time, the power spectral density is calculated through the Welch method to obtain the energy distribution in the frequency domain, avoiding the instability of single-point spectral analysis. The vibration signal slice is further decomposed using wavelet packets and the db6 wavelet basis into multiple equal-width sub-bands. The energy value and energy proportion of each sub-band are calculated, and the energy entropy eigenvalue is introduced to quantify the complexity and disorder of the spectrum. Energy entropy, as a nonlinear characteristic indicator, can significantly improve the model's sensitivity and resolution to weak early fault signals.

[0006] Furthermore, step S200 includes: Step S201: Obtain a historical sampling period of the hydropower equipment monitoring system, collect data packets of a certain historical sampling period, extract the power parameters of the data packets, and calculate the operating condition score according to the following formula: ; Among them, E represents the working condition score, P represents the active power, U represents the voltage, and I represents the current; Step S202: Summarize all the working condition scores and calculate the average value of the working condition scores as E1 and the standard deviation as E2. K working condition levels are preset, and the working condition score range of each working condition level is calculated according to the following formula: ; Among them, R r It is expressed as the working condition score range of the rth working condition level, summarizing the data packets of each working condition level; Step S203: In a data packet of a certain operating condition level, extract the spectrum energy distribution curve of each data packet, construct a spectrum data set, preset a number of candidate cluster numbers, use a spectral clustering algorithm to perform cluster analysis on the spectrum data set according to each candidate cluster number, obtain a clustering result corresponding to each candidate cluster number, perform a goodness of fit evaluation on each cluster result using the Akaike Information Criterion, obtain an AIC value corresponding to each cluster result, select the candidate cluster number corresponding to the cluster result with the smallest AIC value as the cluster number of the operating condition level, and divide the spectrum of the operating condition level into multiple frequency band regions according to the cluster number to construct a spectrum structure template for the operating condition level; Step S204: Obtain all spectrum energy distribution curves in each frequency band, and perform clustering and classification based on the spectrum shape distribution rules, assign different types of labels to each frequency band, and generate a spectrum structure reference template library; Step S205: For a certain frequency band in the spectrum structure reference template library, extract the center frequency within the frequency band from the spectrum energy distribution curve of each data packet. The center frequency is the frequency point with the maximum energy value in the spectrum energy distribution curve. The average value of the center frequency is calculated as F1 and the standard deviation is calculated as F2. The center frequency interval of the frequency band is constructed as [F1-h×F2,F1+h×F2], where h represents a preset tolerance adjustment parameter. By introducing parameters such as voltage, current, and active power, the operating condition score is calculated, and a working condition evaluation index system strongly correlated with the operating status is constructed. The mean and standard deviation of the operating condition score are used for grading. This method can divide large-scale historical data into multiple representative operating condition levels. Compared with traditional methods that rely on manual labeling or empirical judgment, this method can automatically generate standardized operating condition grading standards, significantly improving the system's adaptability and objectivity. At each operating level, a data set is constructed using the spectrum energy distribution curve. Combining the spectral clustering algorithm and the Akaike Information Criterion, the optimal number of clusters is automatically searched. The AIC minimization strategy avoids the error of subjectively setting the number of clusters, making the spectrum structure division more reasonable and enhancing the scientific nature of frequency band identification. Compared with fixed-band or sliding-window segmentation methods, this method has the capabilities of dynamic frequency division and adaptive modeling. An independent spectrum structure template is constructed for each operating condition level, which is classified according to the frequency bands divided by clustering. Different labels are assigned to different frequency bands to build a spectrum structure reference template library, providing a structured retrieval basis for subsequent frequency band matching and anomaly positioning.

[0007] Furthermore, step S300 includes: Step S301: Based on the frequency band center frequency interval constructed in step S205, the vibration signal of the data packet in each frequency band is collected, the vibration signal in a certain frequency band is extracted through a bandpass filter, and the extracted vibration signal is subjected to Hilbert transform to obtain an envelope signal. A number of sampling points are preset, and the values ​​of the envelope signal corresponding to each sampling point are summarized. The average value L1 and standard deviation L2 of the envelope signal are calculated, and the envelope kurtosis is calculated according to the following formula: ; Among them, H represents the envelope kurtosis, L i It is represented by the value of the envelope signal corresponding to the i-th sampling point, and j is represented by the total number of sampling points; Step S302: extract the frequency value and power spectrum density value corresponding to each sampling point, and calculate the spectrum centroid offset value according to the following formula: ; Among them, N represents the spectrum centroid offset value, n i Represented as the frequency value of the i-th sampling point, m i It is represented by the power spectrum density value of the i-th sampling point, and F is represented by the mean center frequency of the current frequency band; Step S303: Obtain the axis trajectory data of the data packet. The axis trajectory data is synchronously collected by bidirectional vibration sensors arranged at key positions of the rotating shaft to form time series displacement data on a two-dimensional plane. Based on the time series displacement data, a set of trajectory coordinate points is extracted. The geometric shape of the axis trajectory is fitted using the least squares method. The length of the major semi-axis corresponding to the axis trajectory is calculated as p1 and the length of the minor semi-axis is calculated as p2. The ellipticity characteristic value of the axis trajectory is calculated according to the following formula: ; Where P represents the ellipticity eigenvalue; Step S304: Summarize the envelope kurtosis, spectrum centroid offset value, and ellipticity eigenvalue of each frequency band to construct a multidimensional feature vector of the data packet, input the multidimensional feature vector of the data packet collected under historical normal operating conditions into a variational autoencoder model for training, perform nonlinear dimensionality reduction and sensitivity enhancement processing on the multidimensional features through unsupervised learning, minimize reconstruction error and KL divergence as optimization goals, learn a low-dimensional representation of the feature distribution in the latent space under normal conditions, and output a latent variable representation vector; Step S305: Based on the trained variational autoencoder model, calculate the sensitivity index of each dimensional feature vector to the reconstruction error, construct a feature sensitivity matrix, preset the number of sensitive features to z, sort the sensitivity of all dimensional features to the reconstruction error, select the first z features with the highest contribution as sensitive features, and construct a sensitive feature vector; Step S306: Extract the latent variable representation vector set corresponding to the sensitive features, calculate the mean v1 and standard deviation v2 of the latent variable of each dimension, and construct the judgment boundary of the latent variable of a certain dimension as [vmin, vmax] = [v1-w×v2, v1+w×v2], where w represents the adjustment coefficient of the latent variable; The system uses envelope kurtosis to capture impact-related faults such as bearing spalling; spectrum centroid offset reflects the deviation between the spectrum centroid and the center frequency of the frequency band, identifying frequency drift-related faults; and ellipticity eigenvalue describes the degree to which the axis trajectory deviates from a circle, identifying mechanical structural anomalies such as shaft eccentricity and flexible deformation. These features comprehensively cover multiple potential fault modes, including mechanical structural anomalies, impact vibration, and frequency drift, effectively enhancing the system's multimodal perception capabilities. Based on VAE, high-dimensional original features are compressed into low-dimensional latent variables. At the same time, by minimizing the reconstruction error + KL divergence optimization method, the potential distribution structure of multi-dimensional features under normal conditions can be learned. This provides a mathematical basis for subsequent judgment of whether a certain state deviates from the "normal sample family" and can efficiently identify unseen anomalies or signs of minor faults.

[0008] Furthermore, step S400 includes: Step S401: Select several consecutive days as a training cycle, extract the time domain feature vector and frequency domain feature vector of each data packet, the time domain feature vector includes the envelope kurtosis and the spectrum centroid offset value, and the frequency domain feature vector includes the power spectrum density, energy proportion, and energy entropy feature value, concatenate the time domain feature vector and the frequency domain feature vector to obtain a feature input vector, and form a training set with the state label of the data packet. Input the feature input vector into the gated attention mechanism model, calculate the frequency domain score vector, and perform back propagation training based on the cross entropy loss based on the difference between the frequency domain score vector and the target state label, thereby optimizing the weight matrix in the gated attention mechanism model, and obtaining a trained gated attention mechanism model; Step S402: Obtain a data packet containing a fault status label, perform numerical processing on the fault status, extract and fuse the feature vector processed by the gated attention mechanism model, input the feature input vector into the mixture density network model, output the conditional probability distribution of the target variable, maximize the log-likelihood value distributed at the fault status label as the optimization target, and use the backpropagation algorithm to update the parameters used in the model for predicting the mixture weight, mean, and standard deviation to obtain a trained mixture density network model; The feature input vector integrates time-domain and frequency-domain features, achieving integrated expression of information from different dimensions. This enhances the model's ability to identify complex operating conditions and multiple types of faults, ensuring that the model can fully capture the inherent patterns of the hydropower unit's operating status. The gated attention mechanism automatically learns the attention weights for each feature dimension based on training data. It dynamically highlights the more representative feature dimensions under different conditions, implementing a modeling strategy that lets important features speak for themselves. Compared to traditional static weighting methods, the gated structure has information regulation capabilities, significantly improving the robustness and generalization ability of the model under complex feature inputs. The mixed density network achieves uncertainty modeling of prediction results by outputting the "probability density function of a certain state belonging to a certain category". It can adapt to fuzzy transitions or overlapping areas between state labels, such as the boundary between minor early faults and normal states. Compared with traditional classification models, it can reflect multi-peak distribution and probability overlapping structure, and is closer to the physical essence of the actual operating state of hydropower equipment.

[0009] Furthermore, step S500 includes: Step S501: In the hydropower equipment monitoring system, the operating condition level of the historical data packets is obtained, the historical data packets of each operating condition level are summarized, the average load rate, average speed, and sensitive feature vector of a historical data packet in a certain operating condition level are collected, and the operating condition feature of the historical data packet is constructed as Q=(X, Y, U), where X, Y, and U represent the average load rate, average speed, and sensitive feature vector of the historical data packet, respectively. The operating condition features of all historical data packets are summarized to obtain a sample set of the operating condition level, and a historical operating condition library is generated; Step S502: Acquire real-time data packets collected in real time, extract the average load rate, average rotation speed, and sensitive feature vectors of the real-time data packets, and calculate the similarity index according to the following formula: ; Among them, G represents the similarity index, α represents the weight diagonal matrix of the working condition characteristics, and Q1 represents the real-time working condition characteristics of the real-time data packet; Step S503: Obtain similarity indexes between the real-time data packet and each operating condition feature in the historical operating condition database. The maximum value of the collected similarity indexes is Gmax. The dynamic classification threshold is calculated according to the following formula: ; Wherein, T represents the dynamic classification threshold, T1 represents the basic threshold, and β represents the adjustment coefficient. If Gmax ≥ T, the real-time data packet is determined to be similar to the historical working condition, and the known working condition processing flow is entered. If Gmax < T, step S601 is executed; Extract three dimensional information, namely, average load rate, average speed, and sensitive feature vectors, from historical data to form a ternary working condition feature. This allows us to establish a representative and discriminative historical working condition sample set classified by working condition level, providing a reliable data foundation for subsequent similarity matching and working condition retrospective identification, and enhancing the efficiency of the model's historical knowledge utilization. By introducing a diagonal matrix of working condition feature weights, different dimensions can be assigned different importance weights, achieving quantitative evaluation of the similarity between real-time working conditions and historical working conditions. It supports flexible adjustment of feature weights according to actual scenarios and realizes personalized working condition identification.

[0010] Furthermore, step S600 includes: Step S601: Based on a spectrum structure reference template library, frequency band division and template matching are performed on the real-time data packet, and the real-time center frequency in each frequency band is extracted from the spectrum energy distribution curve of the real-time data packet; Step S602: Obtain the time domain feature vector and the frequency domain feature vector of the real-time data packet, and concatenate them to obtain a real-time feature input vector, input the real-time feature input vector into the gated attention mechanism model, and obtain a weighted fused feature vector; Step S603: Input the feature vector into the mixture density network model to model the fault risk probability distribution, obtain the mixed weights of several components, and calculate the real-time fault risk probability according to the following formula: ; Among them, S represents the real-time failure risk probability, S t It is represented as the mixing weight of the t-th component, and t' is represented as the number of components judged to be in a normal state; Step S604: inputting the sensitive feature vector of the real-time data packet into the trained variational autoencoder model to obtain a latent variable representation vector; By calling the spectrum structure reference template library, the spectrum signal of the real-time data packet is divided into multiple frequency bands and the center frequency within each band is extracted. This allows for rapid determination of whether the current spectrum morphology deviates from the historical template, enabling structural perception of phenomena such as frequency domain anomalies, operating condition drift, and harmonic distortion. The time-domain feature vector and frequency-domain feature vector of the real-time data packet are concatenated to form a complete feature input. After inputting into the gated attention mechanism model, the model automatically "focuses" on key feature dimensions and outputs a weighted fusion discriminative feature vector. This effectively enhances the model's ability to mine interactions between multi-source features and improves its sensitivity and accuracy to complex signal states. The fused feature vector is input into a mixture density network. The model outputs multiple components and their mixed weights to calculate the real-time fault risk probability. This not only indicates whether it is abnormal but also reflects the magnitude of the risk. Compared with traditional classifiers, this model can characterize multimodal probability structures and is more suitable for handling complex scenarios such as state ambiguity and transition zone conditions. Sensitive features are input into a pre-trained variational autoencoder to obtain its latent variable representation vector. This vector can be used to detect whether the input falls into a known distribution area, reflecting "whether the state belongs to the normal behavior space that the system has mastered", further enhancing the ability to capture new faults or unknown anomalies, and providing strong support for anomaly identification and model generalization.

[0011] Furthermore, step S700 includes: Step S701: If a fault risk probability threshold is preset, and if the real-time fault risk probability is greater than or equal to the fault risk probability threshold, or if any dimension of the latent variable representation vector is not within the discrimination boundary interval constructed in step S306, or if the real-time center frequency is not within the center frequency interval constructed in step S205, then step S702 is executed; Step S702: Calculate the contribution of each input feature to the real-time fault risk probability using a gradient approximation method, and multiply the contribution of each input feature by the impact factor of the input feature to obtain an impact score; Step S703: Preset an impact score threshold, treat the input features corresponding to those exceeding the impact score threshold as abnormal features, and output a prompt message for operation and maintenance personnel to check; At the same time, we examine: whether the real-time fault risk probability exceeds the threshold; whether the latent variable crosses the boundary; whether the center frequency deviates from the template frequency interval; Fault determination can be triggered if any condition is met, reducing the risk of misjudgment of a single indicator, achieving more comprehensive anomaly identification, adapting to various fault modes and operating condition changes, and enhancing the model's adaptability to complex operating conditions. The gradient approximation method is used to calculate the marginal impact of each input feature on the risk probability, reflecting its "risk driving force". The contribution is multiplied by the preset feature impact factor to obtain the impact degree score. This supports a clear explanation of "which features dominate the risk" and makes the risk assessment process traceable, understandable, and verifiable. By setting an impact score threshold, only features with truly high impact and high risk relevance are output, and prompt information containing abnormal features is output, thus realizing an intelligent response closed loop of "identification, positioning, and notification". Operation and maintenance personnel can conduct rapid investigation and maintenance based on this, significantly improving diagnostic efficiency and response speed, effectively reducing false alarm and repair rates, and improving the intelligence level of the system.

[0012] In order to better implement the above method, a hydropower unit status online monitoring system based on data analysis is proposed. The system includes a data acquisition and preprocessing module, a spectrum template module, a variational autoencoder model module, a training cycle module, a real-time operating condition matching module, a real-time monitoring module, and an abnormality judgment module. Data acquisition and preprocessing module: In the hydropower equipment monitoring system, it collects the operating parameters and power parameters of the hydropower unit, packages them into data packets with timestamps, obtains the vibration signals in the data packets, draws the spectrum energy distribution curve, and calculates the energy proportion and energy entropy characteristic value; Spectrum template module: obtains the power parameters of the data packet, calculates the operating condition score, determines the operating condition level of each data packet, builds a spectrum structure reference template library, and calculates the center frequency range of each frequency band; Variational autoencoder model module: obtains the vibration signal within the center frequency range, calculates the envelope kurtosis, spectrum centroid offset value, and ellipticity eigenvalue, trains and generates a variational autoencoder model, constructs sensitive feature vectors based on the variational autoencoder model, and generates the discriminant boundary interval of each dimensional latent variable; Training cycle module: Sets the training cycle, extracts the time domain feature vector and frequency domain feature vector of each data packet, trains and generates a gated attention mechanism model, and trains and generates a mixture density network model based on data packets containing fault status labels; Real-time working condition matching module: constructs the working condition characteristics of historical data packets, generates a historical working condition library, obtains the real-time working condition characteristics of real-time data packets, performs similarity matching with the historical working condition library, collects the maximum value of the similarity index, calculates the dynamic classification threshold, and determines the matching status of the real-time data packets; Real-time monitoring module: Based on the matching status, the real-time data packet is divided into frequency bands and template matched, the real-time center frequency within each frequency band is extracted, and a real-time feature input vector is obtained. The mixed weights of several components are calculated through a gated attention mechanism model and a mixture density network model, and the real-time fault risk probability is calculated. The sensitive feature vector of the real-time data packet is input into the variational autoencoder model to obtain a latent variable representation vector. Abnormal judgment module: Based on the real-time fault risk probability, latent variable representation vector, and real-time center frequency, it determines whether to analyze abnormal features. If abnormal features are to be analyzed, the gradient approximation method is used to calculate the contribution of each input feature to the real-time fault risk probability, and the impact score is calculated. The input features corresponding to those exceeding the impact score threshold are regarded as abnormal features, and prompt information is output for verification by operation and maintenance personnel.

[0013] Furthermore, step S700 includes: the training cycle module includes a gated attention mechanism model unit and a mixture density network model unit: Gated attention mechanism model unit: Select several consecutive days as the training cycle, extract the time domain feature vector and frequency domain feature vector of each data packet, the time domain feature vector includes the envelope kurtosis and the spectrum centroid offset value, and the frequency domain feature vector includes the power spectrum density, energy proportion, and energy entropy feature value. The time domain feature vector and the frequency domain feature vector are spliced ​​to obtain a feature input vector, which is used together with the status label of the data packet to form a training set. The feature input vector is input into the gated attention mechanism model, and a frequency domain score vector is calculated. The difference between the frequency domain score vector and the target status label is used to perform back propagation training based on the cross-entropy loss, thereby optimizing the weight matrix in the gated attention mechanism model, and obtaining a trained gated attention mechanism model. Mixture density network model unit: obtains a data packet containing a fault status label, digitizes the fault status, extracts and fuses the feature vector processed by the gated attention mechanism model, inputs the feature input vector into the mixture density network model, outputs the conditional probability distribution of the target variable, maximizes the log-likelihood value distributed at the fault status label as the optimization target, and uses the backpropagation algorithm to update the parameters in the model used to predict the mixture weight, mean, and standard deviation to obtain a trained mixture density network model.

[0014] Compared with the existing technology, the present invention has the following advantages: it introduces a historical operating condition library and a dynamic classification threshold mechanism. By comparing the similarity between real-time operating condition characteristics and historical data, it can intelligently determine whether the current operating condition belongs to an existing category. If the operating condition is unknown, it jumps to the spectrum structure matching and deep modeling process, which has a highly flexible response capability to new operating conditions. Compared with traditional rule recognition methods, it improves the system's adaptability and scalability in complex environments. The method integrates spectral structure template matching, gated attention mechanism model feature focusing, mixed density network risk modeling, variational autoencoder state mapping and anomaly detection, realizing a complete closed loop from data preprocessing → feature extraction → risk modeling → anomaly interpretation. Compared with traditional single threshold or SVM classification methods, the accuracy and interpretability are significantly enhanced, greatly improving the degree of intelligence and manual operation and maintenance efficiency, and reducing equipment downtime and false detection and repair rates. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a flow chart of a method for online monitoring of the status of a hydropower unit based on data analysis according to the present invention; Figure 2 The figure is a structural diagram of an online monitoring system for the status of a hydropower unit based on data analysis according to the present invention. DETAILED DESCRIPTION

[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0017] See also Figure 1 The present invention provides a technical solution: a method for online monitoring of the status of a hydropower unit based on data analysis, the method comprising: Step S100: In the hydropower equipment monitoring system, the operating parameters and power parameters of the hydropower unit are collected and packaged into a data packet with a time stamp, the vibration signal in the data packet is obtained, the spectrum energy distribution curve is drawn, and the energy proportion and energy entropy characteristic value are calculated; Wherein, step S100 includes: Step S101: Several types of sensors are placed at key locations of the hydropower unit to monitor the operating parameters of the unit and simultaneously collect the power parameters output by the unit. The operating parameters and power parameters are synchronized and processed via an edge acquisition terminal, packaged into a data packet with a timestamp, and sampled according to a preset sampling period. The data packet is then uploaded to the hydropower equipment monitoring system via a communication interface. Step S102: Obtain a data packet from a hydropower equipment monitoring system, collect vibration signals of operating parameters in the data packet, decompose the vibration signal into several intrinsic mode functions by using an empirical mode decomposition method, remove redundant modes of high-frequency noise and background disturbances by analyzing spectral characteristics, retain the dominant modal components of the operating state of the hydropower unit, perform sliding window slicing on the obtained dominant modal component signals according to a preset time window, preset an overlap rate threshold, frame the signal sequence that meets the overlap rate threshold, generate a signal slice sequence within a continuous time period, extract spectral characteristics of each frame slice signal by short-time Fourier transform, further calculate the Mel-frequency cepstrum coefficient, construct a frame-level time-frequency feature matrix, and calculate the power spectral density of each frame slice signal by using the Welch method to obtain the energy distribution of the signal in the frequency domain, and draw a spectral energy distribution curve; Step S103: Slice the vibration signal after frame processing and perform multi-scale frequency domain decomposition using the wavelet packet decomposition method. Select the db6 wavelet basis and set the number of decomposition layers. Decompose each frame signal into several equal-width sub-bands. Calculate the energy value of each sub-band signal and perform normalization processing. Calculate the energy proportion according to the following formula: ; Among them, A a Expressed as the energy ratio of the ath sub-band signal, B a Expressed as the energy value of the ath sub-band signal, B b It is expressed as the energy value of the b-th sub-band signal, c is the total number of sub-band signals, and the energy entropy eigenvalue is calculated according to the following formula: ; Among them, D represents the energy entropy eigenvalue, A d Expressed as the energy proportion of the d-th sub-band signal; For example, a three-dimensional acceleration sensor is installed at the shaft end of a hydropower unit, and voltage and current sensors are deployed on the stator side. Every 0.1 second, a packet containing 1024 points of vibration data and 100 points of electrical parameters is generated and transmitted to the hydropower equipment monitoring system with a unified timestamp. Perform empirical mode decomposition on the vibration signal and decompose it into 6 IMF components. IMF5 and IMF6 are eliminated, and IMF1–IMF4 are retained. 256 points are preset per frame, with 128 points overlapping. A total of 14 frames of vibration signal are cut out. Short-time Fourier transform and Mel-frequency cepstral coefficient extraction are performed on each frame to calculate a 13-dimensional feature vector. Select the db6 wavelet basis, set the decomposition level to 3, perform full-band decomposition on each frame signal, and obtain a total of 8 sub-bands. Assuming that the energy values ​​of the 8 sub-bands are [150.0, 120.0, 100.0, 80.0, 60.0, 40.0, 30.0, 20.0], the calculated proportions are [0.25, 0.20, 0.1667, 0.1333, 0.10, 0.0667, 0.05, 0.0333]. According to the calculation results of the proportion, the energy entropy characteristic value is 2.75.

[0018] Step S200: Obtaining the power parameters of the data packet, calculating the operating condition score, determining the operating condition level of each data packet, constructing a spectrum structure reference template library, and calculating the center frequency interval of each frequency band; Wherein, step S200 includes: Step S201: Obtain a historical sampling period of the hydropower equipment monitoring system, collect data packets of a certain historical sampling period, extract the power parameters of the data packets, and calculate the operating condition score according to the following formula: ; Among them, E represents the working condition score, P represents the active power, U represents the voltage, and I represents the current; Step S202: Summarize all the working condition scores and calculate the average value of the working condition scores as E1 and the standard deviation as E2. K working condition levels are preset, and the working condition score range of each working condition level is calculated according to the following formula: ; Among them, R r It is expressed as the working condition score range of the rth working condition level, summarizing the data packets of each working condition level; Step S203: In a data packet of a certain operating condition level, extract the spectrum energy distribution curve of each data packet, construct a spectrum data set, preset a number of candidate cluster numbers, use a spectral clustering algorithm to perform cluster analysis on the spectrum data set according to each candidate cluster number, obtain a clustering result corresponding to each candidate cluster number, perform a goodness of fit evaluation on each cluster result using the Akaike Information Criterion, obtain an AIC value corresponding to each cluster result, select the candidate cluster number corresponding to the cluster result with the smallest AIC value as the cluster number of the operating condition level, and divide the spectrum of the operating condition level into multiple frequency band regions according to the cluster number to construct a spectrum structure template for the operating condition level; Step S204: Obtain all spectrum energy distribution curves in each frequency band, and perform clustering and classification based on the spectrum shape distribution rules, assign different types of labels to each frequency band, and generate a spectrum structure reference template library; Step S205: For a certain frequency band in the spectrum structure reference template library, extract the center frequency within the frequency band from the spectrum energy distribution curve of each data packet. The center frequency is the frequency point with the maximum energy value in the spectrum energy distribution curve. The average value of the center frequency is calculated as F1 and the standard deviation is calculated as F2. The center frequency interval of the frequency band is constructed as [F1-h×F2,F1+h×F2], where h represents a preset tolerance adjustment parameter. For example, if the active power is 150 kW, the voltage is 400 V, and the current is 0.5 A, the calculated operating condition score is 0.75. Assuming the mean of the operating condition score is 0.752, the standard deviation is 0.047, and the operating condition level is 3, then the operating condition range for level 1 is [0.705, 0.736), the operating condition range for level 2 is [0.736, 0.767), and the operating condition range for level 3 is [0.767, 0.798]. Assume that the spectrum energy distribution of the data packet is as shown in Table 1:

[0019] Table 1 The number of candidate clusters is preset to 2, 3, and 4. For each clustering result, we use the Akaike Information Criterion to evaluate the goodness of fit of the clustering, and calculate the following AIC values: when the number of candidate clusters is 2, the AIC is 350, when the number of candidate clusters is 3, the AIC is 340, and when the number of candidate clusters is 4, the AIC is 355. When the number of candidate clusters is 3, the AIC is the smallest, so 3 is selected as the number of clusters, and the frequency bands are divided into 3 regions: [0, 20], [20, 40], and [40, 60]. The calculated center frequency of frequency band 1 [0, 20) is 10Hz, the center frequency of frequency band 2 [20, 40) is 30Hz, and the center frequency of frequency band 3 [40, 60] is 50Hz. The calculated standard deviation of frequency band 1 is 1, and the preset tolerance adjustment parameter is 0.1. The center frequency range of frequency band 1 is [9.9, 10.1).

[0020] Step S300: Obtain a vibration signal within a center frequency interval, calculate the envelope kurtosis, spectrum centroid offset value, and ellipticity eigenvalue, train and generate a variational autoencoder model, construct a sensitive eigenvector based on the variational autoencoder model, and generate a discriminant boundary interval for each dimensional latent variable; Wherein, step S300 includes: Step S301: Based on the frequency band center frequency interval constructed in step S205, the vibration signal of the data packet in each frequency band is collected, the vibration signal in a certain frequency band is extracted through a bandpass filter, and the extracted vibration signal is subjected to Hilbert transform to obtain an envelope signal. A number of sampling points are preset, and the values ​​of the envelope signal corresponding to each sampling point are summarized. The average value L1 and standard deviation L2 of the envelope signal are calculated, and the envelope kurtosis is calculated according to the following formula: ; Among them, H represents the envelope kurtosis, L i It is represented by the value of the envelope signal corresponding to the i-th sampling point, and j is represented by the total number of sampling points; Step S302: extract the frequency value and power spectrum density value corresponding to each sampling point, and calculate the spectrum centroid offset value according to the following formula: ; Among them, N represents the spectrum centroid offset value, n i Represented as the frequency value of the i-th sampling point, m i It is represented by the power spectrum density value of the i-th sampling point, and F is represented by the mean center frequency of the current frequency band; Step S303: Obtain the axis trajectory data of the data packet. The axis trajectory data is synchronously collected by bidirectional vibration sensors arranged at key positions of the rotating shaft to form time series displacement data on a two-dimensional plane. Based on the time series displacement data, a set of trajectory coordinate points is extracted. The geometric shape of the axis trajectory is fitted using the least squares method. The length of the major semi-axis corresponding to the axis trajectory is calculated as p1 and the length of the minor semi-axis is calculated as p2. The ellipticity characteristic value of the axis trajectory is calculated according to the following formula: ; Where P represents the ellipticity eigenvalue; Step S304: Summarize the envelope kurtosis, spectrum centroid offset value, and ellipticity eigenvalue of each frequency band to construct a multidimensional feature vector of the data packet, input the multidimensional feature vector of the data packet collected under historical normal operating conditions into a variational autoencoder model for training, perform nonlinear dimensionality reduction and sensitivity enhancement processing on the multidimensional features through unsupervised learning, minimize reconstruction error and KL divergence as optimization goals, learn a low-dimensional representation of the feature distribution in the latent space under normal conditions, and output a latent variable representation vector; Step S305: Based on the trained variational autoencoder model, calculate the sensitivity index of each dimensional feature vector to the reconstruction error, construct a feature sensitivity matrix, preset the number of sensitive features to z, sort the sensitivity of all dimensional features to the reconstruction error, select the first z features with the highest contribution as sensitive features, and construct a sensitive feature vector; Step S306: Extract the latent variable representation vector set corresponding to the sensitive features, calculate the mean v1 and standard deviation v2 of the latent variable of each dimension, and construct the judgment boundary of the latent variable of a certain dimension as [vmin, vmax] = [v1-w×v2, v1+w×v2], where w represents the adjustment coefficient of the latent variable; For example, based on the center frequency interval of frequency band 1 being [9.9, 10.1), the collected vibration signal is extracted by bandpass filter and obtained as [0.02, 0.03, 0.05, 0.04, 0.02, 0.01, 0.03, 0.05, 0.02, 0.01]. The envelope signal obtained by Hilbert transform is [0.021, 0.035, 0.048, 0.041, 0.025, 0.015, 0.031, 0.049, 0.027, 0.012]. The calculated envelope kurtosis is 3.85. The frequency value of sampling point 1 is 9.95, and the power spectrum density value is 6; the frequency value of sampling point 2 is 10, and the power spectrum density value is 9; the frequency value of sampling point 3 is 10.05, and the power spectrum density value is 7. The calculated spectrum centroid offset value is -0.0013; The two-dimensional displacement point coordinates of sampling point 1 are (0.2, 0.1), the two-dimensional displacement point coordinates of sampling point 2 are (0.15, 0.12), and the two-dimensional displacement point coordinates of sampling point 3 are (0.18, 0.11). The major semi-axis and the minor semi-axis are 0.22 and 0.13, respectively, obtained by least squares fitting. The calculated ellipticity eigenvalue is 1.69. The feature vector is [3.85, -0.0013, 1.69]. The historical data is aggregated to form a complete training set, which is input into the variational autoencoder for unsupervised learning; The sensitivity of the envelope kurtosis to the reconstruction error is 0.92, the sensitivity of the spectrum centroid offset value to the reconstruction error is 0.76, and the sensitivity of the ellipticity eigenvalue to the reconstruction error is 0.89. The number of sensitive features is set to 2, and the envelope kurtosis and ellipticity eigenvalues ​​are sensitive features. Assume that the mean of the latent variable of envelope kurtosis is 3.5 and the standard deviation is 0.2, the mean of the latent variable of ellipticity eigenvalue is 1.6 and the standard deviation is 0.1, and the preset adjustment coefficient is 3. Then the judgment boundary of envelope kurtosis is [2.9, 4.1], and the judgment boundary of ellipticity eigenvalue is [1.3, 1.9].

[0021] Step S400: Setting a training cycle, extracting the time domain feature vector and frequency domain feature vector of each data packet, training and generating a gated attention mechanism model, and training and generating a mixture density network model based on the data packet containing the fault status label; Wherein, step S400 includes: Step S401: Select several consecutive days as a training cycle, extract the time domain feature vector and frequency domain feature vector of each data packet, the time domain feature vector includes the envelope kurtosis and the spectrum centroid offset value, and the frequency domain feature vector includes the power spectrum density, energy proportion, and energy entropy feature value, concatenate the time domain feature vector and the frequency domain feature vector to obtain a feature input vector, and form a training set with the state label of the data packet. Input the feature input vector into the gated attention mechanism model, calculate the frequency domain score vector, and perform back propagation training based on the cross entropy loss based on the difference between the frequency domain score vector and the target state label, thereby optimizing the weight matrix in the gated attention mechanism model, and obtaining a trained gated attention mechanism model; Step S402: Obtain a data packet containing a fault status label, perform numerical processing on the fault status, extract and fuse the feature vector processed by the gated attention mechanism model, input the feature input vector into the mixture density network model, output the conditional probability distribution of the target variable, and maximize the log-likelihood value distributed at the fault status label as the optimization target. Use the backpropagation algorithm to update the parameters in the model used to predict the mixture weight, mean, and standard deviation to obtain a trained mixture density network model.

[0022] Step S500: constructing the operating condition characteristics of the historical data packet, generating a historical operating condition library, obtaining the real-time operating condition characteristics of the real-time data packet, performing similarity matching with the historical operating condition library, collecting the maximum similarity index, calculating the dynamic classification threshold, and determining the matching status of the real-time data packet; Wherein, step S500 includes: Step S501: In the hydropower equipment monitoring system, the operating condition level of the historical data packets is obtained, the historical data packets of each operating condition level are summarized, the average load rate, average speed, and sensitive feature vector of a historical data packet in a certain operating condition level are collected, and the operating condition feature of the historical data packet is constructed as Q=(X, Y, U), where X, Y, and U represent the average load rate, average speed, and sensitive feature vector of the historical data packet, respectively. The operating condition features of all historical data packets are summarized to obtain a sample set of the operating condition level, and a historical operating condition library is generated; Step S502: Acquire real-time data packets collected in real time, extract the average load rate, average rotation speed, and sensitive feature vectors of the real-time data packets, and calculate the similarity index according to the following formula: ; Among them, G represents the similarity index, α represents the weight diagonal matrix of the working condition characteristics, and Q1 represents the real-time working condition characteristics of the real-time data packet; Step S503: Obtain similarity indexes between the real-time data packet and each operating condition feature in the historical operating condition database. The maximum value of the collected similarity indexes is Gmax. The dynamic classification threshold is calculated according to the following formula: ; Wherein, T represents the dynamic classification threshold, T1 represents the basic threshold, and β represents the adjustment coefficient. If Gmax ≥ T, the real-time data packet is determined to be similar to the historical working condition, and the known working condition processing flow is entered. If Gmax < T, step S601 is executed; For example, there are three historical data packets, the working condition characteristics of the first historical data packet are (85.2, 500, [3.9, 1.72]), the working condition characteristics of the second historical data packet are (84.5, 505, [4.0, 1.68]), and the working condition characteristics of the third historical data packet are (86.0, 498, [3.85, 1.70]), and a historical working condition library is generated; The real-time working condition feature is (85.0, 502, [3.88, 1.71]). The similarity index of the first historical data packet is calculated to be 0.9796, the similarity index of the second historical data packet is 0.892, and the similarity index of the third historical data packet is 0.9687. The maximum value of the similarity index is selected and the dynamic classification threshold is calculated to be 0.99; 0.9687<0.99, then execute step S601.

[0023] Step S600: Based on the matching status, the real-time data packet is divided into frequency bands and template matched, the real-time center frequency within each frequency band is extracted, and a real-time feature input vector is obtained. The mixed weights of several components are calculated through a gated attention mechanism model and a mixture density network model, and the real-time fault risk probability is calculated. The sensitive feature vector of the real-time data packet is input into a variational autoencoder model to obtain a latent variable representation vector. Step S600 includes: Step S601: Based on a spectrum structure reference template library, frequency band division and template matching are performed on the real-time data packet, and the real-time center frequency in each frequency band is extracted from the spectrum energy distribution curve of the real-time data packet; Step S602: Obtain the time domain feature vector and the frequency domain feature vector of the real-time data packet, and concatenate them to obtain a real-time feature input vector, input the real-time feature input vector into the gated attention mechanism model, and obtain a weighted fused feature vector; Step S603: Input the feature vector into the mixture density network model to model the fault risk probability distribution, obtain the mixed weights of several components, and calculate the real-time fault risk probability according to the following formula: ; Among them, S represents the real-time failure risk probability, S t It is represented as the mixing weight of the t-th component, and t' is represented as the number of components judged to be in a normal state; Step S604: inputting the sensitive feature vector of the real-time data packet into the trained variational autoencoder model to obtain a latent variable representation vector; For example, assuming the real-time data packet spectrum energy curve is as shown in Table 2:

[0024] Table 2 Extract the center frequency of each frequency band, and the center frequency of frequency band F1 is 40, the center frequency of frequency band F2 is 145, the center frequency of frequency band F3 is 230, the center frequency of frequency band F4 is 360, and the center frequency of frequency band F5 is 440; The real-time feature input vector obtained by splicing is [0.25, 0.03, 0.15, 3.9, 40, 0.35, 145, 0.42, 230, 0.58]. The key features are extracted by combining the gating mechanism with the attention mechanism, and the output weighted fusion feature vector is [0.22, 0.031, 0.51, 0.87]. Input into the mixed density network model, and the output is shown in Table 3:

[0025] Table 3 The calculated real-time failure risk probability is 0.75; The sensitive feature vector of the real-time data packet is [3.88, 1.71], which is input into the variational autoencoder model to obtain [0.45, -0.32].

[0026] Step S700: Based on the real-time fault risk probability, latent variable representation vector, and real-time center frequency, determine whether to analyze abnormal features. If abnormal features are to be analyzed, use the gradient approximation method to calculate the contribution of each input feature to the real-time fault risk probability, calculate the impact score, and identify input features that exceed the impact score threshold as abnormal features. Output a prompt message for verification by operation and maintenance personnel. Step S700 includes: Step S701: If a fault risk probability threshold is preset, and if the real-time fault risk probability is greater than or equal to the fault risk probability threshold, or if any dimension of the latent variable representation vector is not within the discrimination boundary interval constructed in step S306, or if the real-time center frequency is not within the center frequency interval constructed in step S205, then step S702 is executed; Step S702: Calculate the contribution of each input feature to the real-time fault risk probability using a gradient approximation method, and multiply the contribution of each input feature by the impact factor of the input feature to obtain an impact score; Step S703: Preset an impact score threshold, treat the input features corresponding to those exceeding the impact score threshold as abnormal features, and output a prompt message for operation and maintenance personnel to check; For example, the preset fault risk probability threshold is 0.6, judgment condition 1: 0.75>0.6, judgment condition 1 is satisfied, judgment condition 2: 0.45∉[0.1,0.4], -0.32∉[-0.2,0.2], judgment condition 2 is satisfied, judgment condition 3: 230∉[150,210], judgment condition 3 is satisfied, and step S702 is executed; According to the real-time feature input vector [0.25, 0.03, 0.15, 3.9, 40, 0.35, 145, 0.42, 230, 0.58], the impact score is calculated and the data is shown in Table 4:

[0027] Table 4 The preset impact score threshold is 0.05, and the abnormal features identified are kurtosis, F3 frequency, and F3 energy.

[0028] In order to better implement the above method, a hydropower unit status online monitoring system based on data analysis is proposed. The system includes a data acquisition and preprocessing module, a spectrum template module, a variational autoencoder model module, a training cycle module, a real-time operating condition matching module, a real-time monitoring module, and an abnormality judgment module. Data acquisition and preprocessing module: In the hydropower equipment monitoring system, it collects the operating parameters and power parameters of the hydropower unit, packages them into data packets with timestamps, obtains the vibration signals in the data packets, draws the spectrum energy distribution curve, and calculates the energy proportion and energy entropy characteristic value; Spectrum template module: obtains the power parameters of the data packet, calculates the operating condition score, determines the operating condition level of each data packet, builds a spectrum structure reference template library, and calculates the center frequency range of each frequency band; Variational autoencoder model module: obtains the vibration signal within the center frequency range, calculates the envelope kurtosis, spectrum centroid offset value, and ellipticity eigenvalue, trains and generates a variational autoencoder model, constructs sensitive feature vectors based on the variational autoencoder model, and generates the discriminant boundary interval of each dimensional latent variable; Training cycle module: Sets the training cycle, extracts the time domain feature vector and frequency domain feature vector of each data packet, trains and generates a gated attention mechanism model, and trains and generates a mixture density network model based on data packets containing fault status labels; Among them, the training cycle module includes the gated attention mechanism model unit and the mixed density network model unit: Gated attention mechanism model unit: Select several consecutive days as the training cycle, extract the time domain feature vector and frequency domain feature vector of each data packet, the time domain feature vector includes the envelope kurtosis and the spectrum centroid offset value, and the frequency domain feature vector includes the power spectrum density, energy proportion, and energy entropy feature value. The time domain feature vector and the frequency domain feature vector are spliced ​​to obtain a feature input vector, which is used together with the status label of the data packet to form a training set. The feature input vector is input into the gated attention mechanism model, and a frequency domain score vector is calculated. The difference between the frequency domain score vector and the target status label is used to perform back propagation training based on the cross-entropy loss, thereby optimizing the weight matrix in the gated attention mechanism model, and obtaining a trained gated attention mechanism model. Mixture density network model unit: obtains a data packet containing a fault status label, digitizes the fault status, extracts and fuses the feature vector processed by the gated attention mechanism model, inputs the feature input vector into the mixture density network model, outputs the conditional probability distribution of the target variable, maximizes the log-likelihood value distributed at the fault status label as the optimization target, and uses the backpropagation algorithm to update the parameters in the model used to predict the mixture weight, mean, and standard deviation to obtain a trained mixture density network model.

[0029] Real-time working condition matching module: constructs the working condition characteristics of historical data packets, generates a historical working condition library, obtains the real-time working condition characteristics of real-time data packets, performs similarity matching with the historical working condition library, collects the maximum value of the similarity index, calculates the dynamic classification threshold, and determines the matching status of the real-time data packets; Real-time monitoring module: Based on the matching status, the real-time data packet is divided into frequency bands and template matched, the real-time center frequency within each frequency band is extracted, and a real-time feature input vector is obtained. The mixed weights of several components are calculated through a gated attention mechanism model and a mixture density network model, and the real-time fault risk probability is calculated. The sensitive feature vector of the real-time data packet is input into the variational autoencoder model to obtain a latent variable representation vector. Abnormal judgment module: Based on the real-time fault risk probability, latent variable representation vector, and real-time center frequency, it determines whether to analyze abnormal features. If abnormal features are to be analyzed, the gradient approximation method is used to calculate the contribution of each input feature to the real-time fault risk probability, and the impact score is calculated. The input features corresponding to those exceeding the impact score threshold are regarded as abnormal features, and prompt information is output for verification by operation and maintenance personnel.

[0030] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

Claims

1. A method for online monitoring of the status of a hydropower unit based on data analysis, characterized in that: Methods include: Step S100: In the hydropower equipment monitoring system, the operating parameters and power parameters of the hydropower unit are collected and packaged into a data packet with a time stamp, the vibration signal in the data packet is obtained, the spectrum energy distribution curve is drawn, and the energy proportion and energy entropy characteristic value are calculated; Step S200: Obtaining the power parameters of the data packet, calculating the operating condition score, determining the operating condition level of each data packet, constructing a spectrum structure reference template library, and calculating the center frequency interval of each frequency band; Step S300: Obtain a vibration signal within a center frequency interval, calculate the envelope kurtosis, spectrum centroid offset value, and ellipticity eigenvalue, train and generate a variational autoencoder model, construct a sensitive eigenvector based on the variational autoencoder model, and generate a discriminant boundary interval for each dimensional latent variable; Step S400: Setting a training cycle, extracting the time domain feature vector and frequency domain feature vector of each data packet, training and generating a gated attention mechanism model, and training and generating a mixture density network model based on the data packet containing the fault status label; Step S500: constructing the operating condition characteristics of the historical data packet, generating a historical operating condition library, obtaining the real-time operating condition characteristics of the real-time data packet, performing similarity matching with the historical operating condition library, collecting the maximum similarity index, calculating the dynamic classification threshold, and determining the matching status of the real-time data packet; Step S600: Based on the matching status, the real-time data packet is divided into frequency bands and template matched, the real-time center frequency within each frequency band is extracted, and a real-time feature input vector is obtained. The mixed weights of several components are calculated through a gated attention mechanism model and a mixture density network model, and the real-time fault risk probability is calculated. The sensitive feature vector of the real-time data packet is input into a variational autoencoder model to obtain a latent variable representation vector. Step S700: Determine whether to analyze abnormal features based on the real-time fault risk probability, latent variable representation vector, and real-time center frequency. If abnormal features are to be analyzed, the gradient approximation method is used to calculate the contribution of each input feature to the real-time fault risk probability, and the impact score is calculated. The input features corresponding to those exceeding the impact score threshold are regarded as abnormal features, and a prompt message is output for verification by the operation and maintenance personnel.

2. The method for online monitoring of the status of a hydropower unit based on data analysis according to claim 1 is characterized in that: The step S100 includes the following steps: Step S101: Several types of sensors are placed at key locations of the hydropower unit to monitor the operating parameters of the unit and simultaneously collect the power parameters output by the unit. The operating parameters and power parameters are synchronized and processed via an edge acquisition terminal, packaged into a data packet with a timestamp, and sampled according to a preset sampling period. The data packet is then uploaded to the hydropower equipment monitoring system via a communication interface. Step S102: Obtain a data packet from a hydropower equipment monitoring system, collect vibration signals of operating parameters in the data packet, decompose the vibration signal into several intrinsic mode functions by using an empirical mode decomposition method, remove redundant modes of high-frequency noise and background disturbances by analyzing spectral characteristics, retain the dominant modal components of the operating state of the hydropower unit, perform sliding window slicing on the obtained dominant modal component signals according to a preset time window, preset an overlap rate threshold, frame the signal sequence that meets the overlap rate threshold, generate a signal slice sequence within a continuous time period, extract spectral characteristics of each frame slice signal by short-time Fourier transform, further calculate the Mel-frequency cepstrum coefficient, construct a frame-level time-frequency feature matrix, and calculate the power spectral density of each frame slice signal by using the Welch method to obtain the energy distribution of the signal in the frequency domain, and draw a spectral energy distribution curve; Step S103: Slice the vibration signal after frame processing and perform multi-scale frequency domain decomposition using the wavelet packet decomposition method. Select the db6 wavelet basis and set the number of decomposition layers. Decompose each frame signal into several equal-width sub-bands. Calculate the energy value of each sub-band signal and perform normalization processing. Calculate the energy proportion according to the following formula: ; Among them, A a Expressed as the energy ratio of the ath sub-band signal, B a Expressed as the energy value of the ath sub-band signal, B b It is expressed as the energy value of the b-th sub-band signal, c is the total number of sub-band signals, and the energy entropy eigenvalue is calculated according to the following formula: ; Among them, D represents the energy entropy eigenvalue, A d It is expressed as the energy proportion of the d-th sub-band signal.

3. The method for online monitoring of the status of a hydropower unit based on data analysis according to claim 2 is characterized in that: The step S200 includes the following steps: Step S201: Obtain a historical sampling period of the hydropower equipment monitoring system, collect data packets of a certain historical sampling period, extract the power parameters of the data packets, and calculate the operating condition score according to the following formula: ; Among them, E represents the working condition score, P represents the active power, U represents the voltage, and I represents the current; Step S202: Summarize all the working condition scores and calculate the average value of the working condition scores as E1 and the standard deviation as E2. K working condition levels are preset, and the working condition score range of each working condition level is calculated according to the following formula: ; Among them, R r It is expressed as the working condition score range of the rth working condition level, summarizing the data packets of each working condition level; Step S203: In a data packet of a certain operating condition level, extract the spectrum energy distribution curve of each data packet, construct a spectrum data set, preset a number of candidate cluster numbers, use a spectral clustering algorithm to perform cluster analysis on the spectrum data set according to each candidate cluster number, obtain a clustering result corresponding to each candidate cluster number, perform a goodness of fit evaluation on each cluster result using the Akaike Information Criterion, obtain an AIC value corresponding to each cluster result, select the candidate cluster number corresponding to the cluster result with the smallest AIC value as the cluster number of the operating condition level, and divide the spectrum of the operating condition level into multiple frequency band regions according to the cluster number to construct a spectrum structure template for the operating condition level; Step S204: Obtain all spectrum energy distribution curves in each frequency band, and perform clustering and classification based on the spectrum shape distribution rules, assign different types of labels to each frequency band, and generate a spectrum structure reference template library; Step S205: For a certain frequency band in the spectrum structure reference template library, extract the center frequency within the frequency band from the spectrum energy distribution curve of each data packet, where the center frequency is the frequency point with the maximum energy value in the spectrum energy distribution curve. The average value of the center frequency is calculated to be F1 and the standard deviation is calculated to be F2. The center frequency interval of the frequency band is constructed as [F1-h×F2,F1+h×F2], where h represents a preset tolerance adjustment parameter.

4. The method for online monitoring of the status of a hydropower unit based on data analysis according to claim 3 is characterized in that: The step S300 includes the following steps: Step S301: Based on the frequency band center frequency interval constructed in step S205, the vibration signal of the data packet in each frequency band is collected, the vibration signal in a certain frequency band is extracted through a bandpass filter, and the extracted vibration signal is subjected to Hilbert transform to obtain an envelope signal. A number of sampling points are preset, and the values ​​of the envelope signal corresponding to each sampling point are summarized. The average value L1 and standard deviation L2 of the envelope signal are calculated, and the envelope kurtosis is calculated according to the following formula: ; Among them, H represents the envelope kurtosis, L i It is represented by the value of the envelope signal corresponding to the i-th sampling point, and j is represented by the total number of sampling points; Step S302: extract the frequency value and power spectrum density value corresponding to each sampling point, and calculate the spectrum centroid offset value according to the following formula: ; Among them, N represents the spectrum centroid offset value, n i Represented as the frequency value of the i-th sampling point, m i It is represented by the power spectrum density value of the i-th sampling point, and F is represented by the mean center frequency of the current frequency band; Step S303: Obtain the axis trajectory data of the data packet. The axis trajectory data is synchronously collected by bidirectional vibration sensors arranged at key positions of the rotating shaft to form time series displacement data on a two-dimensional plane. Based on the time series displacement data, a set of trajectory coordinate points is extracted. The geometric shape of the axis trajectory is fitted using the least squares method. The length of the major semi-axis corresponding to the axis trajectory is calculated as p1 and the length of the minor semi-axis is calculated as p2. The ellipticity characteristic value of the axis trajectory is calculated according to the following formula: ; Where P represents the ellipticity eigenvalue; Step S304: Summarize the envelope kurtosis, spectrum centroid offset value, and ellipticity eigenvalue of each frequency band to construct a multidimensional feature vector of the data packet, input the multidimensional feature vector of the data packet collected under historical normal operating conditions into a variational autoencoder model for training, perform nonlinear dimensionality reduction and sensitivity enhancement processing on the multidimensional features through unsupervised learning, minimize reconstruction error and KL divergence as optimization goals, learn a low-dimensional representation of the feature distribution in the latent space under normal conditions, and output a latent variable representation vector; Step S305: Based on the trained variational autoencoder model, calculate the sensitivity index of each dimensional feature vector to the reconstruction error, construct a feature sensitivity matrix, preset the number of sensitive features to z, sort the sensitivity of all dimensional features to the reconstruction error, select the first z features with the highest contribution as sensitive features, and construct a sensitive feature vector; Step S306: Extract the latent variable representation vector set corresponding to the sensitive features, calculate the mean v1 and standard deviation v2 of the latent variable of each dimension, and construct the judgment boundary of the latent variable of a certain dimension as [vmin, vmax] = [v1-w×v2, v1+w×v2], where w represents the adjustment coefficient of the latent variable.

5. The method for online monitoring of the status of a hydropower unit based on data analysis according to claim 4 is characterized in that: The step S400 includes the following steps: Step S401: Select several consecutive days as a training cycle, extract the time domain feature vector and frequency domain feature vector of each data packet, the time domain feature vector includes the envelope kurtosis and the spectrum centroid offset value, and the frequency domain feature vector includes the power spectrum density, energy proportion, and energy entropy feature value, concatenate the time domain feature vector and the frequency domain feature vector to obtain a feature input vector, and form a training set with the state label of the data packet. Input the feature input vector into the gated attention mechanism model, calculate the frequency domain score vector, and perform back propagation training based on the cross entropy loss based on the difference between the frequency domain score vector and the target state label, thereby optimizing the weight matrix in the gated attention mechanism model, and obtaining a trained gated attention mechanism model; Step S402: Obtain a data packet containing a fault status label, perform numerical processing on the fault status, extract and fuse the feature vector processed by the gated attention mechanism model, input the feature input vector into the mixture density network model, output the conditional probability distribution of the target variable, and maximize the log-likelihood value distributed at the fault status label as the optimization target. Use the backpropagation algorithm to update the parameters in the model used to predict the mixture weight, mean, and standard deviation to obtain a trained mixture density network model.

6. The method for online monitoring of the status of a hydropower unit based on data analysis according to claim 5 is characterized in that: The step S500 includes the following steps: Step S501: In the hydropower equipment monitoring system, the operating condition level of the historical data packets is obtained, the historical data packets of each operating condition level are summarized, the average load rate, average speed, and sensitive feature vector of a historical data packet in a certain operating condition level are collected, and the operating condition feature of the historical data packet is constructed as Q=(X, Y, U), where X, Y, and U represent the average load rate, average speed, and sensitive feature vector of the historical data packet, respectively. The operating condition features of all historical data packets are summarized to obtain a sample set of the operating condition level, and a historical operating condition library is generated; Step S502: Acquire real-time data packets collected in real time, extract the average load rate, average rotation speed, and sensitive feature vectors of the real-time data packets, and calculate the similarity index according to the following formula: ; Among them, G represents the similarity index, α represents the weight diagonal matrix of the working condition characteristics, and Q1 represents the real-time working condition characteristics of the real-time data packet; Step S503: Obtain similarity indexes between the real-time data packet and each operating condition feature in the historical operating condition database. The maximum value of the collected similarity indexes is Gmax. The dynamic classification threshold is calculated according to the following formula: ; Wherein, T represents the dynamic classification threshold, T1 represents the basic threshold, and β represents the adjustment coefficient. If Gmax ≥ T, the real-time data packet is determined to be similar to the historical working condition and the known working condition processing flow is entered. If Gmax < T, step S601 is executed.

7. The method for online monitoring of the status of a hydropower unit based on data analysis according to claim 6 is characterized in that: The step S600 includes the following steps: Step S601: Based on a spectrum structure reference template library, frequency band division and template matching are performed on the real-time data packet, and the real-time center frequency in each frequency band is extracted from the spectrum energy distribution curve of the real-time data packet; Step S602: Obtain the time domain feature vector and the frequency domain feature vector of the real-time data packet, and concatenate them to obtain a real-time feature input vector, input the real-time feature input vector into the gated attention mechanism model, and obtain a weighted fused feature vector; Step S603: Input the feature vector into the mixture density network model to model the fault risk probability distribution, obtain the mixed weights of several components, and calculate the real-time fault risk probability according to the following formula: ; Among them, S represents the real-time failure risk probability, S t It is represented as the mixing weight of the t-th component, and t' is represented as the number of components judged to be in a normal state; Step S604: Input the sensitive feature vector of the real-time data packet into the trained variational autoencoder model to obtain a latent variable representation vector.

8. The method for online monitoring of the status of a hydropower unit based on data analysis according to claim 7 is characterized in that: The step S700 includes the following steps: Step S701: If a fault risk probability threshold is preset, and if the real-time fault risk probability is greater than or equal to the fault risk probability threshold, or if any dimension of the latent variable representation vector is not within the discrimination boundary interval constructed in step S306, or if the real-time center frequency is not within the center frequency interval constructed in step S205, then step S702 is executed; Step S702: Calculate the contribution of each input feature to the real-time fault risk probability using a gradient approximation method, and multiply the contribution of each input feature by the impact factor of the input feature to obtain an impact score; Step S703: Preset an impact score threshold, treat the input features corresponding to those exceeding the impact score threshold as abnormal features, and output prompt information for operation and maintenance personnel to check.

9. A hydropower unit status online monitoring system based on data analysis, used to implement the hydropower unit status online monitoring method based on data analysis as described in any one of claims 1 to 8, characterized in that: The system includes a data acquisition and preprocessing module, a spectrum template module, a variational autoencoder model module, a training cycle module, a real-time working condition matching module, a real-time monitoring module, and an abnormality judgment module; The data acquisition and preprocessing module: in the hydropower equipment monitoring system, collects the operating parameters and power parameters of the hydropower unit, packages them into data packets with time stamps, obtains the vibration signals in the data packets, draws the spectrum energy distribution curve, and calculates the energy proportion and energy entropy characteristic value; The spectrum template module obtains the power parameters of the data packet, calculates the working condition score, determines the working condition level of each data packet, builds a spectrum structure reference template library, and calculates the center frequency interval of each frequency band; The variational autoencoder model module obtains the vibration signal within the center frequency range, calculates the envelope kurtosis, the spectrum centroid offset value, and the ellipticity eigenvalue, trains and generates a variational autoencoder model, constructs a sensitive feature vector based on the variational autoencoder model, and generates the discriminant boundary interval of each dimensional latent variable; The training cycle module sets a training cycle, extracts the time domain feature vector and frequency domain feature vector of each data packet, trains and generates a gated attention mechanism model, and trains and generates a mixture density network model based on the data packet containing the fault status label; The real-time working condition matching module: constructs the working condition characteristics of the historical data packet, generates a historical working condition library, obtains the real-time working condition characteristics of the real-time data packet, performs similarity matching with the historical working condition library, collects the maximum value of the similarity index, calculates the dynamic classification threshold, and determines the matching status of the real-time data packet; The real-time monitoring module: based on the matching status, performs frequency band division and template matching on the real-time data packet, extracts the real-time center frequency in each frequency band, obtains the real-time feature input vector, calculates the mixed weights of several components through the gated attention mechanism model and the mixture density network model, calculates the real-time fault risk probability, and inputs the sensitive feature vector of the real-time data packet into the variational autoencoder model to obtain the latent variable representation vector; The abnormality judgment module determines whether to analyze abnormal features based on the real-time fault risk probability, the latent variable representation vector, and the real-time center frequency. If abnormal features are to be analyzed, the gradient approximation method is used to calculate the contribution of each input feature to the real-time fault risk probability, and the impact score is calculated. The input features corresponding to the impact score threshold that exceeds the impact score threshold are regarded as abnormal features, and prompt information is output for verification by operation and maintenance personnel.

10. The online monitoring system for hydropower unit status based on data analysis according to claim 9 is characterized in that: The training cycle module includes a gated attention mechanism model unit and a mixture density network model unit: The gated attention mechanism model unit: selects several consecutive days as a training cycle, extracts the time domain feature vector and frequency domain feature vector of each data packet, wherein the time domain feature vector includes the envelope kurtosis and the spectrum centroid offset value, and the frequency domain feature vector includes the power spectrum density, energy proportion, and energy entropy feature value, concatenates the time domain feature vector and the frequency domain feature vector to obtain a feature input vector, and forms a training set with the state label of the data packet, inputs the feature input vector into the gated attention mechanism model, calculates the frequency domain score vector, and performs back propagation training based on the cross entropy loss through the difference between the frequency domain score vector and the target state label, thereby optimizing the weight matrix in the gated attention mechanism model, and obtaining a trained gated attention mechanism model; The mixed density network model unit: obtains a data packet containing a fault status label, performs numerical processing on the fault status, extracts and fuses the feature vector processed by the gated attention mechanism model, inputs the feature input vector into the mixed density network model, outputs the conditional probability distribution of the target variable, maximizes the log-likelihood value distributed at the fault status label as the optimization target, and uses the backpropagation algorithm to update the parameters in the model for predicting the mixed weight, mean and standard deviation to obtain a trained mixed density network model.

Citation Information

Cited By

  • Intelligent vibration reduction control method and system for centrifugal pump

    CN121165509A

  • Operation and maintenance data transmission optimization system based on intelligent industrial sewage treatment equipment

    CN121547476A

  • Part self-adaptive machining control, machining procedure division and machining condition determination method

    CN121722038A