Turbine fault diagnosis method, diagnosis device, computer readable storage medium and computer equipment
Patent Information
- Application Number
- CN202610785952.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-09-25
AI Technical Summary
本发明解决的技术问题是:如何解决现有基于图神经网络的汽轮机故障诊断方法存在的层间表征不一致、诊断稳定性差、抗干扰能力弱、早期微弱故障识别不灵敏等技术问题
可动态适配汽轮机工况波动和传感器噪声干扰,显著提升故障诊断的稳定性、准确性和抗干扰能力,能够精准识别汽轮机不同类型、不同程度的故障,适配工业工程应用场景,实用性强、泛化能力突出,可有效降低汽轮机非计划停机概率,提升生产系统的安全性和运行效率。
Smart Images

Figure CN122817993A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of steam turbine fault diagnosis technology, and more specifically, to a steam turbine fault diagnosis method, diagnostic device, computer-readable storage medium, and computer equipment. Background Technology
[0002] As a core power source in industries such as power, chemical, and metallurgy, the stability of steam turbine operation directly determines the safety and efficiency of the entire production system. With the increasing level of industrial intelligence, data-driven fault diagnosis methods have become mainstream. Among them, Graph Neural Networks (GNNs), with their powerful structured data processing capabilities, are widely used in steam turbine fault diagnosis. The core idea is to model sensor data from various monitoring points of the steam turbine as graph data, learn the graph representation through the message passing mechanism of the GNN, and then achieve fault type identification and location.
[0003] In practical applications, the operating status of a steam turbine is monitored by various types of sensors (vibration, temperature, pressure, oil, etc.) distributed on key components such as the rotor, bearings, cylinders, and steam seals. These sensors have clear physical relationships, and fault propagation follows specific patterns. Existing steam turbine fault diagnosis methods based on graph neural networks typically use sensor data as graph node features and the physical relationships between sensors as graph edges. They learn the graph representation layer by layer using models such as GAT and GCN, and finally achieve fault diagnosis through a classifier.
[0004] However, existing methods suffer from a key technical bottleneck: inconsistency in representations between graph neural network layers. Specifically, as the number of network layers increases, local node information is continuously aggregated, while global topological information is gradually diluted. This leads to significant deviations in the relative similarity of nodes, the weight distribution of edges, and global topological features among the graph representations output by different network layers, resulting in the following problems: First, poor diagnostic stability; diagnostic accuracy fluctuates significantly under turbine operating condition fluctuations (such as load changes) and sensor noise interference. Second, insensitivity in early weak fault identification; weak fault features are easily diluted or distorted during inter-layer transmission. Third, local overfitting; excessive focus on local node information while ignoring the global physical connections between sensors and fault propagation patterns leads to insufficient model generalization ability.
[0005] Currently, existing technologies mainly employ residual connections and jumper structures to mitigate the problem of inconsistencies in interlayer representations. However, these methods only reduce information loss between layers and cannot fundamentally guarantee the relative similarity and topological consistency of interlayer representations. Some methods use attention regularization to suppress attention weight convergence, but they do not consider the application scenarios of multiple operating conditions and fault types in steam turbines, making it difficult to adapt to interference from sensor noise and operating condition fluctuations, and thus failing to meet the actual diagnostic needs of industry.
[0006] Chinese invention patent CN111832353B discloses a turbine rotor fault diagnosis method based on EMD and BA-optimized SVM. This method does not use graph neural network modeling and cannot utilize the physical correlation information between sensors, resulting in limited diagnostic accuracy and generalization ability. Chinese invention patent CN108627345A discloses a turbine system-level fault diagnosis method and system, which only achieves qualitative diagnosis by constructing a fault propagation model. It does not solve the problem of inconsistent representation between graph neural network layers and cannot adapt to the requirements of high-precision, real-time fault diagnosis.
[0007] As steam turbines develop towards higher parameters, larger capacity, and higher speeds, the types of faults are becoming increasingly complex, placing higher demands on the stability, accuracy, and anti-interference capabilities of fault diagnosis. Existing fault diagnosis methods based on graph neural networks suffer from numerous problems due to inconsistencies in inter-layer representations, failing to meet the needs of practical industrial applications. Therefore, developing a method that can effectively maintain the consistency of inter-layer representations in graph neural networks and improve the stability and accuracy of fault diagnosis has become a pressing technical problem for those skilled in the art. Summary of the Invention
[0008] (a) The technical problem to be solved by the present invention The technical problem solved by this invention is: how to solve the technical problems of existing turbine fault diagnosis methods based on graph neural networks, such as inconsistent interlayer representations, poor diagnostic stability, weak anti-interference ability, and insensitive early weak fault identification.
[0009] (II) The technical solution adopted in this invention A method for diagnosing steam turbine faults, the method comprising: Historical monitoring data of the turbine operation process is acquired, and the historical monitoring data is used to create a graph model to obtain graph data. A graph neural network model is constructed. The message passing mechanism of each graph attention layer of the graph neural network model includes a neighbor selection step guided by global topology, an attention calculation step based on fault sensitivity weighting, and a representation update step based on temporal attention enhancement. The total loss function of the graph neural network model includes a local similarity constraint loss function, a global topology constraint loss function, a fault feature alignment constraint loss function, and a classification loss function. The graph neural network model is trained using the graph data, with the goal of minimizing the total loss function. Real-time monitoring data of the turbine operation process is acquired, and the real-time monitoring data is modeled as graph data and then input into a trained graph neural network model to obtain the representations of each layer and the consistent representation. Based on the representations of each layer and the consistent representation, the fault prediction result is obtained.
[0010] Optionally, the neighbor selection step based on global topology guidance includes: calculating the comprehensive correlation degree based on the initial graph edge weight matrix and the node similarity represented by the previous layer graph, retaining valid neighbor nodes and removing redundant neighbor nodes according to the comprehensive correlation degree; The fault-sensitive weighted attention calculation step includes: calculating the node attention score based on the fault sensitivity label and the basic attention score; The representation update step based on temporal attention enhancement includes: introducing a temporal attention function, and dynamically balancing neighbor aggregation information and node-specific feature information based on the correlation of historical monitoring data.
[0011] Optionally, the local similarity constraint loss function The calculation formula is: ; In the formula, K represents the number of network layers. Let be the number of node pair combinations. Let be the cosine similarity between nodes i and j in the k-th layer.
[0012] Optionally, the global topology constraint loss function is calculated using the following formula: ; In the formula, The global topology constraint loss function. Let S be the normalized adjacency matrix of the k-th layer, S be the topological similarity weight matrix, and ⊙ be the Hadamard product. It is the square of the F-norm.
[0013] Optionally, the calculation formula for the fault feature alignment constraint loss function is as follows: ; In the formula, C represents the number of fault types. For fault weights, Let c be the set of nodes corresponding to the c-th type of fault. It represents the feature center of the c-th type of fault in the k-th layer.
[0014] Optionally, the classification loss function The calculation formula is: ; In the formula, M is the number of samples. For one-hot real labels, To predict probabilities.
[0015] Optionally, the total loss function for: ; In the formula, To constrain the weights, For the adaptive coefficient of the working condition, , , These are the weighting coefficients.
[0016] Optionally, obtaining the fault prediction result based on the representations of each layer and the consistent representation includes: Inter-layer consistency voting diagnosis is performed on the representations of each layer to obtain preliminary diagnostic results; The consistent representations of each layer are gated and weighted to obtain a global representation. The global representation is then input into a fault feature separation and enhancement process. Based on the result of the fault feature separation and enhancement process, an accurate diagnostic result is obtained. Fault prediction results are obtained based on the preliminary diagnosis results and the precise diagnosis results.
[0017] Optionally, the historical monitoring data includes vibration signals, temperature signals, pressure signals, and oil signals monitored by multiple sensors. Methods for generating graphical data by performing graph modeling on the historical monitoring data include: After performing outlier removal and noise suppression on the historical monitoring data, feature extraction is performed to obtain multi-dimensional node feature vectors. Using each sensor as a graph node and the multidimensional node feature vector as the feature of each node, an edge weight matrix of the graph is constructed based on the physical distance between sensors and the correlation of fault propagation, and the fault sensitivity label of each node is marked.
[0018] (III) Beneficial Effects The turbine fault diagnosis method disclosed in this invention has the following technical advantages compared with existing methods: It can dynamically adapt to turbine operating condition fluctuations and sensor noise interference, significantly improving the stability, accuracy, and anti-interference capability of fault diagnosis. It can accurately identify different types and degrees of turbine faults, adapt to industrial engineering application scenarios, and has strong practicality and outstanding generalization ability. It can effectively reduce the probability of unplanned turbine downtime and improve the safety and operating efficiency of the production system. Attached Figure Description
[0019] Figure 1 This is a flowchart of the main steps of a turbine fault diagnosis method according to one or more embodiments.
[0020] Figure 2 This is a schematic diagram of a graph neural network model according to one or more embodiments.
[0021] Figure 3 This is a schematic diagram of a two-branch diagnostic reasoning process for a turbine fault diagnosis method according to one or more embodiments. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0023] Before describing the various embodiments of this application in detail, the technical concept of this application is first briefly described: Existing turbine fault diagnosis methods based on graph neural networks suffer from inconsistent inter-layer representations, poor diagnostic stability, weak anti-interference capabilities, and insensitivity in identifying early, minor faults. Therefore, the turbine fault diagnosis method provided in this application has the key improvement of refining the message passing mechanism and loss function of the graph neural network model. The message passing mechanism breaks the traditional "equal aggregation of all neighbors" model, achieving collaborative learning of local fault features and global topological information. This solves the technical problems of local overfitting and global information loss in traditional message passing mechanisms, adapting to the complex fault propagation characteristics of turbines. The collaborative optimization system of consistency loss and classification loss balances inter-layer representation consistency and fault diagnosis accuracy, avoiding fault feature distortion caused by excessive constraints, and completely solving the problems of inter-layer representation disconnect and local overfitting in traditional methods. The specific principles of the turbine fault diagnosis method of this application are described below with reference to more embodiments.
[0024] Specifically, such as Figure 1 , Figure 2 and Figure 3 As shown, the turbine fault diagnosis method in this embodiment includes the following steps: Step S10: Obtain historical monitoring data of the turbine operation process, perform graph modeling on the historical monitoring data, and obtain graph data.
[0025] Step S20: Construct the graph neural network model. The message passing mechanism of each graph attention layer of the graph neural network model includes a neighbor selection step guided by global topology, an attention calculation step based on fault sensitivity weighting, and a representation update step based on temporal attention enhancement. The total loss function of the graph neural network model includes a local similarity constraint loss function, a global topology constraint loss function, a fault feature alignment constraint loss function, and a classification loss function.
[0026] Step S30: Train the graph neural network model using graph data, with the goal of minimizing the total loss function.
[0027] Step S40: Obtain real-time monitoring data of the turbine operation process, model the real-time monitoring data as graph data and input it into the trained graph neural network model to obtain the representations of each layer and the consistent representation, and obtain the fault prediction results based on the representations of each layer and the consistent representation.
[0028] In one or more embodiments, historical monitoring data includes vibration signals, temperature signals, pressure signals, and oil signals monitored by multiple sensors. After outlier removal and noise suppression processing, the historical monitoring data is used for feature extraction to obtain multi-dimensional node feature vectors. Using each sensor as a graph node and the multi-dimensional node feature vector as the feature of each node, an edge weight matrix is constructed based on the physical distance between sensors and the correlation of fault propagation, and fault sensitivity labels are used to annotate each node.
[0029] For example, multi-source sensor data from key components such as the turbine rotor, bearings, cylinders, and steam seals are collected to ensure data coverage across the entire operating range of the turbine, from 30% to 110% of its rated load. Outlier removal: Outliers in the sensor data are removed using the 3σ criterion. Data exceeding the range [μ-3σ, μ+3σ] are replaced using linear interpolation, where μ is the mean of the sensor data and σ is the standard deviation. Noise suppression: A combination of Variational Mode Decomposition (VMD) and Kalman filtering is used to denoise the sensor time-series data. The VMD decomposition is set to 6 modes and a penalty factor of 2000, while the Kalman filtering is set to a process noise variance of 10. -6 Observation noise variance 10 -4 Eliminate the effects of electromagnetic interference and mechanical noise. Feature extraction: For the denoised time series data, extract time-domain features (12 dimensions such as peak, valley, mean, and variance), frequency-domain features (40 dimensions of frequency amplitude in the range of 0-500Hz), and time-frequency joint features (12 dimensions such as wavelet packet energy entropy) to construct a 64-dimensional node feature vector.
[0030] Graph modeling: Each sensor is treated as a graph node, with node features defined by the aforementioned 64-dimensional feature vector. Based on the physical distance between sensors and the correlation of fault propagation, an edge weight matrix W (m×m, where m is the total number of sensors, 20-30) is constructed. The closer the physical distance and the stronger the correlation of fault propagation, the larger the edge weight. Simultaneously, based on the sensor's fault detection capability, a fault sensitivity label s is assigned to each node. i (The range is between 0 and 1, and the fault-sensitive node is the vibration sensor s) i =0.8-1.0, non-sensitive nodes such as ambient temperature sensors s i =0.2-0.5).
[0031] In one or more embodiments, the neighbor selection step guided by global topology includes: calculating a comprehensive correlation degree based on the initial graph edge weight matrix and the node similarity represented by the previous layer graph; retaining valid neighbor nodes and removing redundant neighbor nodes based on the comprehensive correlation degree. For example, the threshold for the comprehensive correlation degree is set to 0.3. Based on the physical correlation weights between sensors and the similarity of the previous layer representation, valid neighbor nodes are selected, and redundant neighbors with low correlation degrees are removed to avoid local information interference and to align with the turbine fault propagation path (e.g., for rotor nodes, only sensor nodes of adjacent components are selected).
[0032] The attention calculation steps based on fault sensitivity weighting include: calculating node attention scores based on fault sensitivity labels and basic attention scores. Introducing fault sensitivity labels into the attention score calculation increases the attention weight between fault-sensitive nodes, highlighting the transmission of fault-related messages and enhancing the ability to extract fault features. For example, fault sensitivity labels are incorporated when calculating the attention scores between nodes. i s j , This emphasizes message passing between fault-sensitive nodes.
[0033] The representation update steps based on temporal attention enhancement include: introducing a temporal attention function to dynamically balance neighbor aggregation information and node-specific feature information based on the correlation of historical monitoring data. A temporal attention enhancement module is added, which, combined with the characteristics of turbine sensor time-series data, captures fault mutation features (such as the impact peak of vibration signals) in the time-series data through the temporal attention function, adaptively balancing neighbor aggregation information and node-specific feature information, effectively suppressing interference from sensor noise and operating condition fluctuations. For example, a temporal attention function TA(t) is introduced. i Based on the correlation of sensor time-series data, a dynamic balance is achieved between neighbor aggregation information and node-specific feature information, TA(t) i The larger the value, the more emphasis is placed on neighbor aggregation information (during sudden fault changes), and vice versa (during stable operating conditions).
[0034] The aforementioned message passing mechanism breaks away from the traditional "equal aggregation of all neighbors" model, enabling collaborative learning of local fault characteristics and global topology information. It solves the technical problems of local overfitting and global information loss in traditional message passing mechanisms, and is adapted to the complex fault propagation characteristics of steam turbines.
[0035] In one or more embodiments, the local similarity constraint loss function The calculation formula is: ; In the formula, K represents the number of network layers. Let be the number of node pair combinations. Let be the cosine similarity between nodes i and j in the k-th layer.
[0036] The optimization objective is to make the cosine similarity of any pair of nodes (i,j) as close as possible across different layers of the network (e.g., layer 1, layer 2, ..., layer K). The result is achieved by minimizing... This forces the local relative relationships (who is more similar to whom) between nodes to remain stable during the propagation of the network at each layer, avoiding drastic changes in node representations as the number of layers increases. A fault sensitivity weight is introduced, taking into account the differences in fault detection capabilities of different turbine sensors (e.g., rotor vibration sensors are more sensitive to faults), to impose a more severe penalty on the inter-layer similarity differences of fault-sensitive nodes, ensuring the consistency of fault characteristics across layers and avoiding invalid constraints.
[0037] In one or more embodiments, the global topology constraint loss function is calculated as follows: ; In the formula, The global topology constraint loss function. Let S be the normalized adjacency matrix of the k-th layer, S be the topological similarity weight matrix, and ⊙ be the Hadamard product. It is the square of the F-norm.
[0038] The core function of this loss function is to constrain the graph topology of each layer of the network to remain consistent and to approximate the global target topology. (1) Preventing topology drift: In deep networks (such as deep GNNs), as information propagates and transforms, the adjacency relationships or node representations of each layer may gradually deviate from the original graph structure. This loss forces the normalized adjacency matrix of each layer to move towards the target topology. (2) Maintain global structural consistency: By sharing the same target matrix across all layers, ensure that the topological cognition of the network does not change drastically from shallow to deep layers. (3) Integrate topological priors: Matrix S is usually constructed from the original graph structure or some topological similarity metric (such as Jaccard similarity, cosine similarity, etc.), explicitly encoding the prior knowledge of the graph structure into the loss function. Global topological constraint loss forces the network to maintain consistency with the global graph structure at each layer by penalizing the "difference between the normalized adjacency matrix of each layer and the target topological matrix", thus avoiding the loss of topological information or structural drift in deep networks.
[0039] In one or more embodiments, the formula for calculating the fault feature alignment constraint loss function is: ; In the formula, C represents the number of fault types. For fault weights, Let c be the set of nodes corresponding to the c-th type of fault. It represents the feature center of the c-th type of fault in the k-th layer.
[0040] The loss function is essentially an intra-class compactness constraint, similar to clustering loss. (1) Clustering of the same type: It forces the feature representations of nodes of the same fault type in each layer of the network to be closely distributed around their class center, avoiding that similar samples are too scattered in the feature space. (2) Sensitivity to harm: Through weights It reflects the difference in the severity of different faults. For fault types with extremely high severity, even if a few nodes deviate from the center, a large loss value will be generated, thus forcing the network to prioritize the feature alignment accuracy of high-severity faults. (3) Cross-layer consistency: The constraint is not only effective in a single layer, but runs through all K layers of the network, ensuring that from shallow to deep layers, fault features always have good intra-class clustering and inter-class distinguishability, preventing feature drift or confusion caused by deep networks. The fault feature alignment constraint loss forces the network to tightly cluster samples of the same fault type in each layer by weighting the "distance from the feature of the same fault node to its class center", and imposes a stronger alignment constraint on high-severity faults.
[0041] By integrating global topological constraints and fault feature alignment constraints, the physical correlation between sensors and the fault propagation law are preserved through topological similarity weights, and the fault feature center alignment mechanism ensures that the node features of the same fault type in each layer are clustered around the corresponding fault center, thus avoiding fault feature offset between layers.
[0042] In one or more embodiments, the classification loss function The calculation formula is: ; In the formula, M is the number of samples. For one-hot real labels, To predict probabilities.
[0043] The constructed total loss function for: ; In the formula, The constraint weights can be adaptively adjusted between 0.8 and 1.2. , , These are weighting coefficients, with a total of 1. For example, they can be 0.45, 0.3, and 0.25 respectively. The adaptive coefficients for operating conditions are as follows: ; In the formula, P is the real-time load of the steam turbine. For rated load, The value range is 0.8-1.2.
[0044] By adding an adaptive coefficient for operating conditions, the consistency constraint strength is dynamically adjusted according to the real-time load of the steam turbine. When the operating conditions are close to the rated value, the constraints are strengthened, and when the operating conditions fluctuate greatly, the constraints are appropriately relaxed. This adapts to the actual characteristics of the steam turbine operating under varying loads and avoids constraint failure caused by operating condition fluctuations.
[0045] The above total loss function constructs a collaborative optimization system of "consistency loss + classification loss", which takes into account the consistency of inter-layer representation and the accuracy of fault diagnosis, avoids the distortion of fault features caused by excessive constraints, and completely solves the problems of inter-layer representation disconnect and local overfitting in traditional methods.
[0046] A fully connected layer and a Softmax classifier are added after the last layer of the graph neural network model to output the fault prediction probability. Simultaneously, an inter-layer consistency verification module is added to calculate the similarity consistency index ΔS and topological consistency index ΔA of each layer in real time, which are used for subsequent diagnostic inference and sensor status monitoring.
[0047] After the graph neural network model is built, it is trained and optimized.
[0048] 1. Dataset partitioning: The preprocessed graph data is divided into training, validation, and test sets in a 7:2:1 ratio. The training set is used for training model parameters, and the validation set is used for adjusting weight coefficients. , , , The test set is used to verify the final diagnostic performance of the model, including network hyperparameters (learning rate, number of iterations, etc.).
[0049] 2. Training parameter settings: The Adam optimizer is used, and the initial learning rate is set to 10. -3 The cosine annealing strategy is used for dynamic adjustment; the number of iterations is set to 300-500 rounds; and the early stopping strategy is set to stop training if the validation set loss does not decrease for 20 consecutive rounds to avoid overfitting.
[0050] 3. Optimization process: using the total loss function With the goal of minimization, the model parameters (graph attention layer parameters, consistency loss function weights, fully connected layer parameters, etc.) are updated through backpropagation. At the same time, the working condition adaptive coefficient τ(P) and the neighbor selection threshold are adjusted through the validation set to ensure that the model can maintain the consistency of inter-layer representation under different working conditions.
[0051] 4. Model Validation: After training, the model performance is validated using a test set. The diagnostic accuracy, recall, and F1 score are calculated. At the same time, the inter-layer consistency indexes ΔS≤0.05 and ΔA≤0.1 are verified. If they are not met, the weight coefficients are readjusted until the performance requirements are met.
[0052] 5. Backbone Network Freezing and Representation Extraction (Two-Stage Training Transition): After the first stage of graph neural network training meets the inter-layer consistency index and accuracy requirements, all parameters of the backbone network, such as the graph attention layer, are frozen, solidifying them into a global feature extractor. Subsequently, the training set data is input into this frozen network to extract multi-scale consistent graph representations of each sample, providing input features for multi-scale weight optimization, dictionary matrix learning, and terminal classifier fitting in the subsequent diagnostic system.
[0053] In one or more embodiments, obtaining fault prediction results based on each layer's representation and consistent representation includes: performing inter-layer consistency voting diagnosis on each layer's representation to obtain preliminary diagnosis results; performing gated weighted fusion on the consistent representations of each layer to obtain a global representation; inputting the global representation into fault feature separation and enhancement processing; obtaining accurate diagnosis results based on the result of fault feature separation and enhancement processing; and obtaining fault prediction results based on the preliminary diagnosis results and accurate diagnosis results.
[0054] For example, real-time monitoring data of the turbine operation process, after being collected and preprocessed, is modeled as graph data and input into a trained graph neural network model. Graph representations that satisfy consistency constraints at each layer are extracted, and abnormal representation layers are removed through a consistency verification module. Removing abnormal representation layers ensures that the graph representations used for diagnosis have high reliability, providing a foundation for accurate diagnosis.
[0055] For example, the representation fusion strategy of "gated weighting + fault scale adaptation" merges consistent representations from each layer into a global representation. The lower-level representation focuses on capturing early, weak fault features, while the higher-level representation focuses on capturing severe fault features, achieving the synergistic utilization of features at different fault scales. The weighting coefficients are: ,in For scale weights.
[0056] For example, a fault feature enhancement module is designed. This module separates fault features from noise features through dictionary learning and, combined with a fault-sensitive amplification factor, adaptively amplifies weak fault features to address the problem of early weak fault features being indistinct and easily masked by noise. The fused global representation is input into the fault feature enhancement module. In the offline fitting stage, the feature dictionary matrix of faults and noise is iteratively solved and constructed using the training set representation. In the online inference stage, based on the pre-trained dictionary matrix, fault features and noise features are separated, and, combined with the fault-sensitive amplification factor, weak fault features are amplified to improve feature discriminativeness.
[0057] The design incorporates a dual-branch diagnostic inference structure, combining inter-layer consistency voting diagnosis with enhanced feature-based precise diagnosis. It introduces a diagnostic confidence threshold to ensure the accuracy of fault diagnosis and provide effective early warning for suspected faults. Simultaneously, it outputs inter-layer consistency indicators, enabling synchronous monitoring of sensor operating status and adapting to the full lifecycle diagnostic needs of steam turbines, from early minor faults to severe faults.
[0058] For example, the two-branch diagnostic reasoning process is as follows: (1) Branch 1: Inter-layer consistency voting diagnosis, weighted voting is performed on the fault prediction results represented by each layer, with the weight being the consistency index of each layer (the smaller ΔS is, the greater the weight), to obtain the preliminary diagnosis results.
[0059] (2) Branch 2: Enhanced feature for accurate diagnosis. In the fitting stage, the fault features of the training set after the dictionary enhancement are input into the downstream independent classifier for secondary supervised training to fit the final fault category; in the online inference stage, the real-time enhanced fault features are input into the trained classifier to obtain accurate diagnosis results.
[0060] (3) Result fusion and early warning: The diagnostic results of the two branches are fused. If the maximum diagnostic confidence is ≥0.85, the fault type is determined. If 0.7≤confidence <0.85, the suspected fault and early warning information are output. If the confidence is <0.7, the sensor data is abnormal and the sensor needs to be calibrated. At the same time, the interlayer consistency indexes ΔS and ΔA are output to monitor the model operation status and sensor reliability.
[0061] The following detailed description of this solution, with reference to a specific embodiment, further illustrates the present solution. This embodiment uses a 300MW condensing steam turbine as the application object. This turbine is equipped with 24 monitoring sensors (8 vibration sensors, 6 temperature sensors, 5 pressure sensors, and 5 oil level sensors). Fault types include four typical faults: rotor imbalance, rotor rubbing, bearing wear, and cylinder leakage, as well as normal operating conditions, totaling five operating conditions. The specific implementation steps are as follows: I. Data Acquisition and Preprocessing from Multi-Source Sensors 1.1 Data Acquisition: Real-time operating data from 24 sensors are collected through the turbine online monitoring system. The sampling frequency is 1000Hz and the sampling duration is 10s. Each sensor collects 10,000 time-series data points. The collected operating conditions cover 30%, 50%, 70%, 90%, and 110% of the rated load (300MW). 50 sets of samples are collected for each operating condition, including 200 sets of normal state samples and 150 sets of samples for each of the four fault types, for a total of 800 sets of samples.
[0062] 1.2 Outlier Removal: Calculate the mean μ and standard deviation σ of each sensor data. For vibration sensors, μ = 0.02-0.05 mm and σ = 0.003-0.008 mm. For temperature sensors, μ = 80-120℃ and σ = 2-5℃. Use the 3σ criterion to remove outliers and replace them with linear interpolation.
[0063] 1.3 Noise Suppression: Perform VMD decomposition on the time-series data of each sensor (6 modes, penalty factor 2000, convergence accuracy 10). -7 ), filter IMF components with correlation coefficients ≥ 0.6, and then perform Kalman filtering (state transition matrix A = 1, observation matrix H = 1, process noise variance Q = 10). -6 Observation noise variance R=10 -4 ), thus completing the noise reduction.
[0064] 1.4 Feature Extraction: Extract 12-dimensional time-domain features (peak, valley, mean, etc.), 40-dimensional frequency-domain features (0-500Hz, with an amplitude value taken every 12.5Hz), and 12-dimensional wavelet packet joint features to construct a 64-dimensional node feature vector.
[0065] 1.5 Graph Modeling: Using 24 sensors as nodes, each node's feature is a 64-dimensional vector; based on sensor physical distances and fault propagation patterns, a 24×24 edge weight matrix W is constructed. The edge weights for sensors within the same bearing housing are 0.8-1.0, for adjacent components are 0.5-0.7, and for non-adjacent components are 0.1-0.3; fault sensitivity labels s are then used. i Vibration sensor i =0.8-1.0, temperature and pressure sensor s i =0.4-0.7, oil sensor s i =0.2-0.5.
[0066] II. Constructing a Graph Neural Network Model with Inter-Layer Consistency Constraints 2.1 Model Structure: A 4-layer graph attention layer is constructed, with each layer outputting a 64-dimensional graph representation; the input layer consists of a 24×64 node feature matrix and a 24×24 edge weight matrix, the middle 3 layers are improved graph attention layers, and the output layer is a fully connected layer + Softmax classifier.
[0067] 2.2 Improved message passing mechanism settings: Neighbor filtering threshold θ=0.3, temporal attention function TA(t) i The variance is calculated based on the sensor time-series data. The larger the variance (fault mutation), the higher the TA(t) value. i The larger the ) (0.6-0.9), the smaller the variance (stable operating conditions), TA(t) i The smaller (0.1-0.5).
[0068] 2.3 Setting the consistency loss function: , , Constraint weights Adaptive coefficient for operating conditions Based on real-time load dynamic adjustment, at rated load (300MW), At 30% rated load, Fault weights Rotor rubbing fault Bearing wear, rotor imbalance Cylinder leak Normal state .
[0069] 2.4 Model Initialization: The parameters of the graph attention layer are initialized using Xavier normal distribution, and the parameters of the fully connected layer are initialized using He normal distribution. The weight coefficients of the consistency loss function and the classification loss function are... .
[0070] III. Model Training and Optimization 3.1 Dataset partitioning: The 800 sets of image data were partitioned into a training set (560 sets), a validation set (160 sets), and a test set (80 sets) in a ratio of 7:2:1.
[0071] 3.2 Training parameters: Adam optimizer, initial learning rate 10. -3 The cosine annealing strategy is used, with 400 iterations and an early stopping strategy (the validation set loss stops if it does not decrease for 20 consecutive iterations).
[0072] 3.3 Optimization Process: Calculate the total loss in each iteration. Backpropagation updates the model parameters, and adjustments are made using the validation set. and The values of are selected to ensure that the diagnostic accuracy of the validation set is ≥97%, and the inter-layer consistency indexes ΔS≤0.05 and ΔA≤0.1.
[0073] 3.4 Model Validation: After training, the model performance is validated using a validation set, and the neighbor selection threshold θ and the working condition adaptive coefficient are adjusted. Finally, θ was determined to be 0.35. The adjustment range was 0.8-1.2, and the model validation accuracy reached 97.5%.
[0074] IV. Fault Diagnosis Reasoning 4.1 Real-time data processing: Collect real-time sensor data from the steam turbine, repeat the preprocessing steps 1.2-1.4, and construct real-time graph data.
[0075] 4.2 Consistency Representation Extraction: Input the real-time graph data into the trained model and extract 4 layers of graph representation. Calculate ΔS=0.032 and ΔA=0.085 through the consistency verification module. Both meet the threshold requirements. Remove abnormal representation layers (there are no abnormal layers in this embodiment).
[0076] 4.3 Multi-scale representation fusion: A gated weighted fusion strategy is adopted, with smaller weights at lower levels and larger weights at higher levels. The fusion yields a global representation.
[0077] 4.4 Fault Feature Enhancement: Fault features and noise features are separated by dictionary learning, and weak fault features are amplified by an amplification factor of 1.5-2.0.
[0078] 4.5 Dual-branch diagnostic reasoning: Branch 1 voting diagnosis accuracy is 96.8%, Branch 2 enhanced feature diagnosis accuracy is 98.2%, and the final diagnosis accuracy after fusion is 98.2%; among them, the recognition accuracy of slight rotor imbalance (fault degree ≤0.01mm) is 92.5%, which is significantly higher than the traditional GNN method (78.3%).
[0079] 4.6 Output Results: Outputs fault type (e.g., rotor imbalance), diagnostic confidence level (e.g., 0.92), and inter-layer consistency index (ΔS, ΔA). Simultaneously monitors sensor status and indicates no sensor abnormalities. The diagnostic results meet actual industrial needs.
[0080] The verification in this embodiment shows that the method of the present invention can effectively maintain the consistency of representation between layers of the graph neural network, significantly improve the accuracy, stability and anti-interference ability of turbine fault diagnosis, adapt to the actual operating conditions of 300MW condensing steam turbine, and can be extended to fault diagnosis scenarios of steam turbines of different capacities and types.
[0081] Compared with existing technologies, this solution, relying on a designed consistency loss function, an improved message passing mechanism, and a multi-scale diagnostic system, achieves the following significant benefits for applications involving multiple operating conditions, multiple sensors, and complex fault types in steam turbines: 1. Overcoming the technical bottleneck of inconsistent interlayer representations and improving diagnostic stability: Through triple consistency constraints, the graph representations of each layer are forced to maintain consistency in local similarity, global topology and fault characteristics. Under the scenario of 30%-110% rated load fluctuation of steam turbine, the fluctuation range of diagnostic accuracy is controlled within 3%, which is more than 60% more stable than the traditional GNN method (fluctuation range 8%-12%).
[0082] 2. Enhance fault feature capture capability and improve diagnostic accuracy: The improved message transmission mechanism highlights the message transmission of fault-sensitive nodes. Combined with the fault feature enhancement module, the accuracy of identifying early minor faults (such as slight rotor imbalance and slight bearing wear) is improved by more than 45%, and the overall fault diagnosis accuracy reaches 98.2%, which is better than the existing graph neural network method (accuracy of 88%-92%).
[0083] 3. Excellent anti-interference capability and adaptability to complex industrial conditions: Through VMD+Kalman filtering for noise reduction, timing attention enhancement and adaptive constraints for operating conditions, it effectively suppresses interference from sensor noise and operating condition fluctuations. Even in scenarios with noise intensity ≤10%, the diagnostic accuracy remains above 95%, making it suitable for complex industrial environments.
[0084] 4. Enables full fault lifecycle diagnosis with strong practicality: The multi-scale diagnostic system can be adapted to the full lifecycle identification of early weak faults to serious faults. The dual-branch inference structure combined with confidence warning can not only accurately diagnose faults, but also provide timely warnings of suspected faults and sensor anomalies, reducing the probability of unplanned equipment downtime.
[0085] 5. Strong generalization ability and adaptability to various turbine types: The operating condition adaptive coefficients, fault sensitivity labels and other parameters in the model can be flexibly adjusted according to the operating parameters of different capacities (such as 300MW, 600MW) and different types (condensing, back pressure) of turbines, without the need to retrain the model, and the generalization ability is outstanding.
[0086] 6. Strong engineering applicability: The entire diagnostic process can be embedded into the existing steam turbine online monitoring system without the need for large-scale modification of existing equipment. Data preprocessing and diagnostic reasoning can be automated, making operation convenient and cost controllable. It can effectively improve the safety and operating efficiency of the production system and has extremely high engineering application value and industrialization prospects.
[0087] The specific embodiments of the present invention have been described in detail above. Although some embodiments have been shown and described, those skilled in the art should understand that modifications and improvements can be made to these embodiments without departing from the principles and spirit of the present invention as defined by the claims and their equivalents, and such modifications and improvements should also be within the protection scope of the present invention.
Claims
1. A method for diagnosing steam turbine faults, characterized in that, The method includes: Historical monitoring data of the turbine operation process is acquired, and the historical monitoring data is used to create a graph model to obtain graph data. A graph neural network model is constructed. The message passing mechanism of each graph attention layer of the graph neural network model includes a neighbor selection step guided by global topology, an attention calculation step based on fault sensitivity weighting, and a representation update step based on temporal attention enhancement. The total loss function of the graph neural network model includes a local similarity constraint loss function, a global topology constraint loss function, a fault feature alignment constraint loss function, and a classification loss function. The graph neural network model is trained using the graph data, with the goal of minimizing the total loss function. Real-time monitoring data of the turbine operation process is acquired, and the real-time monitoring data is modeled as graph data and then input into a trained graph neural network model to obtain the representations of each layer and the consistent representation. Based on the representations of each layer and the consistent representation, the fault prediction result is obtained.
2. The turbine fault diagnosis method according to claim 1, characterized in that, The neighbor selection step based on global topology guidance includes: calculating the comprehensive correlation degree based on the initial graph edge weight matrix and the node similarity represented by the upper layer graph, retaining valid neighbor nodes and removing redundant neighbor nodes according to the comprehensive correlation degree; The fault-sensitive weighted attention calculation step includes: calculating the node attention score based on the fault sensitivity label and the basic attention score; The representation update step based on temporal attention enhancement includes: introducing a temporal attention function, and dynamically balancing neighbor aggregation information and node-specific feature information based on the correlation of historical monitoring data.
3. The turbine fault diagnosis method according to claim 1, characterized in that, The local similarity constraint loss function The calculation formula is: ; In the formula, K represents the number of network layers. Let be the number of node pair combinations. Let be the cosine similarity between nodes i and j in the k-th layer.
4. The turbine fault diagnosis method according to claim 3, characterized in that, The formula for calculating the global topology constraint loss function is as follows: ; In the formula, The global topology constraint loss function. Let S be the normalized adjacency matrix of the k-th layer, S be the topological similarity weight matrix, and ⊙ be the Hadamard product. It is the square of the F-norm.
5. The turbine fault diagnosis method according to claim 4, characterized in that, The formula for calculating the fault feature alignment constraint loss function is as follows: ; In the formula, C represents the number of fault types. For fault weights, Let c be the set of nodes corresponding to the c-th type of fault. It represents the feature center of the c-th type of fault in the k-th layer.
6. The turbine fault diagnosis method according to claim 5, characterized in that, The classification loss function The calculation formula is: ; In the formula, M is the number of samples. For one-hot real labels, To predict probabilities.
7. The turbine fault diagnosis method according to claim 6, characterized in that, The total loss function for: ; In the formula, To constrain the weights, For the adaptive coefficient of the working condition, , , These are the weighting coefficients.
8. The turbine fault diagnosis method according to claim 1, characterized in that, The fault prediction result obtained based on the representations of each layer and the consistent representation includes: Inter-layer consistency voting diagnosis is performed on the representations of each layer to obtain preliminary diagnostic results; The consistent representations of each layer are gated and weighted to obtain a global representation. The global representation is then input into a fault feature separation and enhancement process. Based on the result of the fault feature separation and enhancement process, an accurate diagnostic result is obtained. Fault prediction results are obtained based on the preliminary diagnosis results and the precise diagnosis results.
9. The turbine fault diagnosis method according to claim 1, characterized in that, The historical monitoring data includes vibration signals, temperature signals, pressure signals, and oil signals monitored by multiple sensors. Methods for generating graphical data by creating graphical models from the historical monitoring data include: After performing outlier removal and noise suppression on the historical monitoring data, feature extraction is performed to obtain multi-dimensional node feature vectors. Using each sensor as a graph node and the multidimensional node feature vector as the feature of each node, an edge weight matrix of the graph is constructed based on the physical distance between sensors and the correlation of fault propagation, and the fault sensitivity label of each node is marked.
Citation Information
Patent Citations
Steam turbine system-level fault diagnosing method and system
CN108627345A
A Steam Turbine Rotor Fault Diagnosis Method Based on EMD and BA-Optimized SVM
CN111832353B