Method and system for unit-wide fault diagnosis and root cause tracing under all operating conditions of risk classification
Patent Information
- Application Number
- CN202611081697.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-21
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-07-21
AI Technical Summary
基于阈值与逻辑规则的监测报警系统:技术人员尝试优化参数阈值、增加连锁逻辑减少误报,但仍无法区分瞬态过程的正常参数波动与真实故障前兆,误报率高达40%以上,无根因溯源能力,仅能实现简单报警
全工况无盲区覆盖:通过自适应工况识别和特征提取,实现冷态启动、升负荷、基本负荷、停机等全工况诊断,诊断覆盖率大幅提升,解决了动态工况诊断盲区问题。
Smart Images

Figure CN122595164B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power generation equipment condition monitoring and fault diagnosis technology, specifically involving a risk-classified method and system for fault diagnosis and root cause tracing of units under all operating conditions. Background Technology
[0002] Combined cycle (CCPP) gas turbine units have become a core component of modern power systems due to their high efficiency, low emissions, and excellent peak-shaving performance. These units consist of a gas turbine, a waste heat boiler, a steam turbine, and multiple auxiliary systems coupled together. Their operating conditions change frequently with the grid load, and they exhibit diverse fault modes and overlapping symptoms, placing high demands on the timeliness, accuracy, and root cause analysis capabilities of fault diagnosis.
[0003] Currently, in order to solve the fault diagnosis problem of CCPP units, those skilled in the art have successively developed monitoring and alarm systems based on thresholds and logical rules, single-model diagnostic methods based on data-driven approaches, diagnostic methods based on expert systems, and diagnostic models oriented towards specific operating conditions. Each technology has undergone targeted optimization efforts, but there are still insurmountable shortcomings: Monitoring and alarm systems based on thresholds and logic rules: Technicians have tried to optimize parameter thresholds and add interlocking logic to reduce false alarms, but they still cannot distinguish between normal parameter fluctuations in transient processes and real fault precursors. The false alarm rate is as high as 40% or more, and there is no root cause tracing capability. They can only achieve simple alarms.
[0004] Data-driven single-model diagnostic methods: Technicians improve diagnostic results by expanding training datasets and optimizing machine learning model structures, but they rely on high-quality labeled training data, have poor adaptability to unseen faults and dynamic operating conditions, and the model is a "black box" structure with no interpretability of diagnostic results.
[0005] Diagnostic methods based on expert systems: Technicians improve diagnostic capabilities by continuously supplementing the rule base, but rule construction and maintenance are cumbersome, knowledge updates are lagging, and they cannot cover complex coupled faults. When faced with new models or abnormal operating conditions that are not covered, their generalization ability is weak, and the fault identification rate is less than 50%.
[0006] Diagnostic models for specific operating conditions: Technicians have optimized thermal analysis and efficiency calculation models for steady-state operating conditions, which has improved the diagnostic accuracy under steady-state conditions. However, they cannot adapt to dynamic operating conditions such as start-up, shutdown, and load changes. The diagnostic system has "operating condition blind spots" and the full operating condition diagnostic coverage is less than 70%, which makes it impossible to achieve full life cycle health management of the unit.
[0007] In summary, the core common problems of existing technologies are: poor coverage of all operating conditions and lack of dynamic operating condition diagnostic capabilities; isolated diagnostic methods that fail to achieve complementary advantages and combine intelligence and interpretability; lack of fault risk classification, resulting in a massive number of alarms drowning out real danger signals; and insufficient root cause tracing capabilities, making it impossible to locate the initial stage of the fault. Therefore, there is an urgent need in this field for a fault diagnosis solution for CCPP units that covers all operating conditions, integrates multiple methods, classifies risks, accurately traces root causes, and possesses self-evolving capabilities.
[0008] To this end, we propose a risk-based method and system for fault diagnosis and root cause tracing of units under all operating conditions. Summary of the Invention
[0009] The present invention aims to solve at least one of the technical problems existing in the prior art, and to provide a method and system for risk-classified unit full-condition fault diagnosis and root cause tracing.
[0010] One aspect of the present invention provides a risk-level, full-condition fault diagnosis and root cause tracing method for generating units, comprising the following steps: S1: Collecting raw data across all dimensions and performing data preprocessing; performing intelligent identification of all operating conditions based on the preprocessed data; and extracting and optimizing features from the identified operating condition data to construct a feature vector set; S2: Periodically scanning the feature vector set to determine whether it triggers a safety red line rule; if the safety red line rule is triggered, an emergency response command is output; if the safety red line rule is not triggered, the feature vector set is transmitted to a secondary risk parallel diagnosis channel; S3: Parallel initiating a data-driven anomaly detection and pattern matching diagnosis process and a knowledge-driven root cause reasoning and strategy optimization diagnosis process to implement dual collaborative diagnosis and output diagnostic results including anomaly type hypothesis, root cause hypothesis, and corresponding confidence level; S4: Employing a hybrid decision-making architecture combining strong evidence priority and dynamic weighting to fuse the diagnostic results to obtain a comprehensive diagnostic conclusion and generate a structured, graded diagnostic report; S5: Collecting operation and maintenance verification results to update the fault case library, optimize the knowledge graph, and fine-tune model parameters to achieve system self-evolution. Further, in step S1, the data preprocessing includes data cleaning and hard synchronization, wherein the data cleaning includes: based on 3 The criteria include outlier detection and removal, short-term missing data completion based on linear interpolation, and marking and blanking of long-term missing data; the hard synchronization includes: full data time synchronization through hardware clock.
[0011] Furthermore, in step S1, the full-condition intelligent identification includes: using a condition classifier with the unit load command and the generator outlet circuit breaker status as the core basis, and combining the load change rate threshold to determine the cold start, load increase, basic load operation, and shutdown conditions. The load increase threshold is ≥1.5% of rated load / minute, and the basic load threshold is ≤0.5% of rated load / minute.
[0012] Furthermore, in step S1, the feature extraction and optimization of the identified working condition data includes: for different working conditions, key features are adaptively extracted using a preset feature extraction template, and a feature vector set is constructed after Z-score standardization.
[0013] Furthermore, in step S2, the safety red line rule is an IF-THEN hard logic rule related to unit safety, including at least one of unit overspeed, overtemperature, and shutdown.
[0014] Furthermore, in step S3, the data-driven anomaly detection and pattern matching diagnosis process constructs a healthy cluster model through the DBSCAN clustering algorithm and uses cosine similarity calculation to achieve anomaly pattern matching; the knowledge-driven root cause reasoning and strategy optimization diagnosis process performs root cause causal reasoning through the unit's multi-physics knowledge graph and uses a reinforcement learning agent to generate optimization strategies, wherein the reinforcement learning agent uses the operation and maintenance verification results as a reward signal to iterate the strategy.
[0015] Furthermore, step S4 specifically includes: prioritizing high-confidence diagnostic results with a confidence level > 0.95 and no conflict; otherwise, dynamically calculating the normalized weights of data-driven and knowledge-driven approaches based on the data noise level and the credibility coefficient of expert experience, and performing weighted voting fusion; simultaneously constructing a risk matrix based on the severity of the fault consequences and the speed of fault development to quantify the risk level, and finally generating a structured hierarchical diagnostic report.
[0016] Another aspect of the present invention provides a risk-classified unit full-condition fault diagnosis and root cause tracing system for implementing the risk-classified unit full-condition fault diagnosis and root cause tracing method described above. The system includes: a data perception preprocessing module, a risk diversion diagnosis module, a parallel collaborative diagnosis module, a dynamic fusion decision module, and a self-learning evolution module. These modules are sequentially and bidirectionally connected via an industrial communication protocol. The data perception preprocessing module performs data acquisition and preprocessing, full-condition intelligent identification, and condition feature extraction and optimization based on the identification results to generate a feature vector set. The risk diversion diagnosis module periodically scans the feature vector set and performs safety red line rule judgment on it. If the judgment result is triggered, an emergency response command is generated and output; if the judgment result is not triggered, the... The feature vector set is transmitted to the secondary risk parallel diagnosis channel; the parallel collaborative diagnosis module is used to initiate data-driven anomaly detection and pattern matching diagnosis process and knowledge-driven root cause reasoning and strategy optimization diagnosis process in parallel to implement dual collaborative diagnosis and output diagnosis results including anomaly type hypothesis, root cause hypothesis and corresponding confidence level; the dynamic fusion decision module is used to generate comprehensive diagnosis conclusions by adopting a hybrid decision architecture that combines strong evidence priority and dynamic weight fusion, and generate structured hierarchical diagnosis reports based on the comprehensive diagnosis conclusions; the self-learning evolution module forms a closed-loop feedback connection with the parallel collaborative diagnosis module and the dynamic fusion decision module, and is used to collect operation and maintenance verification results, and update the fault case library, optimize the knowledge graph and fine-tune the model parameters based on the operation and maintenance verification results to achieve the self-evolution of the system.
[0017] Furthermore, the parallel collaborative diagnostic module includes a data-driven diagnostic unit and a knowledge-driven diagnostic unit that are independent of each other. The data-driven diagnostic unit is equipped with a DBSCAN clustering algorithm program and a fault case library, while the knowledge-driven diagnostic unit is equipped with a knowledge graph management system and a reinforcement learning agent program.
[0018] Furthermore, the hardware carrier of the system is an industrial server, a data acquisition terminal, and a diagnostic terminal; the software carrier is an embedded program and industrial diagnostic system software; and the industrial communication protocol is at least one of OPC DA, OPC UA, and OPC API.
[0019] The beneficial effects of this invention are as follows: Full-condition coverage without blind spots: Through adaptive condition identification and feature extraction, full-condition diagnosis is achieved, including cold start, load increase, basic load, and shutdown. The diagnostic coverage is greatly improved, and the problem of blind spots in dynamic condition diagnosis is solved.
[0020] Accurate fault risk classification and improved response speed: Emergency faults achieve millisecond-level response of ≤100 milliseconds through hard logic rules, while non-emergency faults enter in-depth analysis, effectively avoiding the flood of alarms that overwhelm dangerous signals, and the risk classification has extremely high accuracy.
[0021] Significantly improved diagnostic accuracy and root cause localization capability: Through data and knowledge-driven parallel collaborative diagnosis, combined with dynamic weight fusion, the false alarm rate of diagnosis is greatly reduced, the accuracy of root cause localization is greatly improved, and the diagnostic results have a visualized reasoning path, solving the problem of the lack of interpretability of the "black box" model.
[0022] Enhanced self-learning and evolution capabilities of the system: Through the closed loop of operation and maintenance feedback, the case library, knowledge graph, and diagnostic model are continuously optimized, and the ability to identify new faults is constantly improved with the accumulation of operation and maintenance data, eliminating the need for frequent manual model adjustments and reducing operation and maintenance costs.
[0023] Guided tiered operation and maintenance enhances unit operation safety: Tiered diagnostic reports provide targeted handling suggestions, enabling immediate handling of emergency faults and advance planning of maintenance for medium / low-risk faults, effectively reducing the number of unplanned unit shutdowns and improving unit operation stability and safety. Attached Figure Description
[0024] Figure 1 This is a flowchart of the steps of the risk-classified unit full-condition fault diagnosis and root cause tracing method according to an embodiment of the present invention. Figure 2 This is a flowchart of a data-driven diagnosis method for risk-level unit full-condition fault diagnosis and root cause tracing in an embodiment of the present invention. Figure 3 This is a schematic diagram of the knowledge graph structure of the risk-level unit full-condition fault diagnosis and root cause tracing method in an embodiment of the present invention. Figure 4 This is a schematic diagram of the risk-classified unit full-condition fault diagnosis and root cause tracing system according to an embodiment of the present invention. Figure 5 This is a flowchart illustrating the risk-based unit full-condition fault diagnosis and root cause tracing method according to an embodiment of the present invention. Detailed Implementation
[0025] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0026] Please refer to the following: Figure 1The risk-level unit full-condition fault diagnosis and root cause tracing method provided by the specific embodiments of the present invention includes the following steps: S1, collecting raw data from all dimensions and performing data preprocessing, performing intelligent identification of all operating conditions based on the preprocessed data, and extracting and optimizing features from the identified operating condition data to construct a feature vector set; S2, periodically scanning the feature vector set to determine whether it triggers the safety red line rule. If the safety red line rule is triggered, an emergency response command is output; if the safety red line rule is not triggered, the feature vector set is transmitted to the secondary risk level. Parallel diagnostic channels; S3, Initiate data-driven anomaly detection and pattern matching diagnostic processes and knowledge-driven root cause reasoning and strategy optimization diagnostic processes in parallel to implement dual collaborative diagnosis and output diagnostic results including anomaly type hypothesis, root cause hypothesis, and corresponding confidence level; S4, Employ a hybrid decision-making architecture combining strong evidence priority and dynamic weighting to fuse the diagnostic results to obtain a comprehensive diagnostic conclusion and generate a structured hierarchical diagnostic report; S5, Collect operation and maintenance verification results to update the fault case library, optimize the knowledge graph, and fine-tune model parameters to achieve system self-evolution.
[0027] Specifically, the core of this method lies in: separating urgent and non-urgent faults through a risk diversion mechanism to ensure millisecond-level response to safety-related faults; achieving complementary advantages between anomaly detection and root cause reasoning by parallel initiation of both data-driven and knowledge-driven diagnostic processes; and finally, intelligently fusing the two diagnostic results through a hybrid decision-making architecture, balancing the authority of high-confidence evidence with adaptability to dynamic environments, thereby constructing a fault diagnosis system that covers all operating conditions, is accurate and efficient, and has self-evolving capabilities. The following will elaborate on the above core processes of this invention in conjunction with specific implementation details.
[0028] The aforementioned full-dimensional raw data refers to data collected from multiple perspectives, including physical signals, business environment, and time span, during intelligent diagnostics of industrial equipment.
[0029] In this embodiment, in step S1, the data preprocessing includes data cleaning and hard synchronization. The data cleaning includes: based on 3D... The criteria include outlier detection and removal, short-term missing data completion based on linear interpolation, and marking and blanking of long-term missing data; the hard synchronization includes: full data time synchronization through hardware clock.
[0030] In this embodiment, in step S1, the full-condition intelligent identification includes: using a condition classifier with the unit load command and the generator outlet circuit breaker status as the core basis, and combining the load change rate threshold to determine the cold start, load increase, basic load operation, and shutdown conditions. The load increase threshold is ≥1.5% of rated load / minute, and the basic load threshold is ≤0.5% of rated load / minute.
[0031] The term "full operating conditions" refers to the complete set of all possible operating states of the unit, from startup, shutdown, normal operation, low-load peak shaving to extreme abnormalities.
[0032] In this embodiment, in step S1, the feature extraction and optimization of the identified operating condition data includes: for different operating conditions, adaptively extracting key features using a preset feature extraction template, and constructing a feature vector set after Z-score standardization. Specifically, in step S1, firstly, the data acquisition server communicates with the factory-level distributed control system (DCS) and / or safety instrumented system (SIS) via the OPC DA (OLE Data Access for Process Control) protocol and / or OPC UA (Open Platform Unified Communication Architecture) protocol to collect more than 2000 process parameters in real time, including but not limited to temperature, pressure, and flow rate; simultaneously, the data acquisition server also obtains vibration data of bearings of key rotating equipment (such as gas turbines and / or steam turbines) in the X-axis and / or Y-axis directions from the vibration monitoring system via the application programming interface (API); in addition, the data acquisition server also obtains multiple performance indicators, such as heat rate and compressor efficiency, from the performance calculation module via the API, wherein the number of performance indicators is at least 20; thereby obtaining full-dimensional raw data. Then, based on 3 The criteria involve removing outliers from the collected raw data, using linear interpolation to fill in short-term missing segments, and relying on a hardware clock to achieve microsecond-level time synchronization of multi-source data. Subsequently, a load condition classifier monitors load commands and their rate of change in real time. When the load command exceeds 95% of the rated value and the rate of change is less than 0.5% for 10 consecutive minutes, the current state is marked as "basic load operation." Based on this, the system calls a preset feature extraction template to calculate 50 key features from the preprocessed data, and performs Z-score standardization to eliminate dimensional differences, ultimately generating the feature vector set representing the current operating state of the equipment. .
[0033] In this embodiment, in step S2, the safety red line rule is an IF-THEN hard logic rule related to unit safety, including at least one of unit overspeed, overtemperature and shutdown.
[0034] Specifically, in step S2, the feature vector set is scanned at 100-millisecond intervals. If the feature vector set triggers any of the preset safety red line rules for unit overspeed, overtemperature, or shutdown, the system immediately classifies the current state as an emergency fault of the highest risk level and activates the first rapid diagnostic path. Under this path, the system is configured to skip subsequent complex iterative analysis, directly issue emergency response commands to the DCS control system, and simultaneously trigger the highest-level red alarm on the human-machine interface to achieve millisecond-level emergency response and closed-loop control. Conversely, if the feature vector set does not trigger any safety red line rules, the system classifies the current state as a non-emergency state and transmits it to the secondary risk parallel collaborative diagnostic channel for subsequent in-depth analysis and refined diagnosis.
[0035] Please refer to the following as well. Figure 2 and Figure 3 Furthermore, in step S3, the data-driven anomaly detection and pattern matching diagnosis process constructs a healthy cluster model through the DBSCAN clustering algorithm and uses cosine similarity calculation to achieve anomaly pattern matching; the knowledge-driven root cause reasoning and strategy optimization diagnosis process performs root cause causal reasoning through the unit's multi-physics knowledge graph and uses a reinforcement learning agent to generate optimization strategies, wherein the reinforcement learning agent uses the operation and maintenance verification results as reward signals to iterate the strategy.
[0036] The core of the secondary risk parallel diagnostic channel lies in the parallel initiation of data-driven anomaly detection and pattern matching diagnostic processes and knowledge-driven root cause reasoning and strategy optimization diagnostic processes to achieve dual collaborative diagnosis. These two processes operate independently and in parallel, enabling analysis of equipment status from different dimensions and achieving complementary advantages.
[0037] Specifically, in step S3, the secondary risk parallel collaborative diagnosis channel initiates a data-driven anomaly detection and pattern matching diagnosis process and a knowledge-driven root cause reasoning and strategy optimization diagnosis process. The specific steps of the data-driven anomaly detection and pattern matching diagnosis process are as follows: Based on the historical health baseline database and fault case database, calculate the cosine similarity between the real-time feature vector and the feature vector in the fault case database under the same operating condition, and output the top 3 fault cases with the highest similarity and their matching degree; simultaneously, input the real-time feature vector into the DBSCAN clustering health cluster model of the current operating condition, calculate its distance D to the nearest core sample point, and if D exceeds a preset threshold... If an anomaly is detected, it is classified as an unknown anomaly and an anomaly index is quantified. The final output includes a list of anomaly type hypotheses, a quantified matching degree, and an unknown anomaly index. The specific steps of the knowledge-driven root cause reasoning and strategy optimization diagnostic process are as follows: Based on a multi-physics knowledge graph of the unit (stored in a graph database, with 0-1 confidence weights assigned to the may_cause relationship), with nodes representing equipment, measurement points, faults, and symptoms, and relationships represented by has_part, has_symptom, and may_cause, fault chain tracing reasoning is performed starting from nodes with significant anomaly symptoms. Simultaneously, a reinforcement learning agent (state being an active alarm set, action being fault node selection, and reward being maintenance verification results) optimizes the diagnostic strategy. Finally, the root cause hypothesis with the highest probability, confidence level, and visualized causal reasoning path are output. The advantage of this process is that its diagnostic results are highly interpretable, clearly demonstrating the propagation path and causal relationships of the fault.
[0038] The data-driven and knowledge-driven processes described above are not simply sequential or selective, but rather employ a parallel and collaborative strategy: both are initiated simultaneously and operate independently, ultimately outputting diagnostic results that include anomaly type hypotheses, root cause hypotheses, and corresponding confidence levels. This parallel and collaborative design allows the "pattern discovery" capability of the data-driven process and the "causal reasoning" capability of the knowledge-driven process to be utilized simultaneously, providing a complementary and complete evidentiary basis for subsequent fusion decision-making, thereby overcoming the inherent limitations of single diagnostic methods.
[0039] Furthermore, step S4 specifically includes: prioritizing high-confidence diagnostic results with a confidence level > 0.95 and no conflict; otherwise, dynamically calculating the normalized weights of data-driven and knowledge-driven approaches based on the data noise level and the credibility coefficient of expert experience, and performing weighted voting fusion; simultaneously constructing a risk matrix based on the severity of the fault consequences and the speed of fault development to quantify the risk level, and finally generating a structured hierarchical diagnostic report.
[0040] Step S4 employs a hybrid decision-making architecture combining strong evidence priority and dynamic weighting to fuse and comprehensively evaluate the diagnostic results output by the data-driven and knowledge-driven diagnostic processes. The design principle of this hybrid decision-making architecture is to ensure that high-confidence evidence receives absolute priority while also considering the system's adaptability and robustness in dynamic and complex environments. The fusion strategy in step S4 specifically includes: first, the system performs conflict detection and confidence assessment on the two sets of diagnostic results. If the confidence level of any result is >0.95 and does not contradict the other result, it is determined to be "strong evidence," and this high-confidence result is directly adopted as the comprehensive diagnostic conclusion, thereby ensuring the decisiveness and authority of the decision when the evidence is sufficient and clear.
[0041] If no strong evidence meeting the above conditions is found, the system automatically switches to dynamic weight fusion mode. In this mode, the system dynamically calculates the normalized weights of the two sets of diagnostic results based on the current data noise level and the confidence coefficient of expert experience. Specifically, data-driven diagnostic weights are used. =Confidence level × (1 - data noise level), knowledge-driven diagnostic weights = Confidence Level × Expert Experience Credibility Coefficient. Here, the data noise level reflects the quality of the currently collected data; when data fluctuates drastically or has significant gaps, the confidence level of data-driven diagnosis will be automatically lowered. The expert experience credibility coefficient is dynamically maintained based on historical verification results, reflecting the reliability of the knowledge graph under different operating conditions. By weighted voting fusion of the normalized weights, the system can adaptively adjust the dependence on the two sets of diagnostic results under different operating environments, achieving the optimal fusion effect.
[0042] Building upon this foundation, the system constructs a fault risk matrix based on the severity of fault consequences (S, 1-5 points) and the speed of fault development (L, 1-5 points). Risk levels are categorized using a risk value R = S × L (R ≥ 20 for emergency, 15 ≤ R < 20 for high risk, 10 ≤ R < 15 for medium risk, and R < 10 for low risk), thus quantifying and classifying the risk of diagnostic conclusions. Finally, a structured, graded diagnostic report is generated, including operating conditions, fault / anomaly descriptions, risk levels, root cause analysis, recommended handling priorities, and reasoning basis. This report is then pushed to maintenance personnel via WeChat / SMS. This hybrid decision-making architecture, through a dual mechanism of "strong evidence priority" and "dynamic weight fusion," ensures both the reliability and authority of key diagnostic conclusions while endowing the system with flexibility and intelligence to cope with complex and ever-changing field environments.
[0043] In this embodiment, the specific steps of step S5 are as follows: Collect maintenance verification results through the "Confirm / Correction / False Alarm" buttons on the structured hierarchical diagnostic report interface; if the verification result is "Confirm," package the feature vector, diagnostic result, and final conclusion of this diagnosis to generate a new fault case and update it to the fault case library; optimize the knowledge graph based on the verification results, increase the weight of the may_cause relation in the correct reasoning path, and decrease the weight of the incorrect reasoning path; fine-tune the threshold and similarity matching weight vector of the DBSCAN clustering model; add the positive / negative rewards of the maintenance verification to the experience replay pool of the reinforcement learning agent, and train and update the neural network strategy offline.
[0044] Please refer to the following as well. Figure 4 and Figure 5Another aspect of the present invention provides a risk-classified unit full-condition fault diagnosis and root cause tracing system for implementing the risk-classified unit full-condition fault diagnosis and root cause tracing method described above. The system includes: a data perception preprocessing module 41, a risk diversion diagnosis module 42, a parallel collaborative diagnosis module 43, a dynamic fusion decision module 44, and a self-learning evolution module 45. These modules are sequentially and bidirectionally connected via an industrial communication protocol. The data perception preprocessing module 41 performs data acquisition and preprocessing, full-condition intelligent identification, and condition feature extraction and optimization based on the identification results to generate a feature vector set. The risk diversion diagnosis module 42 periodically scans the feature vector set and performs safety red line rule judgment on it. If the judgment result is triggered, an emergency response command is generated and output; if the judgment result is not triggered, an emergency response command is generated and output. The feature vector set is then transmitted to the secondary risk parallel diagnosis channel; the parallel collaborative diagnosis module 43 is used to initiate the data-driven anomaly detection and pattern matching diagnosis process and the knowledge-driven root cause reasoning and strategy optimization diagnosis process in parallel to implement dual collaborative diagnosis and output diagnosis results including anomaly type hypothesis, root cause hypothesis and corresponding confidence level; the dynamic fusion decision module 44 is used to generate a comprehensive diagnosis conclusion using a hybrid decision architecture that combines strong evidence priority and dynamic weight fusion, and generate a structured hierarchical diagnosis report based on the comprehensive diagnosis conclusion; the self-learning evolution module 45 forms a closed-loop feedback connection with the parallel collaborative diagnosis module and the dynamic fusion decision module, and is used to collect operation and maintenance verification results, and update the fault case library, optimize the knowledge graph and fine-tune the model parameters based on the operation and maintenance verification results to achieve the self-evolution of the system.
[0045] In this embodiment, a DCS control system 46 and an operation and maintenance terminal 47 are also included.
[0046] In this embodiment, the parallel collaborative diagnosis module 43 includes a data-driven diagnosis unit and a knowledge-driven diagnosis unit that are independent of each other. The data-driven diagnosis unit is equipped with a DBSCAN clustering algorithm program and a fault case library, while the knowledge-driven diagnosis unit is equipped with a knowledge graph management system and a reinforcement learning agent program.
[0047] In this embodiment, the hardware carrier of the system is an industrial server, a data acquisition terminal, and a diagnostic terminal; the software carrier is an embedded program and industrial diagnostic system software; and the industrial communication protocol is at least one of OPC DA, OPC UA, and OPC API.
[0048] Specifically, the hardware carriers of the system include the data acquisition server (with a 1-second acquisition cycle), the diagnostic server (equipped with a multi-core processor and supporting parallel computing) and the operation and maintenance terminal 47 (engineer's operating terminal) deployed in the information management area. All hardware components communicate via industrial Ethernet, using OPC DA / OPC UA and API protocols. The software carriers include the data acquisition embedded program, the DBSCAN clustering algorithm program, the reinforcement learning agent program, the knowledge graph management system, and the diagnostic report generation system, all of which are deployed in the diagnostic server.
[0049] To make the objectives, technical solutions, and advantages of this invention clearer, the following will be combined with... Figure 5 The method and system for fault diagnosis and root cause tracing of units under all operating conditions based on the risk classification are described in detail below.
[0050] The specific steps are as follows: First, the data acquisition server in the data sensing preprocessing module 41 acquires over 2000 process parameters such as temperature, pressure, and flow rate from the DCS / SIS system via the OPC DA / OPC UA protocol; it also acquires X / Y vibration data of the gas turbine / steam turbine bearings from the vibration monitoring system via the API interface; and it acquires over 20 performance indicators such as heat rate and compressor efficiency from the performance calculation module. The data then undergoes 3... Outlier cleaning is performed according to criteria, and short-term missing data is filled in using linear interpolation. Time synchronization is achieved through a hardware clock. The operating condition classifier identifies the current operating condition as "basic load operation" (load > 95% of rated value, and change rate < 0.5% for 10 consecutive minutes) based on load commands and change rates. Fifty key features are extracted using feature extraction templates and then standardized using Z-scores to generate a feature vector set. .
[0051] Secondly, the risk diversion diagnosis module 42 starts a high-priority security guardian process to scan the feature vector set at a 100-millisecond cycle. If the safety red line rules such as speed and temperature are not triggered, it is determined to be a non-emergency state and enters the secondary risk parallel collaborative diagnosis channel.
[0052] Then, the parallel collaborative diagnosis module 43 initiates the data-driven anomaly detection and pattern matching diagnosis process and the knowledge-driven root cause reasoning and strategy optimization diagnosis process in parallel to implement dual collaborative diagnosis. The specific steps are as follows: The data-driven diagnosis unit calculates... The cosine similarity score between the case and the fault case database under basic load conditions is used to output the top two most similar cases: F2023-045 (abnormal increase in compressor inlet filter pressure difference, similarity 0.92) and F2023-112 (loose IGV linkage mechanism, similarity 0.87). Meanwhile, the DBSCAN clustering model calculates D=1.2× The result was determined to be abnormal, with an anomaly index of 65 and a confidence level of 0.90. The knowledge-driven diagnostic unit started from abnormal symptom nodes such as "increased exhaust temperature dispersion" and "decreased compressor efficiency", traversed the knowledge graph to perform causal reasoning, and the reinforcement learning agent combined with operation and maintenance experience to optimize the reasoning strategy, outputting the root cause hypothesis "carbon deposits in burner #3 and #6 nozzles", with a confidence level of 0.78. The reasoning path was "high exhaust temperature dispersion → uneven combustion → carbon deposits in fuel nozzles".
[0053] Subsequently, the dynamic fusion decision module 44 performs dynamic weight fusion and comprehensive judgment through the fusion center. The data noise level is 0.1, and the expert experience reliability coefficient is 0.9. The calculated results are as follows: =0.90×(1-0.1)=0.81, =0.78×0.9=0.702, after normalization and weighted voting, a comprehensive conclusion is formed: "Carbon buildup in the nozzles of gas turbine #3 and #6 burners"; according to the fault risk matrix calculation, the severity of the fault consequence S=3, the fault development speed L=4, and the risk value R=12, which is judged as medium risk; a structured report is generated and pushed to the operation and maintenance engineer. The report includes the handling suggestion: "Strengthen the monitoring of exhaust temperature dispersion and conduct borehole inspection on burners #3 and #6 during planned shutdowns".
[0054] Finally, the maintenance engineer planned to perform borehole probing on the burner after shutdown, confirming severe carbon buildup in nozzle #6. He then clicked "Confirm" on the diagnostic report interface and filled in the verification results. The self-learning evolution module 45 generated new cases from this diagnostic data and updated them to the fault case library. It increased the may_cause weight of "fuel nozzle carbon buildup → uneven combustion" in the knowledge graph, fine-tuned the DBSCAN clustering model threshold, and added positive rewards to the reinforcement learning agent's experience replay pool, completing model training optimization.
[0055] In summary, the embodiments disclosed herein have at least the following technical effects: Full-condition coverage without blind spots: Through adaptive condition identification and feature extraction, full-condition diagnosis is achieved, including cold start, load increase, basic load, and shutdown. The diagnostic coverage is greatly improved, and the problem of blind spots in dynamic condition diagnosis is solved.
[0056] Accurate fault risk classification and improved response speed: Emergency faults achieve millisecond-level response of ≤100 milliseconds through hard logic rules, while non-emergency faults enter in-depth analysis, effectively avoiding the flood of alarms that overwhelm dangerous signals, and the risk classification has extremely high accuracy.
[0057] Significantly improved diagnostic accuracy and root cause localization capability: Through data and knowledge-driven parallel collaborative diagnosis, combined with dynamic weight fusion, the false alarm rate of diagnosis is greatly reduced, the accuracy of root cause localization is greatly improved, and the diagnostic results have a visualized reasoning path, solving the problem of the lack of interpretability of the "black box" model.
[0058] Enhanced self-learning and evolution capabilities of the system: Through the closed loop of operation and maintenance feedback, the case library, knowledge graph, and diagnostic model are continuously optimized, and the ability to identify new faults is constantly improved with the accumulation of operation and maintenance data, eliminating the need for frequent manual model adjustments and reducing operation and maintenance costs.
[0059] Guided tiered operation and maintenance enhances unit operation safety: Tiered diagnostic reports provide targeted handling suggestions, enabling immediate handling of emergency faults and advance planning of maintenance for medium / low-risk faults, effectively reducing the number of unplanned unit shutdowns and improving unit operation stability and safety.
[0060] It is understood that the above embodiments are merely exemplary implementations used to illustrate the principles of the present invention, and the present invention is not limited thereto. For those skilled in the art, various modifications and improvements can be made without departing from the spirit and essence of the present invention, and these modifications and improvements are also considered to be within the scope of protection of the present invention.
Claims
1. A risk-classified method for fault diagnosis and root cause tracing of generator units under all operating conditions, characterized in that, Includes the following steps: S1: Collect all-dimensional raw data of the unit and perform data preprocessing. Based on the preprocessed data, perform intelligent identification of all operating conditions and extract and optimize the identified operating condition data to construct a feature vector set. The unit is composed of a gas turbine, a waste heat boiler, a steam turbine and multiple auxiliary machine systems coupled together. S2: Periodically scan the feature vector set to determine whether it triggers the safety red line rule. If the safety red line rule is triggered, output an emergency response command. If the safety red line rule is not triggered, transmit the feature vector set to the secondary risk parallel diagnosis channel. S3: Initiate data-driven anomaly detection and pattern matching diagnostic processes and knowledge-driven root cause reasoning and strategy optimization diagnostic processes in parallel to implement dual collaborative diagnosis and output diagnostic results including anomaly type hypothesis, root cause hypothesis and corresponding confidence level. S4: A hybrid decision-making architecture combining strong evidence priority and dynamic weighting is adopted to fuse the diagnostic results to obtain a comprehensive diagnostic conclusion and generate a structured hierarchical diagnostic report. S5: Collect operation and maintenance verification results to update the fault case library, optimize the knowledge graph, and fine-tune model parameters to achieve system self-evolution.
2. The risk-classified unit full-condition fault diagnosis and root cause tracing method according to claim 1, characterized in that, In step S1, the data preprocessing includes data cleaning and hard synchronization. The data cleaning includes: based on 3 Outlier detection and removal criteria, short-term missing data completion based on linear interpolation, and labeling and blanking of long-term missing data; The hard synchronization includes: achieving full data time synchronization through a hardware clock.
3. The risk-classified unit full-condition fault diagnosis and root cause tracing method according to claim 1, characterized in that, In step S1, the full-condition intelligent identification includes: using a condition classifier with the unit load command and the generator outlet circuit breaker status as the core basis, and combining the load change rate threshold to determine the cold start, load increase, basic load operation, and shutdown conditions. The load increase threshold is ≥1.5% of rated load / minute, and the basic load threshold is ≤0.5% of rated load / minute.
4. The risk-classified unit full-condition fault diagnosis and root cause tracing method according to claim 1, characterized in that, In step S1, the feature extraction and optimization of the identified working condition data includes: for different working conditions, key features are adaptively extracted using a preset feature extraction template, and a feature vector set is constructed after Z-score standardization.
5. The risk-classified unit full-condition fault diagnosis and root cause tracing method according to claim 1, characterized in that, In step S2, the safety red line rule is an IF-THEN hard logic rule related to unit safety, including at least one of unit overspeed, overtemperature and shutdown.
6. The risk-classified unit full-condition fault diagnosis and root cause tracing method according to claim 1, characterized in that, In step S3, The data-driven anomaly detection and pattern matching diagnosis process constructs a healthy cluster model through the DBSCAN clustering algorithm and uses cosine similarity calculation to achieve anomaly pattern matching. The knowledge-driven root cause reasoning and strategy optimization diagnostic process uses the multi-physics knowledge graph of the unit to perform root cause causal reasoning and uses a reinforcement learning agent to generate optimization strategies. The reinforcement learning agent uses the operation and maintenance verification results as a reward signal to iterate the strategy.
7. The risk-classified unit full-condition fault diagnosis and root cause tracing method according to claim 1, characterized in that, Step S4 specifically includes: Prioritize high-confidence diagnostic results with a confidence level > 0.95 and no conflicts. Otherwise, dynamically calculate the normalized weights of data-driven and knowledge-driven approaches based on the data noise level and the credibility coefficient of expert experience, and perform weighted voting fusion. At the same time, construct a risk matrix based on the severity of the fault consequences and the speed of fault development to quantify the risk level, and finally generate a structured hierarchical diagnostic report.
8. A risk-classified unit full-condition fault diagnosis and root cause tracing system, characterized in that, The system is used to implement the risk classification method for unit full-condition fault diagnosis and root cause tracing according to any one of claims 1 to 7. The system includes: a data perception preprocessing module, a risk diversion diagnosis module, a parallel collaborative diagnosis module, a dynamic fusion decision module, and a self-learning evolution module. The above modules are connected in a bidirectional communication manner through an industrial communication protocol. The data perception preprocessing module is used to perform data acquisition and preprocessing, intelligent identification of all working conditions, and extraction and optimization of working condition features based on the identification results, so as to generate a feature vector set; The risk diversion diagnosis module is used to periodically scan the feature vector set and perform safety red line rule judgment on it. If the judgment result is triggered, an emergency response instruction is generated and output. If the judgment result is not triggered, the feature vector set is transmitted to the secondary risk parallel diagnosis channel. The parallel collaborative diagnosis module is used to initiate data-driven anomaly detection and pattern matching diagnosis process and knowledge-driven root cause reasoning and strategy optimization diagnosis process in parallel to implement dual collaborative diagnosis and output diagnosis results including anomaly type hypothesis, root cause hypothesis and corresponding confidence level. The dynamic fusion decision module is used to generate a comprehensive diagnostic conclusion using a hybrid decision architecture that combines strong evidence priority with dynamic weight fusion, and to generate a structured hierarchical diagnostic report based on the comprehensive diagnostic conclusion. The self-learning evolution module forms a closed-loop feedback connection with the parallel collaborative diagnosis module and the dynamic fusion decision module. It is used to collect operation and maintenance verification results, and update the fault case library, optimize the knowledge graph and fine-tune the model parameters based on the operation and maintenance verification results to achieve the self-evolution of the system.
9. The risk-classified unit full-condition fault diagnosis and root cause tracing system according to claim 8, characterized in that, The parallel collaborative diagnostic module includes a data-driven diagnostic unit and a knowledge-driven diagnostic unit that are independent of each other. The data-driven diagnostic unit is equipped with a DBSCAN clustering algorithm program and a fault case library, while the knowledge-driven diagnostic unit is equipped with a knowledge graph management system and a reinforcement learning agent program.
10. The risk-classified unit full-condition fault diagnosis and root cause tracing system according to claim 8, characterized in that, The hardware carrier of the system is an industrial server, a data acquisition terminal, and a diagnostic terminal; the software carrier is an embedded program and industrial diagnostic system software; and the industrial communication protocol is at least one of OPC DA, OPC UA, and OPC API.
Citation Information
Patent Citations
Intelligent fault diagnosis and dynamic early warning system and method for SCADA (supervisory control and data acquisition) system
CN121857652A
Intelligent operation and maintenance method for environment monitoring station building based on multi-source data fusion and collaborative diagnosis
CN122048275A