Power generation equipment maintenance management method based on RCM and large and small model cooperation
By combining RCM with a collaborative architecture of big and small models, and utilizing multi-source data and knowledge graphs, intelligent predictive maintenance of power generation equipment has been achieved. This solves the problems of unreasonable resource allocation and inaccurate fault prediction in traditional maintenance management, and improves the operating efficiency and safety of the equipment.
Patent Information
- Application Number
- CN202511650786.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-12
- Publication Date
- 2026-02-10
AI Technical Summary
Traditional power generation equipment maintenance and management relies on periodic inspections and experience-based judgment, leading to unreasonable resource allocation and inaccurate fault prediction, which affects equipment operating efficiency and safety. There is an urgent need for an intelligent predictive maintenance method that integrates RCM and FMEA.
By collecting multi-source data through a monitoring platform, combining the RCM framework and failure mode impact analysis, using small models for data compression and key information extraction, introducing large models for in-depth analysis, and embedding a knowledge graph enhancement mechanism, structured maintenance work orders are generated to achieve intelligent fault prediction and decision-making.
It significantly improves the accuracy of power generation equipment fault early warning and the level of intelligence in maintenance decision-making, reduces unplanned downtime and resource waste, and improves equipment reliability and operating efficiency.
Smart Images

Figure CN121504431A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power equipment operation and maintenance management technology, and in particular to a power generation equipment maintenance management method based on the collaboration of RCM and large and small models. Background Technology
[0002] With the development of the modern power industry, power generation equipment such as boilers, steam turbines, and generators face increasingly complex operating environments and high-efficiency requirements. The high efficiency and reliability of this equipment are crucial for power production. However, traditional power generation equipment maintenance and management largely rely on periodic inspections and experience-based judgment. These methods often suffer from fixed cycles, unreasonable resource allocation, and inaccurate fault prediction, leading to over-maintenance or under-maintenance, which in turn affects the equipment's operational efficiency and safety. To address these issues, predictive maintenance has become key to modern power generation equipment operation and maintenance. By collecting equipment operating data, inspection records, and historical maintenance data, and combining this with advanced analytical methods for equipment health status assessment and fault early warning, it is possible to predict potential equipment faults in a timely manner. This allows for repairs before faults occur, avoiding unplanned downtime, reducing maintenance costs, and extending equipment lifespan.
[0003] Currently, Reliability Center Maintenance (RCM) and Failure Mode and Effects Analysis (FMEA) are widely used in equipment risk assessment and failure prediction, effectively quantifying the severity, probability of occurrence, and detectability of failures, guiding the determination of maintenance priorities. However, traditional FMEA methods rely on human experience and struggle to handle and analyze complex operating conditions and large amounts of real-time data. Therefore, combining advanced artificial intelligence technologies, machine learning models, and big data analytics to achieve dynamic analysis and real-time prediction has become a trend for improving the efficiency and accuracy of power generation equipment maintenance management.
[0004] Therefore, there is an urgent need for a new predictive maintenance method that integrates the advantages of RCM and FMEA, and combines dynamic analysis of multi-source data, intelligent reasoning of large models, knowledge graph enhancement, and optimization based on field feedback. This method can achieve the integration of efficient data processing and intelligent reasoning, significantly improve the accuracy of fault warnings for power generation equipment, reduce unplanned downtime and resource waste, and provide strong technical support for the safe, economical, and continuous operation of power equipment. Summary of the Invention
[0005] To address the aforementioned technical issues, this application provides a power generation equipment maintenance management method based on the collaboration of RCM and large / small models, which aims to improve the accuracy of power generation equipment fault prediction and the level of intelligence in maintenance decision-making.
[0006] Firstly, this application provides a power generation equipment maintenance management method based on the collaboration of RCM and size model, the method comprising: Step S1: Collect multi-source data of the power generation equipment in real time through the monitoring platform, including operating data, inspection texts and historical maintenance records; Step S2: Under the RCM framework, use the Failure Mode and Effects Analysis method to quantify the severity, probability of occurrence, and detectability of potential failure modes, and calculate the risk priority number. Step S3: Prioritize different failure modes based on the risk priority number, and establish a causal reasoning model for equipment power generation by combining the fault tree analysis method to identify the causes of failure of the power generation equipment. Step S4: Use a small model to compress prompt words and extract key information from the multi-source data to obtain semantic feature vectors, and introduce a large model fine-tuned by DPO to perform in-depth analysis on the semantic feature vectors to obtain the failure prediction probability of the power generation equipment. Step S5: Embed a knowledge graph retrieval enhancement mechanism during the large model reasoning process to dynamically call up equipment structure, typical fault modes and historical maintenance records to form a causal reasoning chain; Step S6: Automatically generate structured maintenance work orders based on the large model inference results, including fault causes, maintenance steps, priorities, and resource allocation.
[0007] Compared with the prior art, the beneficial effects of the present invention are at least as follows: This application presents a power equipment maintenance management method based on the collaboration of Reliability Center Model (RCM) and large and small models. By constructing a closed-loop system encompassing multi-source data fusion, intelligent risk assessment, fault prediction, and decision generation, it achieves a leapfrog upgrade in power equipment operation and maintenance management from traditional periodic inspections and reactive maintenance to digital and intelligent predictive maintenance. First, this application deeply integrates the theoretical framework of Reliability Center Model (RCM) with modern artificial intelligence technology. Through Failure Mode and Effects Analysis (FMEA), it systematically identifies and quantitatively assesses potential equipment faults, and utilizes risk priority numbers to provide a scientific basis for maintenance strategies. This effectively overcomes the problems of strong subjectivity and incomplete coverage caused by reliance on human experience in traditional operation and maintenance, significantly improving the accuracy and systematic nature of equipment risk assessment. Second, it adopts a large and small model collaborative architecture. The small model efficiently compresses prompts and extracts key information from multi-source data, significantly reducing data processing complexity. Simultaneously, it utilizes a large model, fine-tuned through direct preference optimization, for deep semantic analysis and cross-modal fusion. This ensures the real-time requirements of fault prediction while significantly improving the accuracy of fault prediction through the powerful reasoning capabilities of the large model, solving the technical challenge of balancing efficiency and accuracy with a single model.
[0008] This application further enhances the model's reasoning process by embedding a knowledge graph retrieval mechanism, dynamically integrating equipment structural knowledge, typical failure modes, and historical maintenance cases to construct a complete causal reasoning chain. This not only enhances the interpretability of the model's decisions but also organically integrates domain expertise with the data-driven model, making fault diagnosis results more reliable and aligned with engineering realities. Finally, based on the intelligent reasoning results, structured maintenance work orders are automatically generated, clearly defining the fault causes, maintenance steps, priorities, and resource allocation. This achieves an automated closed loop from fault prediction to maintenance execution, significantly improving the standardization and efficiency of maintenance operations. Attached Figure Description
[0009] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is a schematic diagram of an embodiment of the power generation equipment maintenance management method based on RCM and size model collaboration in this application. Figure 2 This is a diagram illustrating the collaborative architecture between the small and large models in the embodiments of this application; Figure 3 This is a flowchart illustrating the training and fine-tuning process based on a large model in an embodiment of this application. Figure 4 This is a flowchart of the failure mode and effect analysis in the embodiments of this application; Figure 5 This is a schematic diagram of the maintenance plan in an embodiment of this application. Detailed Implementation
[0011] This application provides a power generation equipment maintenance management method based on RCM and size model collaboration. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0012] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of the power generation equipment maintenance management method based on RCM and size model collaboration in this application includes: Step S1: Collect multi-source data of the power generation equipment in real time through the monitoring platform, including operating data, inspection texts and historical maintenance records.
[0013] Step S1 includes: The operational data includes temperature data, current data, voltage data, vibration data, and pressure data. The multi-source data of the power generation equipment is synchronized in time using a time alignment function, and the multi-source data is preprocessed using missing value imputation and anomaly removal algorithms.
[0014] Specifically, the monitoring platform collects multi-source data from the power generation equipment in real time. Specifically, sensors deployed on the equipment continuously collect operational data such as temperature, current, voltage, vibration, and pressure, which are transmitted to the monitoring platform via network. Simultaneously, it collects inspection texts entered by inspection personnel and historical maintenance records from the database. Due to timestamp differences in the multi-source data, a time alignment function is first used, for example, interpolation based on the earliest or latest timestamp, to unify all data onto the same timeline. Subsequently, the synchronized data undergoes preprocessing: missing values are filled using linear interpolation or the K-nearest neighbor algorithm; outliers are identified and removed using outlier detection methods or the isolated forest algorithm. The preprocessed, well-organized data is stored in the database for subsequent steps.
[0015] Step S2: Under the RCM framework, use the Failure Mode and Effects Analysis (FMEA) method to quantify the severity, probability of occurrence, and detectability of potential failure modes, and calculate the risk priority number.
[0016] Step S2 includes: Failure Mode and Effects Analysis (FMEA) is used to analyze multi-source data of power generation equipment, identify the consequences of failures recorded in historical maintenance logs, identify failure modes, and assign a severity score S1 to each failure mode. Based on the multi-source data of the power generation equipment, the frequency of occurrence of each failure mode within a preset operating cycle is calculated, and an occurrence probability score O is assigned to each failure mode. Based on the sensor configuration, inspection procedures, and diagnostic technologies of the current monitoring platform, the ease of identifying failure modes before or during failures is assessed, and a detectability score D is assigned to each failure mode. The Risk Priority Number (RPN) for each failure mode is calculated based on the first formula: .
[0017] Specifically, within the RCM (Reliability Center Maintenance) framework, Failure Mode and Effects Analysis (FMEA) is used to rank the potential failure modes of power generation equipment. Specifically, based on equipment design drawings, functional principles, and historical operation and maintenance experience, all possible potential failure modes are pre-deduced, such as generator bearing overheating, insulation aging, and cooling water leakage. Subsequently, each potential failure mode is quantitatively assessed using three indicators: Severity: This assesses the ultimate impact of the failure mode on equipment safety, performance, or the environment, scored from 1 to 10. For example, a failure leading to equipment explosion or complete shutdown is scored 10, while a failure causing only minor performance fluctuations is scored 1. Probability of Occurrence: Based on historical operating data and maintenance records, the frequency of occurrence of the failure mode within a specific operating cycle is statistically estimated, scored from 1 to 10. For example, a failure that does not occur year-round is scored 1, while a failure that occurs monthly is scored 10. Detectability: This assesses the ease with which the failure mode can be identified promptly and accurately when or before it occurs, using the current monitoring platform's sensors, inspection procedures, and diagnostic technologies, scored from 1 to 10. For example, a failure with real-time alarms is scored 1, while a failure that requires shutdown and disassembly to be detected is scored 10. Finally, the risk priority number for each failure mode is calculated using the first formula. A higher RPN value indicates a greater overall risk for that failure mode, and therefore a higher maintenance priority. In this way, the risk priority number provides a scientific basis for prioritizing equipment maintenance, thereby helping to make effective maintenance decisions and improve equipment reliability and maintenance efficiency.
[0018] Step S3: Prioritize different failure modes based on risk priority numbers, and establish a causal reasoning model for equipment power generation by combining fault tree analysis method to identify the causes of equipment failure.
[0019] Step S3 includes: Fault modes with a risk priority number greater than a preset threshold are defined as high-risk fault modes; a fault tree is constructed based on high-risk fault modes, and a fault event set is generated based on the set of all basic fault causes. By combining the failure path combination function of Boolean logic, the overall probability of occurrence of high-risk failure modes under multiple failure causes is calculated based on the second formula. The second formula is: ,in, This is a failure path combination function based on Boolean logic. For each basic cause of failure The probability of occurrence; by quantitatively analyzing the overall probability of occurrence and the contribution of each basic fault cause to the fault tree logic structure, the cause of power generation equipment failure is determined.
[0020] Specifically, different failure modes are prioritized based on their risk priority number (RPN). All calculated RPN values are compared to a preset threshold. If a failure mode's RPN value exceeds this threshold, it is defined as a high-risk failure mode. This effectively prioritizes the most potentially risky failure modes, ensuring that those with the greatest impact on equipment operation and safety are addressed first in maintenance management. For equipment failures defined as high-risk failure modes, a fault tree is constructed. This construction first identifies and lists all basic failure causes that could lead to the high-risk failure mode, such as damage to equipment components, changes in the working environment, and operational errors. Based on these causes, a set of failure events is generated, denoted as: ,in, This indicates a possible cause of the failure.
[0021] Next, we will analyze the relationships between various fault causes by combining Boolean logic failure path combination functions. This function is used to simulate the combined effects of different failure causes on high-risk failure modes. For example, when multiple failure causes occur simultaneously, their effects may be cumulative, leading to eventual system failure. The function assigns the probability of occurrence of each basic failure cause. As input, the overall probability of occurrence of the failure mode is calculated. This probability represents the combined likelihood of a high-risk failure occurring in the system under various failure causes. This method comprehensively considers the overall impact of all failure causes on equipment failure modes, thus calculating the overall failure probability of the system. Finally, quantitative analysis, combined with the logical structure of the fault tree, is used to further analyze the contribution of each basic failure cause in the fault tree. By analyzing the contribution, it can be identified which failure causes have the greatest impact when causing equipment failure, thereby helping to determine the cause of power generation equipment failure. For example, if a basic failure cause has a high probability of occurrence and a large contribution to the overall failure mode, then this failure cause is likely the main cause of equipment failure. Through the above process, the root causes of equipment failure can be comprehensively identified and analyzed, providing accurate basis for subsequent maintenance decisions, ensuring that high-risk failure modes can be addressed in a timely manner, thereby effectively reducing the equipment failure rate and improving equipment reliability and safety.
[0022] Furthermore, constructing a fault tree based on high-risk fault modes includes setting each high-risk fault mode as the top event of the fault tree and constructing the fault tree with the top event as the root node.
[0023] Specifically, high-risk failure modes identified through risk assessment are set as the top events and root nodes of the fault tree. Then, a top-down deductive analysis method is used, starting from the top event and decomposing layer by layer using logical symbols such as AND and OR gates to trace all direct causes that could lead to the previous level event. This deductive process continues until all branches are decomposed to indivisible basic failure causes, i.e., bottom events. For example, when "generator bearing overheating shutdown" is identified as a high-risk failure mode (top event), analysis reveals that it is caused by the simultaneous occurrence of "cooling failure" and "lubrication failure" (connected by an AND gate). "Cooling failure" can be further decomposed into "cooling water pump failure" or "cooling pipe blockage" (connected by an OR gate). These indivisible component failures constitute the bottom events. Finally, the set of all bottom events is defined as the fault event set. This allows for the construction of a logically rigorous fault tree model that can be used for subsequent quantitative analysis.
[0024] Step S4: Use a small model to compress prompt words and extract key information from multi-source data to obtain semantic feature vectors. Then, introduce a large model fine-tuned by DPO to perform in-depth analysis on the semantic feature vectors and obtain the failure prediction probability of power generation equipment.
[0025] Step S4 includes: The small model uses an information compression function to compress prompt words and extract key information from multi-source data, resulting in a low-dimensional semantic feature vector. ; Use DPO to fine-tune the large model to obtain the optimized parameters of the large model; convert the semantic feature vectors generated by the small model into... With real-time monitoring data The input is fed into the optimized large model for cross-modal fusion and deep analysis. The large model outputs the predicted probability of power generation equipment failure within a preset time interval through a prediction function. , ,in, For the prediction function, For the preset time interval, These are parameters for a large model.
[0026] Specifically, small models can employ lightweight pre-trained models such as BERT-mini, whose encoder layer maps lengthy inspection texts and high-dimensional operational data into low-dimensional semantic feature vectors. This reduces the dimensionality and complexity of the data while retaining the information most relevant to equipment failure prediction. Next, the larger model receives the semantic feature vectors output by the smaller model. and real-time device data from the monitoring platform. Building upon this, the large model is fine-tuned using Direct Preference Optimization (DPO). DPO optimizes the large model's parameters by combining expert feedback and field data, enabling it to better adapt to the task requirements of specific application scenarios. Specifically, the large model parameters... The update is performed by minimizing an optimization objective function, which enables the large model to achieve more accurate results in predicting equipment failures. The detailed process will be explained later.
[0027] When the semantic feature vector generated by the small model and real-time monitoring data After being input into the large model fine-tuned by DPO, the model undergoes cross-modal fusion and deep analysis. The large model then uses its prediction function... Output device failure prediction probability , It is the power generation equipment in the future at a preset time interval. The probability of a failure occurring within the system reflects the risk of equipment failure at a future point in time, helping decision-makers to take appropriate maintenance or preventative measures in advance. In this way, a smaller model efficiently extracts key information and reduces input complexity, while a larger model performs deep learning and inference, providing accurate failure prediction probabilities. The DPO fine-tuning mechanism ensures that the large model can be optimized for specific equipment and actual operating environments, greatly improving the accuracy and reliability of failure prediction. This collaborative working mode provides strong technical support for predictive maintenance of power generation equipment, helping to reduce unplanned downtime and improve equipment operating efficiency.
[0028] like Figure 2The diagram shows the collaborative architecture of the small and large models. First, the small model processes the initial input prompts, which come from fault reports or field data. For example, a prompt might be: "You are a professional decision-making consultant for the safe operation of flexible thermal power generating units, responsible for comprehensively analyzing unit operating data, fault characteristics, and maintenance records to generate scientifically sound auxiliary decision-making suggestions." The small model then performs compression and key information extraction tasks based on the prompts. For example, for inspection text, it uses a lightweight natural language processing model to extract keywords and semantically vectorize, compressing the original long text into a low-dimensional feature vector containing core information. For operating data, it uses time-series feature extraction and redundant parameter filtering to generate a set of key indicators representing the equipment's operating status. In this way, the small model can significantly reduce the input size while ensuring information integrity, reducing the computational power consumption and latency of the large model, thereby improving overall inference efficiency. Building upon this foundation, the large model performs in-depth analysis by processing compressed information and real-time monitoring data obtained from the small model, outputting more accurate fault prediction results. During the fault prediction process, the large model combines historical and real-time data through cross-modal fusion to enhance prediction accuracy. It identifies potential equipment faults through causal chain reasoning of fault modes, thereby generating decision-making recommendations. The small model inputs compressed prompts into the large model, which processes them and generates model feedback—the fault prediction results and recommendations. This includes information such as possible causes of equipment failure and risk assessments, helping maintenance personnel make informed decisions.
[0029] Furthermore, the large model is fine-tuned using DPO to obtain the optimized parameters of the large model, including: Set the target function: ,in, This is a triplet of prompt words, preferred responses, and poor responses constructed based on on-site feedback and expert knowledge. For hyperparameters, It is the Sigmoid activation function. Let log be the expectation operator, and let log be the natural logarithm. To directly optimize the loss function, DPO fine-tuning is performed on the large model by minimizing the objective function.
[0030] Specifically, in the process of further optimizing the large model, the Direct Preference Optimization (DPO) method is used to fine-tune the large model. The core of DPO fine-tuning lies in using an objective function to guide the optimization of the large model. The objective function is: By minimizing this loss function, the parameters of the large model are optimized to improve the accuracy of the model in fault prediction. In the objective function... Where x represents the device's input data, such as operational data, monitoring data, etc. The preferred response represents the expected performance of the device under normal operating conditions. Poor response represents the performance of the equipment under faulty or abnormal conditions. These triples, constructed from field feedback and expert knowledge, are used to provide training data for the model. The optimal response represents the output that the equipment should have under normal conditions, while the poor response represents the output when the equipment is faulty or abnormal. By inputting these triples into the model, the model can learn the performance of the equipment under different operating conditions and make correct fault predictions. It is a hyperparameter that controls the weighting of the difference between optimal and poor responses in the loss function. This is the Sigmoid activation function, used to map input values to a range of 0 to 1, ensuring that the output probability values have practical meaning. By minimizing this objective function, the model's parameters are updated, thereby optimizing the large model and making it more accurate in predicting the probability of equipment failure within a preset time interval. Specifically, DPO fine-tuning adjusts the parameters of the large model so that it can more accurately output failure prediction probabilities when faced with complex equipment operating data, improving the accuracy and responsiveness of failure prediction. This fine-tuning process, by combining actual feedback from the field and expert knowledge, not only improves the model's adaptability to specific equipment and environments but also enables the large model to more accurately predict equipment failure risks, further enhancing the predictive maintenance capabilities of the equipment.
[0031] like Figure 3 The diagram shows the training and fine-tuning flowchart based on a large model. It includes two main modules: training the large language model and freezing the large language model, and utilizes the concepts of selection and rejection scores. Training the large language model: During training, the model first receives selection and rejection cues, which guide it to make specific predictions. The model learns specific patterns in the input data, continuously adjusting its parameters. The model makes decisions based on selection and rejection scores, generating corresponding outputs, and is then trained and optimized. After training is complete, the model's parameters are fine-tuned through a weight update mechanism to better adapt to the target task. These adjusted weights directly affect the model's responsiveness to selection and rejection scores, ultimately improving its performance. Once the model has completed training and its performance has been optimized through weight updates, the system enters the model freezing phase. In this phase, the model's weights no longer change, and the fixed parameters are used for subsequent inference and decision-making. During model training, a loss function is used to measure the difference between the model's predicted output and the actual output. The loss function is expressed by the formula: The training error of the model is calculated, and the goal of the loss function is to minimize the difference between the predicted and actual results, thereby optimizing the model's performance. Through selection and rejection scores, the model can evaluate the quality of the generated results according to the needs of the task. The selection score represents the credibility of the output chosen by the model, while the rejection score represents the degree to which the model excludes certain outputs. Based on these scores, the model will make more reasonable choices, improving prediction accuracy.
[0032] Step S5: Embed a knowledge graph retrieval enhancement mechanism during the large model reasoning process to dynamically call up equipment structure, typical fault modes and historical maintenance records to form a causal reasoning chain.
[0033] Step S5 includes: Building knowledge graphs Where h represents the component entity of the power generation equipment or the identified typical failure mode, r represents the structural relationship between components and the causal relationship of the failure, and t represents the corresponding maintenance measures; during the failure reasoning process of the large model, a retrieval enhancement generation module is embedded. Based on the context of the large model's reasoning, the retrieval enhancement generation module retrieves enhanced knowledge subgraphs related to the current state and potential failures of the power generation equipment from the knowledge graph G. The fusion function integrates the enhanced knowledge subgraph, the semantic feature vector generated by the small model, and real-time monitoring data to generate the final enhanced input Z. The large model then uses this enhanced input to perform causal chain reasoning on potential faults in power generation equipment. The fusion function is as follows: .
[0034] Specifically, by embedding a knowledge graph retrieval enhancement mechanism during the large model's reasoning process, the model's fault reasoning capability is further improved. First, a knowledge graph is constructed. Here, 'h' represents the physical components of the power generation equipment or identified typical failure modes, such as boilers, generators, sensors, pressure valves, etc., or failure modes such as overheating, overload, and abnormal vibration; 'r' represents the structural relationships between components or the causal relationships of failures, such as the structural relationship between a boiler and steam pipes, or the causal relationship between boiler overheating and water pump failure; and 't' represents the maintenance measures, i.e., the repair operations taken when a specific failure mode or equipment component fails, such as replacing sensors, cleaning filters, or adjusting the temperature control system. By constructing a knowledge graph G containing equipment components, failure modes, and maintenance measures, rich domain knowledge support is provided for equipment failure reasoning. During the failure reasoning process of the large model, the retrieval enhancement generation module retrieves enhanced knowledge subgraphs related to the current equipment state and potential failures from the constructed knowledge graph G based on the context of the large model's reasoning. This enhanced knowledge subgraph contains information on equipment components, failure modes, and their relationships that are closely related to the current operating status of the equipment, providing additional background knowledge for the large model.
[0035] Then use the fusion function: The semantic feature vectors generated from the small model and real-time monitoring data and the retrieved enhanced knowledge subgraph Perform fusion, fusion function This algorithm is specifically designed for cross-modal information fusion. It combines different types of data, such as numerical monitoring data, text-based inspection data, and structured knowledge graph information, to generate a final enhanced input Z. This enhanced input Z is fed into a large model for fault reasoning, analyzing the causal chains of potential power generation equipment faults, and identifying the root causes of equipment failures. By fusing the structural information and causal relationships provided by the knowledge graph, the large model improves the accuracy and interpretability of reasoning, especially when facing complex fault modes and multi-source data, enabling it to better simulate the paths and mechanisms of fault occurrence. Ultimately, using the causal reasoning results of the large model, a detailed diagnosis of potential equipment faults can be derived, helping maintenance personnel identify the causes of faults and take timely and effective maintenance measures. In this way, domain knowledge such as equipment component relationships, fault modes, and maintenance measures are effectively combined with a data-driven deep learning model, improving the accuracy, professionalism, and interpretability of fault reasoning.
[0036] like Figure 4 The diagram shows the Failure Mode and Effects Analysis (FMEA) flowchart. This process, aided by a large model and FMEA analysis, first collects equipment information and performs structural and functional analysis to identify potential failure modes and calculates the Risk Priority Number (RPN) to determine maintenance priorities. By analyzing the probability and detectability of failure modes and combining environmental assessments, the system optimizes maintenance decisions and generates structured maintenance work orders. Finally, maintenance improvements are made based on feedback. This process improves the accuracy of failure prediction and ensures the efficient execution of maintenance work.
[0037] Step S6: Automatically generate structured maintenance work orders based on the large model inference results, including fault causes, maintenance steps, priorities, and resource allocation.
[0038] Step S6 includes: Based on the results obtained by knowledge graph-enhanced reasoning of the large model, a structured predictive maintenance work order is generated. The maintenance work order is represented as: O={C, U, S, R}, where C is the cause of the fault, U is the maintenance step, P is the priority, and R is the resource allocation.
[0039] Specifically, based on the inference results of the large model, structured predictive maintenance work orders are automatically generated. These work orders, derived from the large model's knowledge graph-enhanced inference, detail the cause of the fault, maintenance steps, priorities, and resource allocation, enabling maintenance personnel to perform maintenance work efficiently and accurately. More specifically, after the large model obtains the causal chain of equipment faults through knowledge graph-enhanced inference, the system automatically analyzes the inference results and generates structured maintenance work orders based on different fault modes. The format of a maintenance work order is defined as O={C, U, S, R}, where C represents the cause of the fault. Based on the reasoning results of a large model, the system will list in detail the root causes of the equipment failure. For example, if the fault mode is "boiler overheating," the cause might be "temperature control sensor failure" or "water pump malfunction." The system automatically identifies and fills in the fault cause through fault tree analysis and knowledge graph reasoning. S represents the maintenance steps. Based on the fault cause, the system will automatically generate a series of specific maintenance steps. Each maintenance step is an optimized and directly executable operation, such as "check the temperature control sensor connection," "replace the damaged water pump component," and "calibrate the temperature control system." The maintenance steps are listed sequentially and are customized operations for specific equipment and fault modes. The system assigns a priority to each work order based on the severity and risk of the fault. High-risk fault modes are given higher priority to ensure that maintenance personnel handle the most urgent and important faults first. For example, high-risk faults involving equipment downtime will be marked as "urgent," while faults with less impact may be marked as "routine." The system also assigns a resource allocation: the work order includes the required resource allocation. The system will automatically determine the required maintenance tools, equipment, materials, and personnel based on the fault type and maintenance steps. For example, "replacing a temperature control sensor" may require specific tools and professional technicians, while "cleaning a filter" may only require regular tools and ordinary maintenance personnel. The system will automatically allocate the required resources to the work order based on the resource library.
[0040] The generated repair work orders are structured data, such as Figure 5 The diagram illustrates a maintenance plan, including basic information about the maintenance task, specific content, personnel and tool configuration, and estimated maintenance time. The plan clearly defines the responsibilities and arrangements for each step, ensuring the smooth progress of the maintenance process. Its accuracy and feasibility are confirmed through signatures. The system can print and archive work orders and provide them to the maintenance team. Maintenance personnel can then follow the maintenance steps step-by-step according to the work order, ensuring the orderly progress of the maintenance work. Furthermore, the priority and resource allocation of work orders help maintenance managers rationally arrange maintenance tasks, avoiding resource waste or delays and ensuring that power generation equipment can be restored to normal operation as quickly as possible. Ultimately, all maintenance activities are tracked based on work orders, and the system can also perform closed-loop optimization based on feedback data, further improving the accuracy and efficiency of maintenance strategies.
[0041] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0042] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0043] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for maintenance management of power generation equipment based on the collaboration of RCM and size model, characterized in that, The method includes: Step S1: Collect multi-source data of the power generation equipment in real time through the monitoring platform, including operating data, inspection texts and historical maintenance records; Step S2: Under the RCM framework, use the Failure Mode and Effects Analysis method to quantify the severity, probability of occurrence, and detectability of potential failure modes, and calculate the risk priority number. Step S3: Prioritize different failure modes based on the risk priority number, and establish a causal reasoning model for equipment power generation by combining the fault tree analysis method to identify the causes of failure of the power generation equipment. Step S4: Use a small model to compress prompt words and extract key information from the multi-source data to obtain semantic feature vectors, and introduce a large model fine-tuned by DPO to perform in-depth analysis on the semantic feature vectors to obtain the failure prediction probability of the power generation equipment. Step S5: Embed a knowledge graph retrieval enhancement mechanism during the large model reasoning process to dynamically call up equipment structure, typical fault modes and historical maintenance records to form a causal reasoning chain; Step S6: Automatically generate structured maintenance work orders based on the large model inference results, including fault causes, maintenance steps, priorities, and resource allocation.
2. The method according to claim 1, characterized in that, Step S1 includes: The operating data includes temperature data, current data, voltage data, vibration data, and pressure data. The multi-source data of the power generation equipment is synchronized in time using a time alignment function, and the multi-source data is preprocessed using missing value imputation and anomaly removal algorithms.
3. The method according to claim 1, characterized in that, Step S2 includes: The multi-source data of the power generation equipment is analyzed using the Failure Mode and Effects Analysis (FEMA) method to identify the consequences of failures recorded in the historical maintenance records, identify the failure modes, and assign a severity score S1 to each failure mode. Based on the multi-source data of the power generation equipment, the occurrence frequency of each fault mode within a preset operating cycle is calculated, and each fault mode is assigned an occurrence probability score O. Based on the current monitoring platform's sensor configuration, inspection procedures, and diagnostic technologies, assess the ease of identifying fault modes before or during a fault, and assign a detectability score D to each fault mode. The Risk Priority Number (RPN) for each failure mode is calculated based on the first formula, which is: .
4. The method according to claim 1, characterized in that, Step S3 includes: Fault modes with a risk priority number greater than a preset threshold are defined as high-risk fault modes; A fault tree is constructed based on the high-risk failure modes, and a set of failure events is generated based on the set of all basic failure causes. And, combined with the failure path combination function of Boolean logic, the overall probability of occurrence of the high-risk failure mode under multiple failure causes is calculated based on the second formula. The second formula is: ,in, This is a failure path combination function based on Boolean logic. For each basic cause of failure The probability of occurrence; By quantitatively analyzing the overall probability of occurrence and the contribution of each basic fault cause to the fault tree logic structure, the causes of power generation equipment failures are determined.
5. The method according to claim 4, characterized in that, Constructing a fault tree based on the aforementioned high-risk failure modes includes: Each high-risk failure mode is set as the top event of the fault tree, and the fault tree is constructed with the top event as the root node.
6. The method according to claim 1, characterized in that, Step S4 includes: The small model performs cue word compression and key information extraction on the multi-source data using an information compression function to obtain a low-dimensional semantic feature vector. ; The DPO is used to fine-tune the large model and obtain the optimized parameters of the large model. The semantic feature vector generated by the small model With real-time monitoring data The input is fed into the optimized large model for cross-modal fusion and deep analysis. The large model outputs the predicted probability of power generation equipment failure within a preset time interval through a prediction function. , ,in, For the prediction function, For the preset time interval, These are parameters for a large model.
7. The method according to claim 6, characterized in that, Using DPO to fine-tune the large model, the optimized parameters of the large model include: Set the target function: ,in, This is a triplet of prompt words, preferred responses, and poor responses constructed based on on-site feedback and expert knowledge. For hyperparameters, It is the Sigmoid activation function. For expectation operator, It is the natural logarithm. To directly optimize the loss function, the large model is fine-tuned by minimizing the objective function.
8. The method according to claim 1, characterized in that, Step S5 includes: Building knowledge graphs Where h represents the component entity of the power generation equipment or the identified typical failure mode, r represents the structural relationship between components and the causal relationship of the failure, and t represents the corresponding maintenance measures. During the fault reasoning process of the large model, a retrieval enhancement generation module is embedded. This module retrieves enhanced knowledge subgraphs related to the current state and potential faults of the power generation equipment from the knowledge graph G based on the context of the large model's reasoning. ; The enhanced knowledge subgraph, the semantic feature vector generated by the small model, and the real-time monitoring data are fused together using a fusion function to generate the final enhanced input Z. Based on the enhanced input, the large model realizes causal chain reasoning for potential faults in power generation equipment.
9. The method according to claim 8, characterized in that, The fusion function is: .
10. The method according to claim 1, characterized in that, Step S6 includes: Based on the results obtained by knowledge graph-enhanced reasoning of the large model, a structured predictive maintenance work order is generated. The maintenance work order is represented as: O={C, U, S, R}, where C is the cause of the fault, U is the maintenance step, P is the priority, and R is the resource allocation.
Citation Information
Patent Citations
Platform door intelligent operation and maintenance and health management method and system based on knowledge graph
CN118313811A
Highway traffic incident detection method and system based on big and small model collaboration
CN120071625A
Multi-modal tampering information detection method based on cooperation of large model and small model
CN120354950A
Power plant maintenance strategy making method based on RCM and related device
CN120781220A