Method for maintenance of a fleet of aircraft derivative gas turbines based on reinforcement learning

By using reinforcement learning-based methods to hierarchically model and optimize maintenance decisions for aero-derivative gas turbine clusters, difficulties in the operation and maintenance process were resolved, resulting in reduced maintenance costs and improved equipment safety.

CN117151674BActive Publication Date: 2026-04-14XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-14
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

During the operation and maintenance of aero-derivative gas turbine equipment, monitoring of operating status, fault diagnosis, and operation and maintenance decision-making are difficult, resulting in large maintenance delays, increasing operation and maintenance risks and costs. Intelligent maintenance decision-making is urgently needed to reduce maintenance costs and improve equipment safety.

Method used

A reinforcement learning-based approach is adopted, dividing the cluster decision-making structure into individual, system, and resource layers. Combining the Weibull failure rate model and the two-parameter exponential degradation model, the maintenance decision is optimized through a deep Q-network algorithm, taking into account the maintenance quantity, type, and cost, to form a comprehensive failure rate model, which is then optimized within the reinforcement learning framework.

Benefits of technology

It effectively reduces maintenance costs, improves equipment security, avoids the problems caused by environmental complexity in traditional optimization algorithms, and realizes intelligent decision-making for cluster maintenance and optimization of resource pressure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117151674B_ABST
    Figure CN117151674B_ABST
Patent Text Reader

Abstract

The disclosure discloses a maintenance method for aero derivative gas turbine cluster based on reinforcement learning, which effectively reduces maintenance cost and improves equipment safety, which first divides the whole cluster decision structure into individual layer I, system layer Sys and resource layer R; based on the Weibull failure rate model, the service age delay factor and the failure rate increasing factor are introduced to reflect the influence of maintenance frequency on the working performance of the component; the effective service age is introduced and the service age delay factor is replaced to reflect the change of failure rate under the effective service age; according to the research on the degradation state evaluation based on multi-source information fusion, the actual aero derivative gas turbine degradation index result is obtained to establish a double-parameter exponential degradation model, and the failure rate modeling is combined with the system layer multi-dimensional state space S; for different maintenance activity types, the system layer action is divided into: preventive maintenance and no action a N Two kinds; the action cost and the resource limit cost jointly constitute the cost function, and the expected value is formed; finally, the simulation of the cluster equipment is completed in the reinforcement learning framework.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure pertains to the field of maintenance strategy optimization for aero-derivative gas turbine clusters, and specifically to a maintenance method for aero-derivative gas turbine clusters based on reinforcement learning. Background Technology

[0002] With the increasing number and scale of aero-derivative gas turbine equipment, its safety has become a major concern. During service, aero-derivative gas turbine failures occur frequently, seriously threatening personnel safety. Therefore, a reasonable operation and maintenance management scheme is urgently needed to provide timely maintenance recommendations to guide the safe and stable operation of the equipment. However, the internal structure, operating environment, and reliability requirements of aero-derivative gas turbines make operational status monitoring, fault diagnosis, and maintenance decision-making difficult, leading to significant maintenance delays and increased risks and costs. Therefore, a method is needed to address these issues. Research is needed on a method with significant application potential to provide intelligent maintenance decisions for aero-derivative gas turbine cluster systems, effectively reducing maintenance costs and improving equipment safety.

[0003] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure and may therefore contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0004] To address the aforementioned technical problems, this disclosure reveals a reinforcement learning-based maintenance method for aero-derivative gas turbine clusters, characterized by the following steps:

[0005] S1: Divide the entire cluster decision-making structure into individual layer I, system layer Sys, and resource layer R:

[0006] S2: Based on the Weibull failure rate model, an age delay factor and a failure rate increase factor are introduced to reflect the impact of maintenance frequency on the performance of a certain device in the entire cluster, where the device is a certain aero-derivative gas turbine.

[0007]

[0008] In the formula: i represents the number of repairs, which is a natural number; a i b is the service life delay factor. i The failure rate increasing factor, This represents the failure rate level during the i-th maintenance. Similarly, we can obtain... Where t represents the service time, β is the characteristic lifetime parameter, and ρ is the shape parameter;

[0009] S3: Introducing effective service life and replacing the service life delay factor to reflect the change in failure rate under effective service life, where W is the effective service life after the i-th maintenance. [i] The calculation is as follows:

[0010]

[0011] In the formula: Let C be the cost of the j-th repair, where j ranges from 1 to i. P To replace input costs, α is the adjustment parameter for input costs, γ is the time adjustment parameter, i is the same as in formula (1), and τ i This represents the time interval of the i-th maintenance activity;

[0012] Then the overall failure rate function after the i-th maintenance action is performed. for:

[0013]

[0014] In the formula, N represents the total number of maintenance activities performed. The efficiency level of who is not performing maintenance activities;

[0015] S4: To fully simulate the actual degradation process, a two-parameter exponential model containing a random error term is used to model the degradation:

[0016]

[0017] In the formula, γ(t) represents the degradation state at service time t, φ is a constant term, and θ indicates that ln(θ) follows a mean of μ0 and a variance of φ. The random variable is normally distributed, with ε(t) being the random error term, and has a mean of 0 and a variance of σ. 2 The normal distribution is given by η, which follows a mean of μ1 and a variance of . A normally distributed random variable;

[0018] Furthermore, for all i different devices in the entire cluster, the multidimensional state space S of the system layer is modeled using the integrated failure rate function in step S3:

[0019]

[0020] In the formula, i represents different devices, β i Let ρ be the characteristic lifetime parameter of the i-th device. i Let l be the shape parameter of the i-th device; w (β i , ρ i ,t) represents the overall failure rate function of the i-th device. Let γ(t) represent the degradation status of the i-th device with service time t, where Let represent the parameter set consisting of all parameters other than t in γ(t) corresponding to the i-th device;

[0021] S5: System-level actions are categorized into preventative maintenance based on different maintenance activity types. and no action a N There are two types of actions, both of which require reinforcement learning to make decisions, while the final maintenance action is component replacement. No decision-making is required;

[0022] S6: Combine motion costs and resource constraint costs into a cost function, and form the expected value;

[0023] S7: Simulate the cluster devices within the reinforcement learning framework, and optimize using the deep Q-network algorithm with the goal of minimizing the cost function.

[0024] Preferred,

[0025] In step S1, the individual layer I includes multiple individual devices e i The state of the important components inside each piece of equipment determines the overall performance of the equipment. When the failure rate of an important component is too high, leading to increased operational risks or serious malfunctions, the entire equipment will be unable to perform its tasks. The important internal components include components such as the combustion chamber of the aero-derivative gas turbine and the compressor. The value of i is the number of equipment.

[0026] Preferred,

[0027] The system layer Sys comprises r individual layers, Sys = (I1, I2, ..., I... r Each individual layer performs its own tasks while also being scheduled by the cluster center.

[0028] Preferred,

[0029] The resource layer R includes various operation and maintenance resources, such as monitoring and performance evaluation equipment for the aero-derivative gas turbine prediction and health management system, the number of maintenance support points, and the maximum number of workers at each support point. These operation and maintenance resources are often not unlimited, which will limit the execution of maintenance decisions.

[0030] Preferred,

[0031] In step S2, different model parameters are set for different devices in the Weibull failure rate model to distinguish their failure rates and reflect the impact of maintenance frequency on the failure rate. Different settings are made for ρ and β to distinguish different devices.

[0032] Preferred,

[0033] In step S3, the actual service life of the equipment throughout its entire life cycle is often not its in-situ working time, but rather the effective service life W of the components. [i]This reflects the impact of the type of maintenance activity and the level of cost invested on the equipment recovery level.

[0034] Preferred,

[0035] In step S4, the two-parameter exponential degradation process model containing a random error term distinguishes the degree of degradation between different devices by setting different model parameters. All parameters can be set differently.

[0036] Preferred,

[0037] In step S5, system-level maintenance actions are divided into preventative maintenance. and no action a N There are two types of maintenance actions, both requiring reinforcement learning for decision-making. One involves selecting an action as the maintenance behavior at a fixed monitoring time interval. If a fault occurs between two intervals, replacement maintenance is required. Both preventative and replacement maintenance incur economic and time costs and impact the reinforcement learning environment. The final maintenance action is component replacement. Rc No decision-making is required.

[0038] Preferred,

[0039] In step S6, the cost functions are divided into two categories:

[0040] Action cost:

[0041] Assume that during device operation, devices e in the individual layer within the device cluster are evaluated at time intervals of ΔT (ΔT > 0). i Perform status monitoring and during operation and maintenance, for a duration T Pc And the cost of a single investment, c Pc To perform preventative maintenance, the labor and material costs required for preventative maintenance are included in the single-use cost. The failure risk cost of critical components during condition monitoring is assumed to be v. Fc The unit time delay cost is c Sc MTTR represents the mean time to repair critical components, and C represents the cost of performing a single action at the individual level. P,i Then it can be expressed as:

[0042]

[0043] In the formula: C P,i Let i represent the cost function for an individual i to perform preventative maintenance actions;

[0044] The system-level cost C under the current action set execution. sys The calculation is as follows:

[0045]

[0046] In the formula: The number of devices that require preventative maintenance at time t;

[0047] Resource-constrained costs:

[0048] To avoid increased logistical maintenance pressure and reduced cluster availability due to simultaneous maintenance of multiple devices, system-level optimization is needed to limit the number of devices requiring maintenance at the same time. This can be achieved by reducing the overlap of device maintenance times.

[0049]

[0050] In the formula: d i This represents the maintenance duration of device number i, satisfying d. min >ΔT,t i t k These represent the maintenance times for devices numbered i and k, respectively, with b ranging from 1 to d. i Ol ik This indicates the overlap in maintenance times for devices numbered i and k.

[0051] Based on this, calculate the current system-level action set a. (t) The system-level maintenance overlap cost under execution, considering the cost of maintaining the overlap between two adjacent action sets {a (t-1) a (t) The execution process is described, and assuming M represents the total number of devices involved in maintenance under two action sets at the system layer, then the overall overlapping cost of the system layer under the current action set execution is calculated. for:

[0052]

[0053] In the formula: l represents the system layer overlap cost coefficient, M (t) This represents the total number of devices that performed maintenance actions at time t. This indicates the number of devices that underwent preventative maintenance at time t. The number of devices that undergo replacement and maintenance at time t;

[0054] Considering the overall system-level action costs and resource constraint costs, in the current action set a (t) The cost function c under execution (t) It can be represented as:

[0055]

[0056] Preferred,

[0057] In step S7, the reinforcement learning-based equipment cluster maintenance strategy optimization model processes a series of maintenance behaviors by setting a discount rate γ. The overall objective is:

[0058]

[0059] In the formula: C is the future long-term discount cost under the current maintenance strategy, i is the reinforcement learning time step, and γ is the reinforcement learning discount coefficient;

[0060] The optimization objective of the entire model is to obtain the best maintenance decision by minimizing C;

[0061] In the optimization of cluster operation and maintenance strategies, reinforcement learning optimization models are embedded into actual operation and maintenance scenarios to conduct decision-making research based on device status representation, maintenance behavior space and objective function.

[0062] Therefore, in view of the fact that actual operation and maintenance measures are often based on manual decision-making, which has a certain degree of retrospective nature and poses a risk of major accidents, this disclosure innovatively proposes a reinforcement learning-based maintenance method for aero-derivative gas turbine clusters. This method provides intelligent maintenance decisions for aero-derivative gas turbine cluster systems, thereby effectively reducing maintenance costs and improving equipment safety.

[0063] Compared with the prior art, this disclosure has at least the following beneficial effects:

[0064] 1) Based on the single-machine failure rate function constructed using the Weibull lifetime distribution, the changes in the single-machine post-repair performance caused by the number and type of maintenance and the investment cost are considered, thus forming a comprehensive failure rate change model.

[0065] 2) Divide the entire cluster environment into system layer and individual layer, and establish environment state representation space and behavior space at each layer. Establish cluster operation and maintenance objective function under the constraints of investment cost and resource limitation at different layers.

[0066] 3) The entire model building and objective solution are carried out within the reinforcement learning framework, which fully avoids the problem of large space that is difficult to solve in traditional optimization algorithms due to the complexity of the operation and maintenance environment, and has the potential to solve multi-level and multi-dimensional representation spaces.

[0067] 4) In cluster operation and maintenance, the maintenance plan will be changed from the original single-machine phased fixed-interval maintenance to starting from the "status - maintenance action - real-time value feedback" segment of the entire cluster to find a suitable status action strategy.

[0068] 5) By introducing cost and resource constraints, reinforcement learning-based cluster maintenance optimization can perform maintenance behavior under state guidance, thus avoiding fixed-interval maintenance.

[0069] Therefore, under the condition of meeting the single-machine operation and maintenance cost, model optimization decision-making can effectively improve the resource pressure of cluster maintenance and support points and enhance cluster scheduling capabilities. Attached Figure Description

[0070] Figure 1 This is a schematic diagram of the overall structure in one embodiment of this disclosure;

[0071] Figure 2 This is a schematic diagram of a device cluster scenario in one embodiment of this disclosure;

[0072] Figure 3 This is a schematic diagram illustrating the maintenance of overlapping time in one embodiment of this disclosure;

[0073] Figure 4 This is a schematic diagram of an algorithm in one embodiment of this disclosure;

[0074] Figure 5 This is a schematic diagram of a reinforcement learning decision structure in one embodiment of the present disclosure. Detailed Implementation

[0075] To enable those skilled in the art to understand the technical solutions disclosed herein, the following will describe them in conjunction with embodiments and related appendices. Figures 1 to 5 The technical solutions of various embodiments are described herein, and the described embodiments are only a part of the embodiments of this disclosure, not all of them. The terms "first," "second," etc., used in this disclosure are used to distinguish different objects, not to describe a specific order. Furthermore, "comprising" and "having," and any variations thereof, are intended to be omnipresent and non-exclusive. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, systems, products, or devices.

[0076] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this disclosure. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. Those skilled in the art will understand that the embodiments described herein can be combined with other embodiments.

[0077] In one embodiment, this disclosure discloses a reinforcement learning-based maintenance method for aero-derivative gas turbine clusters, characterized in that the method includes the following steps:

[0078] S1: Divide the entire cluster decision-making structure into individual layer I, system layer Sys, and resource layer R;

[0079] For example, according to S1, the model is verified under simulation conditions, mainly for the operation and maintenance process of a unit containing 3 aero-derivative gas turbines, i.e., the individual layer I = (e1, e2, e3), and the other layers are determined by the actual situation.

[0080] S2: Based on the Weibull failure rate model, an age delay factor and a failure rate increase factor are introduced to reflect the impact of maintenance frequency on the performance of a certain device in the entire cluster, where the device is a certain aero-derivative gas turbine.

[0081]

[0082] In the formula: i represents the number of repairs, which is a natural number; a i b is the service life delay factor. i As the failure rate increasing factor, This represents the failure rate level during the i-th maintenance. Similarly, it can be seen that... Where t represents the service time, β is the characteristic lifetime parameter, and ρ is the shape parameter;

[0083] For example, according to S2 above, there are the following parameters as shown in Table 1:

[0084] Table 1 Individual Failure Rate Parameter Settings

[0085]

[0086] Based on the table above, a failure rate model for three aero-derivative gas turbines can be established, a i This will be discussed in subsequent steps.

[0087] S3: Introducing effective service life and replacing the service life delay factor to reflect the change in failure rate under effective service life, where the effective service life W is at the i-th maintenance. [i] The calculation is as follows:

[0088]

[0089] In the formula: Let C be the total cost of the j-th repair, where j ranges from 1 to i. P To replace input costs, α is the adjustment parameter for input costs, γ is the time adjustment parameter, i is the same as in formula (1), and τ i This represents the time interval of the i-th maintenance activity;

[0090] Then the overall failure rate function after the i-th maintenance action is performed. for:

[0091]

[0092] In the formula, N represents the total number of maintenance activities performed. The failure rate level when no maintenance activities are performed;

[0093] For example: Establish the replaced comprehensive failure rate model according to the parameters shown in Table 2.

[0094] Table 2 Maintenance Parameters

[0095]

[0096] S4: To fully simulate the actual degradation process, a two-parameter exponential model containing a random error term is used to model the degradation:

[0097]

[0098] Degradation modeling is performed according to the parameters shown in Table 3:

[0099] Table 3 Individual Degradation Parameter Settings

[0100]

[0101] And in conjunction with the comprehensive failure rate function in step S3, model the multidimensional state space S of the system layer:

[0102]

[0103] S5: System-level actions are categorized into preventative maintenance based on different maintenance activity types. and no action a N Two kinds;

[0104] S6: Combine motion costs and resource constraint costs into a cost function and form the expected value;

[0105] For example, the specific parameters are shown in Table 4 below:

[0106]

[0107] S7: Simulate the cluster devices within the reinforcement learning framework, and optimize using the deep Q-network algorithm with the goal of minimizing the cost function.

[0108] For example, the method disclosed in this disclosure is compared with the cost of a fixed interval strategy, and the results are shown in Table 5 below:

[0109]

[0110] It can be observed that the reinforcement learning-based maintenance method for aero-derivative gas turbine clusters, due to the adoption of the specific reinforcement learning strategy described in this disclosure, achieves a higher overall cost rate (RMB / thousandth of h) compared to the fixed-interval strategy. -1 and time overlap cost rate / thousand yuan·h -1In both metrics, maintenance costs have been significantly reduced.

[0111] The above embodiments embody the main inventive concept of this disclosure:

[0112] To resolve the conflict of interest between individuals and the cluster in equipment cluster scenarios, the entire cluster decision-making structure is first divided into an individual layer (I), a system layer (Sys), and a resource layer (R). Based on the Weibull failure rate model, an age delay factor and a failure rate increment factor are introduced to reflect the impact of maintenance frequency on component performance. Effective service life is introduced and replaced with the age delay factor to reflect changes in failure rate over the effective service life. A two-parameter exponential degradation model is established based on the degradation state assessment results obtained from multi-source information fusion for actual aero-derivative gas turbine degradation, and combined with the failure rate, it is modeled as a multi-dimensional state space S at the system layer. For different maintenance activity types, system layer actions are categorized as: preventative maintenance... and no action a N Two methods are proposed: combining action costs and resource constraint costs into a cost function to form the expected value; and finally, simulating the cluster devices within a reinforcement learning framework.

[0113] In one embodiment,

[0114] In step S1, each individual layer I includes multiple individual devices e i The state of the important components inside each piece of equipment determines the overall performance of the equipment. When the failure rate of important components is too high, leading to increased operational risks or serious failures, the entire equipment will be unable to perform its work tasks. The important internal components include components such as the combustion chamber of the aero-derivative gas turbine and the compressor. The value of i is the number of equipment.

[0115] In one embodiment,

[0116] The system layer Sys comprises r individual layers, Sys = (I1, I2, ..., I... r Each individual layer performs its own tasks while also being scheduled by the cluster center.

[0117] In one embodiment,

[0118] The resource layer R includes various operation and maintenance resources, such as monitoring and performance evaluation equipment for the aero-derivative gas turbine prediction and health management system, the number of maintenance support points, and the maximum number of workers at each support point. These operation and maintenance resources are often not unlimited, which will limit the execution of maintenance decisions.

[0119] In one embodiment,

[0120] In step S2, the Weibull failure rate model is used, and different model parameters are set for different devices to distinguish their failure rates and reflect the impact of maintenance frequency on the failure rate. Different settings are made for ρ and β to distinguish different devices.

[0121] In one embodiment,

[0122] In step S3, the actual service life of a component over its entire life cycle is often not its in-situ working time, but rather the component's effective service life W. [i] This reflects the impact of the type of maintenance activity and the level of cost invested on the equipment recovery level.

[0123] In one embodiment,

[0124] In step S4, the actual service life of the equipment throughout its entire life cycle is often not its in-situ working time, but rather the effective service life W of the components. [i] This reflects the impact of the type of maintenance activity and the level of cost invested on the equipment recovery level.

[0125] In one embodiment,

[0126] In step 5, system-level maintenance actions are divided into preventative maintenance. and no action a N There are two types of maintenance actions, both requiring reinforcement learning for decision-making. One involves selecting an action as the maintenance behavior at a fixed monitoring interval. If a fault occurs between two intervals, replacement maintenance is required. Both preventative and replacement maintenance incur economic and time costs and impact the reinforcement learning environment. The final maintenance action is component replacement. Rc No decision-making is required.

[0127] In one embodiment,

[0128] In step S6, the cost functions are divided into two categories:

[0129] Action cost:

[0130] Assume that during device operation, devices e in the individual layer within the device cluster are evaluated at time intervals of ΔT (ΔT > 0). i Perform status monitoring and during operation and maintenance, for a duration T Pc And the cost of a single investment, c Pc To perform preventative maintenance, the labor and material costs required for preventative maintenance are included in the single-use cost. The failure risk cost of critical components during condition monitoring is assumed to be v. Fc The unit time delay cost is c Sc MTTR represents the mean time to repair critical components, and C represents the cost of performing a single action at the individual level. P,iThen it can be expressed as:

[0131]

[0132] In the formula: C P,i Let i represent the cost function for an individual i to perform preventative maintenance actions;

[0133] The system-level cost C under the current action set execution. sys The calculation is as follows:

[0134]

[0135] In the formula: The number of devices that require preventative maintenance at time t;

[0136] Resource-constrained costs:

[0137] To avoid increased logistical maintenance pressure and reduced cluster availability due to simultaneous maintenance of multiple devices, system-level optimization is needed to limit the number of devices requiring maintenance at the same time. This can be achieved by reducing the overlap of device maintenance times.

[0138]

[0139] In the formula: d i This represents the maintenance duration of device number i, satisfying d. min >ΔT,t i t k These represent the maintenance times for devices numbered i and k, respectively, with b ranging from 1 to d. i Ol ik This indicates the overlap in maintenance times for devices numbered i and k.

[0140] Based on this, calculate the current system-level action set a. (t) The system-level maintenance overlap cost under execution, considering the cost of maintaining the overlap between two adjacent action sets {a (t-1) a (t) The execution process is described, and assuming M represents the total number of devices involved in maintenance under two action sets at the system layer, then the overall overlapping cost of the system layer under the current action set execution is calculated. for:

[0141]

[0142] In the formula: l represents the system layer overlap cost coefficient, M (t) This represents the total number of devices that performed maintenance actions at time t. This indicates the number of devices that underwent preventative maintenance at time t. The number of devices that undergo replacement and maintenance at time t;

[0143] Considering the overall system-level action costs and resource constraint costs, in the current action set a (t) The cost function c under execution (t) It can be represented as:

[0144]

[0145] In one embodiment,

[0146] In step S7, the reinforcement learning-based equipment cluster maintenance strategy optimization model processes a series of maintenance behaviors by setting a discount rate γ. The overall objective is:

[0147]

[0148] In the formula: C is the future long-term discount cost under the current maintenance strategy, i is the reinforcement learning time step, and γ is the reinforcement learning discount coefficient;

[0149] The optimization objective of the entire model is to obtain the best maintenance decision by minimizing C;

[0150] In the optimization of cluster operation and maintenance strategies, reinforcement learning optimization models are embedded into actual operation and maintenance scenarios to conduct decision-making research based on device status representation, maintenance behavior space and objective function.

[0151] See further Figure 2 , Figure 2 This diagram illustrates a device cluster scenario. The entire cluster decision-making structure is divided into three layers: Individual Layer (I), System Layer (Sys), and Resource Layer (R). At the Individual Layer, the optimization objective is to reduce the maintenance cost of a single device. Individual layer optimization can achieve optimal individual maintenance costs. However, due to resource layer limitations, using optimal parameters at the Individual Layer for maintenance in a device cluster will result in significant economic costs and security risks. Therefore, the System Layer, based on Individual Layer optimization, comprehensively schedules individuals by considering the maintenance resource constraints at the Resource Layer.

[0152] See further Figure 3 As a schematic diagram of the overlapping time maintenance in this disclosure, t in the figure i t k The maintenance times for devices numbered i and k are respectively. This represents three different scenarios on the timeline for the maintenance duration of device number i. Figure 3 This illustrates the cases where j ranges from 1 to 3.

[0153] Please refer to Table 6 for the algorithm flow of this disclosure, which illustrates how to optimize the maintenance strategy using the method proposed in this disclosure, and describes the steps such as algorithm update and sampling:

[0154] Table 6

[0155]

[0156] See further Figure 4 The deep Q-network algorithm structure used in this disclosure is a 3-layer multi-neural network with 64 neurons in each layer, and the ReLU activation function is used in the middle. This can be understood in conjunction with Table 6 above.

[0157] See further Figure 5 As the reinforcement learning decision structure disclosed herein, the left side represents the modeling based on actual operation and maintenance scenarios. It constructs the overall state representation S = (s...) by acquiring degradation data of key components, fault state data, and operation and maintenance data. (1) s (2) , ..., s (t) ), maintain the behavior space A = (a (1) a (2) , ..., a (t) ) and cost matrix C = (c (1) c (2) c (t) ); where s (t) ∈{(L i D i Let )|i=1,2,...} represent the environmental state at time t, and a (t) ∈{a Pc a N} represents the maintenance behavior at the current time step, r (t) =c (t) This represents the immediate reward for the agent's interaction with the environment at time step t; it is then embedded in the reinforcement learning on the right side to learn and maintain the policy π = p(a (t) |s (t) Strategy π needs to continuously interact with the environment during execution to obtain state variables and maintenance cost rewards r. (t) =c (t) This allows for the continuous improvement of π through value iteration or strategy iteration, ultimately yielding the optimal maintenance strategy for the aero-derivative gas turbine cluster in a dynamic environment. Provide reasonable suggestions for equipment operation and maintenance management.

[0158] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0159] The above-described embodiments are only used to illustrate the technical solutions of this disclosure, and are not intended to limit it. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this disclosure.

Claims

1. A maintenance method for aero-derivative gas turbine clusters based on reinforcement learning, characterized in that, The method includes the following steps: S1: Divide the entire cluster decision-making structure into individual layers. System layer and resource layer : S2: Based on the Weibull failure rate model, an age delay factor and a failure rate increase factor are introduced to reflect the impact of maintenance frequency on the performance of a certain device in the entire cluster, where the device is a certain aero-derivative gas turbine. , In the formula: This indicates the number of repairs, and its value is a natural number. As a factor for delaying the age of service, As the failure rate increasing factor, Indicates the first The failure rate level after the second repair can be understood similarly. ,in Indicates service time. For characteristic lifetime parameters, For shape parameters; S3: Introducing effective working years and replacing the service life delay factor to reflect changes in failure rate under effective working years, the first Effective service life after the second repair The calculation is as follows: , In the formula: For the first The cost of the second repair From 1 to , To replace the investment cost, As an adjustment parameter for input costs, For time adjustment parameters, Same as formula (1), Indicates the first The time interval between maintenance activities; Then the first Comprehensive failure rate function after the second maintenance action for: , In the formula, This indicates the total number of maintenance activities performed. The failure rate level before maintenance is performed; S4: To fully simulate the actual degradation process, a two-parameter exponential model containing a random error term is used to model the degradation: , In the formula, Indicates the service life as The state of degradation at that time, For constant terms, express Follow the mean The variance is A normally distributed random variable, The random error term follows a mean of 0 and a variance of . The normal distribution To conform to the mean The variance is A normally distributed random variable; Furthermore, for all i different devices in the entire cluster, the multidimensional state space S of the system layer is modeled using the integrated failure rate function in step S3: , In the formula, Indicates different devices, Let be the characteristic lifetime parameter of the i-th device. For the first Shape parameters of the device; The function representing the overall failure rate of the i-th device ; Indicates the i-th device and its service life as . Degradation status at that time ,in Represents the i-th device The parameter set consisting of all parameters except t; S5: System-level actions are categorized into preventative maintenance based on different maintenance activity types. and no action There are two types of actions, both of which require reinforcement learning to make decisions, while the final maintenance action is component replacement. No decision-making is required; S6: Combine motion costs and resource constraint costs into a cost function and form the expected value; S7: Simulate cluster devices within a reinforcement learning framework, using deep learning. The network algorithm optimizes by minimizing the cost function.

2. The method according to claim 1, characterized in that, In step S1, the individual level Includes multiple individual devices The condition of the critical components inside each piece of equipment determines the overall performance of the equipment. When the failure rate of critical components is too high, leading to increased operational risks or serious malfunctions, the entire equipment will be unable to perform its tasks. These critical internal components include the combustion chamber of the aero-derivative gas turbine and compressor components. The value is the number of devices.

3. The method according to claim 1, characterized in that, System layer include Individual level, Each individual layer performs its own tasks while also being scheduled by the cluster center.

4. The method according to claim 1, characterized in that, resource layer This includes various operation and maintenance resources: monitoring and performance evaluation equipment for the aero-derivative gas turbine predictive and health management system, the number of maintenance support points, and the maximum number of workers at each support point. These operation and maintenance resources are often not unlimited, which will limit the execution of maintenance decisions.

5. The method according to claim 1, characterized in that, In step S2, for the Weibull failure rate model, different model parameters are set for different devices to distinguish their failure rates and reflect the impact of maintenance frequency on the failure rate. This is achieved through... Different settings are configured to differentiate between different devices.

6. The method according to claim 1, characterized in that, In step S3, the actual service life of the equipment throughout its entire life cycle is often not its in-situ working time, but rather the effective service life of its components. This reflects the impact of the type of maintenance activity and the level of cost invested on the equipment recovery level.

7. The method according to claim 1, characterized in that, In step S4, the two-parameter exponential degradation process model containing a random error term distinguishes the degree of degradation between different devices by setting different model parameters. All parameters can be set differently.

8. The method according to claim 1, characterized in that, In step S5, system-level maintenance actions are divided into preventative maintenance. and no action There are two types of maintenance actions. These actions require reinforcement learning for decision-making. At fixed monitoring intervals, one action is selected as the maintenance behavior. If a fault occurs between two intervals, replacement maintenance is required. Both preventative and replacement maintenance incur economic and time costs and impact the reinforcement learning environment. The final maintenance action is component replacement. No decision-making is required.

9. The method according to claim 1, characterized in that, In step S6, the cost functions are divided into two categories: Action cost: Assuming the device is running at... The time interval is for devices in the individual layer within the device cluster. Perform status monitoring and, during the operation and maintenance process, measure the duration. and one-time investment cost To perform preventative maintenance, the labor and material costs required for preventative maintenance are included in the single-use cost. The failure risk cost of critical components during condition monitoring is assumed to be... The unit time delay cost is , This represents the average post-repair time for critical components and the cost of performing a single action at the individual level. Then it can be expressed as: , In the formula: Represents an individual The cost function for performing preventative maintenance actions; The cost incurred at the system level under the current action set execution The calculation is as follows: , In the formula: For a moment The number of devices that require preventative maintenance. Resource-constrained costs: To avoid increased logistical maintenance pressure and reduced cluster availability due to simultaneous maintenance of multiple devices, system-level optimization is needed to limit the number of devices requiring maintenance at the same time. This can be achieved by reducing the overlap of device maintenance times. , In the formula: Indicates the number is The equipment maintenance duration meets the requirements. , They represent the numbers respectively. and Equipment maintenance time, Values ​​range from 1 to , Indicates the number is , Equipment maintenance time overlap; Based on this, calculate the current system-level action set. The system-level maintenance overlap cost under execution, considering the cost of maintaining overlap between two adjacent action sets. Execution process, and assumptions Given the total number of devices involved in maintenance under two action sets at the system layer, the overall overlap cost at the system layer under the current action set execution is calculated. for: , In the formula: This represents the system-level overlap cost coefficient. express The total number of devices performing maintenance actions at all times. Indicates time The number of devices that require preventative maintenance. time The number of devices that require replacement and maintenance at any given time; Considering the overall system-level action costs and resource constraints, in the current action set Cost function under execution It can be represented as: 。 10. The method according to claim 1, characterized in that, In step S7, the equipment cluster maintenance strategy optimization model based on reinforcement learning sets a discount rate. This involves processing a series of maintenance actions, with the overall goal of: , In the formula: To cover the future long-term discount costs under the current maintenance strategy, To enhance learning time steps, To enhance the learning discount factor; The optimization objective of the entire model is to minimize To obtain the best maintenance decisions; Therefore, in the optimization of cluster operation and maintenance strategies, the method embeds a reinforcement learning optimization model into the actual operation and maintenance scenario based on device status representation, maintenance behavior space and objective function, thereby providing quantitative and data support for decision research.

Citation Information

Patent Citations

  • Aero-engine component maintenance method based on time-varying fault rate model

    CN113374543A

  • Land-based equipment cluster maintenance resource intelligent configuration and scheduling strategy making method based on reinforcement learning

    CN114399155A