A method and system for optimizing equipment maintenance strategy of a waste incineration power plant based on reinforcement learning

CN122820190APending Publication Date: 2026-09-25LUZHOU XINGLU ENVIRONMENTAL GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611034731.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,上述方法在垃圾焚烧场景下存在共同的缺陷:奖励塑形无法提供理论上的约束满足保证,训练过程中仍会出现大量约束违反,且惩罚系数难以整定;在安全攸关场景中,仅靠软惩罚无法满足法规对零超标的硬性要求

Benefits of technology

本发明针对垃圾焚烧发电厂高温、高腐蚀及强时变约束条件下维护决策难以兼顾学习效率与安全约束严格满足的问题,实现了维护策略从数据学习到安全执行的全流程闭环控制。与现有技术相比,本发明通过动态约束边界函数刻画设备随退化过程的安全边界收缩机制,使约束由静态阈值转化为状态相关的时变函数,有效提升了约束表达的物理一致性,对策略输出进行逐步可行域投影,实现单步决策层面的安全修正,从执行层面降低约束越界风险;本发明将约束违反从单步控制扩展至长期概率约束优化,实现了约束违反频率在统计意义上的有界控制,同时引入基于贝叶斯神经网络的置信度评估机制,对策略输出提供约束达标概率预测,使系统具备可解释性与风险可量化能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820190A_ABST
    Figure CN122820190A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on reinforcement learning's waste incineration power plant equipment maintenance strategy optimization method and system, comprising: establishing constraint Markov decision model containing equipment degradation characteristic quantity, real-time operating parameter, multistage maintenance action and safety constraint;Dynamic constraint boundary function is constructed and the maximum allowable temperature and stress are output;Pretrain maintenance strategy in combination with offline data and digital twin environment;Based on lyapunov derivative and quadratic programming, online safety correction is carried out;Allowable stress, temperature coupling relationship is dynamically updated using bayesian mechanism;Adopt dual gradient descent control long-term constraint violation probability;Output deployable maintenance strategy with constraint compliance confidence report.The application effectively solves the technical bottleneck problem that traditional reinforcement learning method is difficult to consider learning efficiency and safety reliability under the complex time-varying constraint scene of waste incineration power plant.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data acquisition technology, and in particular to a method and system for optimizing equipment maintenance strategies in waste-to-energy incineration plants based on reinforcement learning. Background Technology

[0002] Waste-to-energy incineration is a mainstream method for reducing and recycling municipal solid waste. Its core equipment includes the incinerator grate, waste heat boiler heating surfaces, and flue gas purification system, which operate under extreme conditions of high temperatures (furnace temperature > 850℃), high corrosiveness (HCl, SO2, alkali metal chlorides), and high fly ash erosion. During operation, the equipment gradually exhibits degradation phenomena such as grate jamming, high-temperature corrosion thinning, coking of heating surfaces, and tube rupture. Improper maintenance can lead not only to unplanned shutdowns and reduced power generation revenue but also potentially to major safety and environmental accidents such as excessive dioxin emissions and furnace overheating.

[0003] In recent years, reinforcement learning-based optimization of equipment maintenance strategies has made some progress in fields such as power and manufacturing. These methods learn optimal maintenance timing and actions through interaction between agents and the environment, aiming to achieve a balance between equipment availability and maintenance costs. However, directly applying reinforcement learning to maintenance decisions in waste-to-energy plants faces a fundamental bottleneck: the strict satisfaction of safety and environmental constraints.

[0004] Specifically, the operation and maintenance of waste incineration plants must adhere to a series of inviolable red lines: The furnace temperature must be maintained above 850℃ and the residence time must be ≥2 seconds to suppress the formation of dioxins; at the same time, the levels of CO, HCl, particulate matter, etc. in the flue gas must not exceed the national standards.

[0005] The local temperature of the grate must not exceed the allowable temperature of the material (e.g., not exceeding 1100℃), the temperature of the superheater tube wall of the waste heat boiler must not exceed the creep limit, and the fluctuation range of the negative pressure in the furnace must be controlled within the safe range.

[0006] When equipment suffers severe degradation, a shutdown for maintenance must be arranged. However, shutdown during maintenance can lead to waste accumulation and power generation interruption. Therefore, the maintenance strategy must be decided under the constraint of continuous operation: partial repair or complete shutdown.

[0007] The above constraints have three significant characteristics: The constraint boundary depends on real-time operating parameters (temperature, pressure, vibration, corrosion rate) and the remaining life of the equipment. It is a continuous, high-dimensional state function, rather than a simple action space limit.

[0008] As equipment ages (such as long-term high-temperature creep of boiler tubes), its safety boundaries, such as the maximum allowable operating temperature and maximum pressure difference, will gradually shrink, and the constraints will exhibit nonlinear drift.

[0009] Even brief overheating or excessive emissions can trigger environmental penalties, catastrophic equipment failures, or even personal injury or death. Therefore, even minor violations of regulations during training are unacceptable.

[0010] Existing reinforcement learning methods mainly fall into two categories regarding constraint handling: one is to use reward shaping to set constraint violations as negative rewards, attempting to guide the agent to avoid unsafe areas; the other is to use Lagrange multiplier methods (such as CPO, RCPO) or Lyapunov-based RL methods to transform constraint optimization into a dual problem. However, these methods share common drawbacks in the waste incineration scenario: reward shaping cannot provide a theoretical guarantee of constraint satisfaction, a large number of constraint violations still occur during training, and the penalty coefficient is difficult to tune; in safety-critical scenarios, soft penalties alone cannot meet the regulatory requirement of zero exceedances. Constraint reinforcement learning methods (such as CPO) assume that the constraint function is known and static, and cannot handle the problem of dynamic drift of constraint boundaries in waste incinerators as the equipment ages; when allowable temperature, allowable stress, and other limit values ​​decrease with corrosion thinning, the algorithm cannot adaptively update the constraint model. Waste incineration plants do not allow online trial and error, as each violation of constraints may lead to damage to real equipment or emission accidents; while when using offline datasets or simulation environments, the deviation between the constraint boundaries and the real aging process makes it difficult for the strategy to converge to a truly safe solution, or even fall into conservative deadlock (e.g., always shut down and give up operation).

[0011] In summary, traditional reinforcement learning methods struggle to balance learning efficiency and safety reliability in complex, time-varying constraints at waste-to-energy plants, representing a key technical problem that remains unsolved and urgently needs to be addressed in the field. Summary of the Invention

[0012] To address the aforementioned technical problems, this invention provides a method and system for optimizing equipment maintenance strategies in waste-to-energy plants based on reinforcement learning.

[0013] To achieve the above objectives, the technical solution adopted by the present invention is as follows: On one hand, this invention discloses a method for optimizing equipment maintenance strategies in waste-to-energy incineration plants based on reinforcement learning, characterized by the following steps: Step 1: Establish a constrained Markov decision model. The state space includes equipment degradation characteristics and real-time operating parameters. The action space consists of multi-level maintenance actions. The reward function is power generation revenue minus maintenance costs. The constraint set includes emission safety constraints and equipment thermal safety constraints. Step 2: Based on the first principles of thermodynamics and material creep fracture data, construct a dynamic constraint boundary function. This function takes the cumulative damage degree of the equipment and the real-time degradation characteristics as inputs and outputs the maximum allowable temperature and the maximum allowable stress. Step 3: Using offline historical operation data and a high-fidelity digital twin environment, the initial maintenance strategy network is pre-trained through offline strategy evaluation and context constraint strategy optimization algorithms. The network outputs maintenance actions that satisfy the dynamic constraint boundary function. Step 4: Design an online safety correction layer based on Lyapunov candidate functions. This correction layer calculates the Lyapunov derivative of the current policy action in real time. If the derivative is greater than zero, it solves the minimum correction action through quadratic programming and forces the system state to drift back to the safe region. Step 5: Introduce an adaptive constraint boundary update module. This module uses online monitoring data from the equipment and Bayesian linear regression to dynamically update the allowable stress-temperature coupling relationship in the dynamic constraint boundary function. Step 6: Using dual gradient descent and Lagrange multiplier adaptive mechanism, the risk tolerance coefficient of constraint violation is adjusted in policy iteration so that the long-term constraint violation probability is bounded to the preset safety threshold. Step 7: Output a deployable maintenance strategy. This strategy receives the current status and outputs maintenance actions at each decision point, while generating a constraint compliance confidence report for the corresponding action.

[0014] Furthermore: Step one includes: The equipment degradation characteristics in the state space of the constrained Markov decision model include the cumulative running time of the incinerator grate, the reduction in wall thickness of the waste heat boiler heating surface, and the high-temperature corrosion rate. Real-time operating parameters in the state space include furnace outlet flue gas temperature, flue gas oxygen content, grate drive motor current, and carbon monoxide concentration in the flue gas. The multi-level maintenance actions in the action space include non-maintenance state maintenance actions, local water flushing and descaling actions, high-temperature anti-corrosion coating spraying and repair actions, and system-wide shutdown and overhaul actions. The power generation revenue in the reward function is calculated based on the product of the real-time grid-connected electricity price of the waste-to-energy plant and the power generation capacity. The maintenance cost in the reward function is calculated based on the weighted sum of the spare parts cost and labor cost corresponding to different actions in the multi-level maintenance action. The emission safety constraints are centralized, namely, the flue gas temperature at the furnace outlet is not less than 850 degrees Celsius and the flue gas residence time is not less than 2 seconds. The equipment thermal safety constraints are centralized, including the maximum allowable wall temperature limit on the exhaust side of the incinerator and the maximum allowable creep stress limit on the superheater tubes of the waste heat boiler.

[0015] Furthermore, step two includes: The dynamic constraint boundary function adopts a nonlinear mapping relationship in the form of cumulative piecewise segments. The first-level mapping maps the cumulative damage degree of the equipment and the real-time degradation characteristic quantity together into the material residual strength reduction coefficient. The cumulative damage degree of the equipment is calculated jointly according to the linear Palmgren-Miner fatigue accumulation law and the creep time fraction law. The real-time degradation characteristic quantity includes the high-temperature corrosion thinning amount on the exhaust side of the incinerator and the creep elongation amount of the superheater tube of the waste heat boiler. The second-order mapping in the dynamic constraint boundary function multiplies the material residual strength reduction factor with the nominal allowable temperature and nominal allowable stress calculated based on the first principle of thermodynamics, to obtain the maximum allowable temperature and maximum allowable stress at the current moment, respectively. The nominal allowable temperature is derived in reverse from the furnace adiabatic combustion temperature and the flue gas convective heat transfer coefficient, and the nominal allowable stress is determined by interpolation of the endurance strength limit from the Lamé formula and the material creep fracture test data.

[0016] Furthermore, step three includes: Offline historical operation data includes time-series data of flue gas temperature at furnace outlet, time-series data of current of incinerator grate drive motor, inspection records of wall thickness reduction of waste heat boiler heating surface, and labels of multi-level maintenance actions performed at the corresponding time during at least one complete maintenance cycle of the waste incineration power plant. The digital twin environment is jointly constructed based on the finite element thermo-mechanical coupling model and the computational fluid dynamics flue gas flow model. The digital twin environment takes the real-time operating parameters in the offline historical operating data as the input boundary conditions and outputs the equipment degradation characteristics and dynamic constraint boundary function values ​​in the simulation state space. The offline strategy evaluation adopts the fitted Q evaluation algorithm. The fitted Q evaluation algorithm uses the state-action-reward-next state quadruple from the offline historical running data to train a state-action value network. The state-action value network is used to evaluate the expected cumulative reward of the maintenance action output by the initial maintenance strategy network in the digital twin environment. The context constraint policy optimization algorithm adopts the constraint policy optimization algorithm framework. In the constraint update step, the estimated value of the constraint violation probability is obtained by comparing the dynamic constraint boundary function with the simulation state output by the digital twin environment in real time. When optimizing the parameters of the initial maintenance policy network, the context constraint policy optimization algorithm forces the simulation constraint violation probability corresponding to the maintenance action output by the initial maintenance policy network to be lower than the preset offline training constraint threshold. The initial maintenance strategy network adopts a parameterized Gaussian strategy network structure. The initial maintenance strategy network takes the equipment degradation features and real-time operating parameters in the state space as inputs, and outputs the mean and variance of the multi-level maintenance action combination. The pre-trained initial maintenance strategy network is obtained by executing the constrained strategy optimization algorithm in the digital twin environment until the strategy network loss function converges.

[0017] Furthermore: Step four includes: The online safety correction layer based on Lyapunov candidate functions constructs a positive definite Lyapunov candidate function with emission safety constraints and equipment thermal safety constraints in the constrained Markov decision model as the boundary. The positive definite Lyapunov candidate function is defined as the square of the weighted Euclidean distance between the current system state and the safety region boundary defined by the dynamic constraint boundary function. The online safety correction layer based on Lyapunov candidate functions performs the following computational process within each decision step: First, the original maintenance action output by the initial maintenance strategy network is received. Then, based on the current system state, the time derivative of the positive definite Lyapunov candidate function induced by the original maintenance action is calculated to obtain the Lyapunov derivative. If the Lyapunov derivative is negative or zero, the online safety correction layer directly outputs the original maintenance action as the final maintenance action. If the Lyapunov derivative is positive, the online safety correction layer starts the quadratic programming solver. The quadratic programming solver takes minimizing the L2 norm square between the correction action and the original maintenance action as the objective function and the safety shrinkage coefficient that forces the Lyapunov derivative to be no greater than negative as the constraint condition to solve for the minimum correction action. The online safety correction layer outputs the minimum correction action as the final maintenance action and applies the final maintenance action to the high-temperature critical equipment of the waste incineration power plant, causing the system state to drift back to the safe region defined by the dynamic constraint boundary function along the direction of the decreasing Lyapunov function value.

[0018] Furthermore, step five includes: The adaptive constraint boundary update module collects real-time online monitoring data of key high-temperature equipment in waste incineration power plants. The online monitoring data includes the actual temperature measured by infrared temperature sensors on the outer wall of the incinerator grate, the actual temperature measured by thermocouples on the outer wall of the superheater tubes of the waste heat boiler, and the carbon monoxide concentration and dioxin precursor concentration measured by the continuous flue gas emission monitoring system. The adaptive constraint boundary update module uses the event that the measured temperature in the online monitoring data exceeds the maximum allowable temperature output by the dynamic constraint boundary function as the trigger condition, and collects the cumulative damage degree of the equipment, real-time degradation characteristics, and measured stress-temperature paired data points at the trigger time; The adaptive constraint boundary update module uses a Bayesian linear regression framework to update the allowable stress-temperature coupling relationship online. This includes using the material's endurance strength limit as the prior mean function of the Gaussian process, using measured stress-temperature paired data points as observation samples, and obtaining the updated posterior mean function and posterior covariance function of the allowable stress-temperature coupling relationship by calculating the mean and variance of the posterior distribution. The adaptive constraint boundary update module uses the posterior mean function as the updated allowable stress-temperature coupling relationship and writes it back into the dynamic constraint boundary function. The dynamic constraint boundary function recalculates the maximum allowable temperature and maximum allowable stress at the current moment based on the updated allowable stress-temperature coupling relationship, so as to realize the closed-loop adaptive adjustment of the dynamic constraint boundary function as the equipment ages.

[0019] Furthermore, step six includes: Before each update of the strategy network parameters during the strategy iteration process, the cumulative constraint violation frequency of the current initial maintenance strategy network on the most recent batch of offline historical running data or online interaction trajectory is calculated. The cumulative constraint violation frequency is the ratio of the total number of decision steps that violate emission safety constraints or equipment thermal safety constraints to the total number of decision steps. A dual gradient descent method is used to maintain a Lagrange multiplier variable. This Lagrange multiplier variable is used to penalize deviations where the cumulative constraint violation frequency exceeds a preset safety threshold. The update rule for the Lagrange multiplier variable is as follows: Subtract the preset safety threshold from the cumulative constraint violation frequency, multiply the difference by the dual step size and accumulate it to the Lagrange multiplier variable, and finally project the Lagrange multiplier variable to the non-negative real number field. The risk tolerance coefficient is defined as a monotonically decreasing function of the Lagrange multiplier variable. The monotonically decreasing function restricts the risk tolerance coefficient to the open interval between 0 and 1. When the Lagrange multiplier variable approaches infinity, the risk tolerance coefficient approaches 0, and when the Lagrange multiplier variable approaches 0, the risk tolerance coefficient approaches 1. The risk tolerance coefficient is embedded in the Lagrange reward function of the constrained Markov decision process. The Lagrange reward function is equal to the reward function minus the product of the Lagrange multiplier variable and the constraint violation indicator function. The Lagrange reward function is then used as the optimization objective for updating the network parameters of the initial maintenance strategy. By adaptively adjusting the Lagrange multiplier variable, the exponential moving average of the long-term cumulative constraint violation frequency is bounded to a preset safety threshold.

[0020] Furthermore, step seven includes: The deployable maintenance strategy consists of a target strategy network with frozen parameters and a lightweight constraint confidence evaluation network. The target strategy network with frozen parameters is obtained by copying the final maintenance action output parameters of the initial maintenance strategy network after it has been corrected by the online safety correction layer. The target strategy network with frozen parameters is not updated with gradients during the deployment phase. The constraint compliance confidence report is generated in real time by the lightweight constraint confidence assessment network at each decision moment. The lightweight constraint confidence assessment network adopts a Bayesian neural network architecture, takes the current system state as input, and outputs the constraint compliance probability value corresponding to the maintenance action output by the target policy network with frozen parameters. The constraint compliance probability value is the predicted probability that the maintenance action simultaneously meets the emission safety constraint and the equipment thermal safety constraint under the current value of the dynamic constraint boundary function. Before deployment, the lightweight constraint confidence assessment network is trained using offline historical operating data and simulation trajectories generated by the digital twin environment for posterior probability distillation. The training label is a binary indicator vector indicating whether the maintenance action at each decision step in the simulation trajectory actually violates emission safety constraints and equipment thermal safety constraints. The lightweight constraint confidence assessment network continuously receives online monitoring data from the device during deployment and updates the posterior parameters of the internal Bayesian neural network online using a sliding window approach, so that the constraint compliance probability value is dynamically adjusted as the device ages. The deployable maintenance strategy displays the maintenance actions output by the target strategy network with frozen parameters and the constraint compliance confidence report output by the lightweight constraint confidence assessment network at each decision moment. When the constraint compliance probability value in the constraint compliance confidence report is lower than the preset confidence threshold, the deployable maintenance strategy simultaneously issues a manual review warning signal.

[0021] On the other hand, this invention also discloses a reinforcement learning-based optimization system for equipment maintenance and repair strategies in waste-to-energy incineration plants, comprising: The system modeling module establishes a constrained Markov decision model. The state space includes equipment degradation characteristics and real-time operating parameters, the action space includes multi-level maintenance actions, the reward function includes power generation revenue minus maintenance costs, and the constraint set includes emission safety constraints and equipment thermal safety constraints. The dynamic constraint module constructs dynamic constraint boundary functions. The functions take the cumulative damage degree of the equipment and the real-time degradation characteristics as inputs and output the maximum allowable temperature and the maximum allowable stress. The strategy pre-training module uses offline historical running data and a digital twin environment to pre-train the initial maintenance strategy network. The network outputs maintenance actions that satisfy the dynamic constraint boundary function. The Lyapunov correction module calculates the Lyapunov derivative of the current policy action in real time. If the derivative is greater than 0, it solves the minimum correction action through quadratic programming and forces the state to drift back to the safe area. The Bayesian update module dynamically updates the allowable stress-temperature coupling relationship in the dynamic constraint boundary function; The dual optimization module employs dual gradient descent and Lagrange multiplier adaptive mechanism to adjust the risk tolerance coefficient of constraint violation during policy iteration, so that the long-term constraint violation probability is bounded to a preset safety threshold. The deployment decision module outputs deployable maintenance strategies. At each decision moment, the strategy receives the current status and outputs maintenance actions, while generating a constraint compliance confidence report for the corresponding actions.

[0022] The technological advancements achieved by this invention compared to existing technologies are as follows: This invention addresses the challenge of balancing learning efficiency and strict safety constraint compliance in maintenance decision-making under the high-temperature, high-corrosion, and strongly time-varying constraints of waste-to-energy plants. It achieves closed-loop control of the entire maintenance strategy process, from data learning to safe execution. Compared to existing technologies, this invention uses dynamic constraint boundary functions to characterize the safety boundary contraction mechanism of equipment during degradation, transforming constraints from static thresholds into state-dependent time-varying functions. This effectively improves the physical consistency of constraint expression and enables stepwise feasible region projection of the strategy output, achieving safe correction at the single-step decision-making level and reducing the risk of constraint violation at the execution level. Furthermore, this invention extends constraint violation control from single-step control to long-term probabilistic constraint optimization, achieving statistically bounded control of constraint violation frequency. Simultaneously, it introduces a confidence assessment mechanism based on Bayesian neural networks to provide constraint compliance probability predictions for the strategy output, giving the system interpretability and quantifiable risk capabilities.

[0023] Therefore, while ensuring efficient offline learning capabilities, this invention achieves zero violation or probabilistic bounded safety control during the policy execution phase. It effectively solves the technical bottleneck problem that traditional reinforcement learning methods struggle to balance learning efficiency and safety reliability in complex time-varying constraint scenarios at waste-to-energy plants, and significantly improves the safety, stability, and engineering deployability of maintenance decisions for high-temperature critical equipment. Attached Figure Description

[0024] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.

[0025] In the attached diagram: Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0026] The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of the present invention will now be described with reference to the accompanying drawings.

[0027] Example 1 like Figure 1As shown, this invention discloses a reinforcement learning-based optimization method for equipment maintenance and repair strategies in waste-to-energy incineration plants, comprising: Step 1: Establish a constrained Markov decision model. The state space includes equipment degradation characteristics and real-time operating parameters, the action space includes multi-level maintenance actions, the reward function includes power generation revenue minus maintenance costs, and the constraint set includes emission safety constraints and equipment thermal safety constraints. Step 2: Construct a dynamic constraint boundary function. The function takes the cumulative damage degree of the equipment and the real-time degradation characteristics as inputs and outputs the maximum allowable temperature and the maximum allowable stress. Step 3: Using offline historical operation data and a digital twin environment, pre-train the initial maintenance strategy network. The network outputs maintenance actions that satisfy the dynamic constraint boundary function. Step 4: Calculate the Lyapunov derivative of the current policy action in real time. If the derivative is greater than 0, solve the minimum correction action through quadratic programming and force the state to drift back to the safe area. Step 5: Dynamically update the allowable stress-temperature coupling relationship in the dynamic constraint boundary function; Step 6: Using dual gradient descent and Lagrange multiplier adaptive mechanism, the risk tolerance coefficient of constraint violation is adjusted in policy iteration so that the long-term constraint violation probability is bounded to the preset safety threshold. Step 7: Output a deployable maintenance strategy. The strategy receives the current status and outputs maintenance actions at each decision point, while generating a constraint compliance confidence report for the corresponding actions.

[0028] Specifically, the core of step one is to abstract the operation and maintenance process of high-temperature key equipment in waste incineration power plants into a constrained Markov decision model with safety constraints. This enables the reinforcement learning agent to uniformly model the dynamic relationship between equipment degradation state, real-time operating conditions, and maintenance actions in a continuous operating environment, and to learn the maintenance decision strategy that maximizes long-term benefits while meeting emission safety constraints and equipment thermal safety constraints.

[0029] A constrained Markov decision model consists of a state space, an action space, state transition relationships, a reward function, and a constraint set. The state space describes the current health status and operating conditions of the equipment in the waste-to-energy plant; the action space describes the maintenance and repair actions that the operation and maintenance system can perform; the reward function quantifies the comprehensive benefits of different maintenance decisions; and the constraint set defines the inviolable safety and environmental boundaries during operation.

[0030] The state space is constructed using a continuous high-dimensional vector form to simultaneously reflect the equipment degradation and evolution process and the real-time thermal state of the incineration system. Since the high-temperature critical equipment in waste-to-energy plants exhibits significant time-varying degradation characteristics, relying solely on a single operating parameter cannot accurately characterize the true health state of the equipment. Therefore, the state space incorporates both equipment degradation characteristics and real-time operating parameters.

[0031] Equipment degradation characteristics are used to characterize the degree of long-term damage accumulation in critical high-temperature equipment, including the cumulative operating time of the incinerator grate, the amount of wall thinning of the waste heat boiler heating surface, and the high-temperature corrosion rate.

[0032] The cumulative runtime of the incinerator grate reflects the degree of aging of the grate after long-term exposure to thermal fatigue and mechanical impact. It can be obtained from the equipment operation log and is defined as follows: In the formula, Indicates the first The cumulative runtime corresponding to each decision moment Indicates the first The duration of each running cycle.

[0033] The reduction in wall thickness of the heating surface of a waste heat boiler reflects the loss of metal materials caused by high-temperature flue gas erosion and chloride corrosion. It is calculated as the difference between the current wall thickness and the initial wall thickness, i.e.: In the formula, Indicates the initial wall thickness of the heated surface. This indicates the measured wall thickness at the current moment.

[0034] High-temperature corrosion rate is used to characterize the rate of material degradation in a high-temperature acidic flue gas environment, and it is obtained based on the rate of change of wall thickness within adjacent testing periods: In the formula, This indicates the high-temperature corrosion rate at the current moment.

[0035] Real-time operating parameters are used to describe the current thermal operating status of the incineration system, including the flue gas temperature at the furnace outlet, the oxygen content in the flue gas, the current of the grate drive motor, and the carbon monoxide concentration in the flue gas.

[0036] The flue gas temperature at the furnace outlet reflects the completeness of waste combustion, directly determining the dioxin suppression effect and the boiler's heat absorption intensity; the oxygen content in the flue gas reflects the combustion air distribution balance; the current of the grate drive motor indirectly reflects the degree of grate jamming and slag accumulation; the carbon monoxide concentration in the flue gas characterizes the degree of incomplete combustion and emission risk. Based on the above variables, in the first... Constructing a complete state vector at each decision moment: In the formula, This indicates the flue gas temperature at the furnace outlet. This indicates the oxygen content of the flue gas. This indicates the current of the grate drive motor. This indicates the concentration of carbon monoxide in the flue gas.

[0037] The action space defines the set of maintenance actions that an agent can perform. Because maintenance activities in waste-to-energy plants have a clear hierarchical structure, the action space is constructed using discrete, multi-level maintenance actions. Different actions correspond to different maintenance intensities, downtime impacts, and maintenance costs.

[0038] The motion space includes the following four types of motion: The non-maintenance state holding action is used to maintain the current operating state; Local water rinsing is used to remove coke and localized ash buildup on heated surfaces. High-temperature anti-corrosion coating spraying repair action is used to reduce the high-temperature chloride salt corrosion rate; A system-wide shutdown and overhaul is used to address severely degraded equipment.

[0039] Therefore, the action space is defined as: in, This indicates that the device remains operational without requiring maintenance. This indicates a localized water rinsing action to remove slag. This indicates the application and repair of a high-temperature anti-corrosion coating. This indicates a system-wide shutdown and major overhaul.

[0040] In the actual decision-making process, the reinforcement learning agent bases its decisions on the current state vector. Output corresponding maintenance actions And obtain the state of the next moment through environmental state transition.

[0041] The reward function is used to quantify the comprehensive economic benefits brought by different maintenance strategies. Its goal is to increase the benefits of continuous equipment operation and reduce maintenance costs while ensuring that safety and environmental protection constraints are met.

[0042] The reward function consists of a power generation revenue term and a maintenance cost term.

[0043] Revenue from electricity generation is calculated based on the real-time grid connection price and the power generation capacity. In the formula, This indicates the real-time on-grid electricity price at the current moment. This indicates the generating capacity of the unit.

[0044] Maintenance costs consist of a weighted average of spare parts costs and labor costs. In the formula, Indicates the cost of spare parts. This indicates the cost of labor hours. and This represents the corresponding weighting coefficient.

[0045] Based on the above definition, the reward function is expressed as: When the system performs a major overhaul, the reward value will be significantly reduced because the power generation drops to zero. Furthermore, when the equipment deteriorates due to prolonged lack of maintenance, it will further increase the subsequent maintenance costs and the risk of constraint violation. Therefore, reinforcement learning agents need to make a dynamic trade-off between continuous operating benefits and equipment health.

[0046] The constraint set is used to define the safety and environmental boundaries that must be strictly met during the operation of a waste-to-energy plant. Since these constraints fall under regulations and equipment lifespan limits, they are embedded in the model as hard constraints.

[0047] Emission safety constraints are used to ensure that the incineration process meets dioxin suppression conditions, and their constraint form is as follows: In the formula, This indicates the temperature of the flue gas at the furnace outlet, in degrees Celsius. This indicates the residence time of flue gas in the high-temperature zone, measured in seconds.

[0048] Equipment thermal safety constraints are used to prevent material failure in high-temperature equipment, including the maximum permissible wall temperature limit on the exhaust side of the incinerator and the maximum permissible creep stress limit for the superheater tubes of the waste heat boiler.

[0049] The grate wall temperature constraint is expressed as: In the formula, This indicates the actual grate wall temperature. This indicates the maximum wall temperature that the current material can withstand.

[0050] The superheater creep stress constraint is expressed as: In the formula, This indicates the current equivalent thermal stress of the superheater tubes. This indicates the allowable creep stress limit of the material.

[0051] During the operation of the constrained Markov decision model, the system checks in real time whether the constraints are satisfied after each action decision made by the reinforcement learning agent. When any constraint is violated, the corresponding state is marked as a high-risk state and enters the subsequent online safety correction layer for safety correction, thereby ensuring that the entire maintenance strategy optimization process always operates within the safe region.

[0052] Specifically, the core of step two lies in establishing a constraint boundary function that can dynamically change with the aging process of the equipment. This allows the reinforcement learning system to no longer use a fixed safety threshold, but to calculate the maximum allowable temperature and maximum allowable stress under the current operating conditions in real time based on the actual damage, corrosion and creep deterioration state of the high-temperature critical equipment. This provides a dynamic safety boundary for subsequent constraint reinforcement learning, online safety correction and risk control.

[0053] Although the constrained Markov decision model constructed in step one has defined emission safety constraints and equipment thermal safety constraints, these constraints are essentially still static constraints. For example, fixed grate wall temperature limits and fixed allowable stress limits only apply to the case where the equipment is in its initial healthy state. However, waste-to-energy plants operate in high-temperature, high-chloride, and high-fly ash scouring environments for extended periods, causing the performance of equipment materials to continuously deteriorate, leading to a continuous contraction of the original safety boundaries.

[0054] For example, in the initial stage of operation, superheater tubes can withstand higher thermal stresses, but as the wall thickness decreases and grain creep expands, their actual stress tolerance gradually decreases; similarly, the maximum allowable wall temperature of the grate will also decrease after long-term thermal fatigue. Therefore, if fixed constraint boundaries are still used, the reinforcement learning agent will mistakenly believe that the equipment still has the original load-bearing capacity, thus creating potential risks of overheating and failure.

[0055] To address this issue, step two introduces a dynamic constraint boundary function to model the current remaining load-bearing capacity of the equipment in real time. The dynamic constraint boundary function adopts a cumulative piecewise nonlinear mapping structure, which consists of two levels of mapping.

[0056] The first-level mapping is used to calculate the material's remaining strength reduction factor based on the equipment's cumulative damage and real-time degradation characteristics.

[0057] The second-level mapping is used to apply the material residual strength reduction factor to the nominal allowable temperature and nominal allowable stress, thereby obtaining the maximum allowable temperature and maximum allowable stress that are actually available at the current moment.

[0058] The entire dynamic constraint boundary function can be expressed as: In the formula, Indicates the first The output of the dynamic constraint boundary function at each decision moment; This indicates the maximum permissible temperature at the current moment; This represents the maximum allowable stress at the current moment.

[0059] The inputs to the dynamic constraint boundary function include the cumulative damage of the equipment and real-time degradation features: In the formula, Indicates the cumulative damage level of the equipment; This indicates the amount of thinning caused by high-temperature corrosion on the exhaust side of the incinerator; This indicates the creep elongation of the superheater tubes in a waste heat boiler.

[0060] The cumulative damage degree of equipment is used to reflect the overall lifespan consumption of materials under long-term thermal cycling and high-temperature stress. Since high-temperature equipment in waste incineration power plants is subject to both thermal fatigue damage and high-temperature creep damage, step two uses a combination of linear Palmgren, Miner's fatigue accumulation rule, and creep time fraction rule to calculate the cumulative damage degree.

[0061] Thermal fatigue damage is represented as follows: In the formula, Indicates the cumulative degree of fatigue damage; Indicates the first The actual number of thermal cycles that occur; This indicates the number of cycles the material can withstand at the corresponding stress amplitude.

[0062] This formula represents the proportional relationship between the actual number of thermal cycles experienced by the equipment and the theoretical lifespan of the material. When When the value approaches 1, it indicates that the material is nearing its fatigue life limit.

[0063] Creep damage is represented using the time fraction rule: In the formula, Indicates the cumulative degree of creep damage; Indicates the device in the Actual operating time under temperature stress conditions; This indicates the theoretical lifespan of the material under the corresponding operating conditions before creep fracture occurs.

[0064] Since fatigue and creep often coexist in high-temperature equipment, the overall cumulative damage degree of the equipment is expressed as: In the formula, and These represent the fatigue damage weight and the creep damage weight, respectively.

[0065] The high-temperature corrosion thinning in the real-time degradation characteristic quantity reflects the cross-sectional loss of the furnace exhaust side material due to chloride corrosion, and its expression is: In the formula, Indicates the initial wall thickness; This indicates the current remaining wall thickness.

[0066] Creep elongation is used to reflect the plastic elongation of superheater tube materials under long-term high-temperature stress, and its expression is: In the formula, Indicates the initial pipe length; Indicates the current supervisor.

[0067] The cumulative damage level and real-time degradation features are input together into the first-level mapping module.

[0068] The first-level mapping uses a piecewise nonlinear degradation model to calculate the material's residual strength reduction factor: In the formula, This represents the reduction factor for the remaining strength of the material, and its value ranges from 0 to 1.

[0069] When the device is in its initial healthy state Approaching 1; as damage accumulates, corrosion intensifies, and creep increases, Continued decline.

[0070] Since material degradation exhibits significant nonlinearity at different life stages, step two uses a piecewise approach to describe the degradation pattern.

[0071] In the mild degradation stage, the material properties decline slowly, and the reduction factor decreases approximately linearly; During the moderate degradation stage, grain boundary oxidation and creep voids expand rapidly, and the deceleration rate increases significantly. In the severe degradation stage, the material enters the accelerated failure zone, and the remaining strength collapses rapidly.

[0072] Therefore, the material residual strength reduction factor is adopted in the following segmented form: In the formula, and Indicates the boundary point of the degradation stage; , , This represents the degradation coefficient.

[0073] After completing the first-level mapping, step two further executes the second-level mapping.

[0074] The purpose of the second-level mapping is to calculate the theoretical ultimate load-bearing capacity of the equipment based on the first principles of thermodynamics, and then combine this with the degree of material degradation to obtain the current real safety boundary.

[0075] The nominal allowable temperature is derived from the thermal balance relationship of the furnace. During waste incineration, the heat released in the furnace is transferred to the heated surfaces through flue gas convection. Therefore, the theoretical allowable temperature of the equipment can be derived by inversely based on the relationship between adiabatic combustion temperature and convective heat transfer.

[0076] The furnace heat balance relationship is expressed as follows: In the formula, Indicates the amount of heat transferred per unit time; Indicates the convective heat transfer coefficient of flue gas; Indicates the heat exchange area; Indicates the flue gas temperature; This indicates the wall temperature of the equipment.

[0077] Based on the relationship between the material design limits and thermal equilibrium of the equipment, the nominal allowable temperature can be obtained: In the formula, This indicates the adiabatic combustion temperature in the furnace.

[0078] The nominal allowable temperature represents the theoretically highest operating temperature that the equipment can operate under ideal health conditions.

[0079] The nominal allowable stress is calculated based on the mechanical relationship of high-temperature pressure vessels.

[0080] Under high-temperature steam pressure, the circumferential thermal stress of the superheater tubes in a waste heat boiler satisfies Lamé's formula: In the formula, Indicates the pressure inside the pipe; Indicates the outer radius; Indicates the inner radius.

[0081] Since the high-temperature creep strength of the material changes with temperature, step two further incorporates the creep strength limit from the material creep fracture test data for interpolation correction.

[0082] Once the endurance strength data points at different temperatures are known, the theoretical allowable stress limit corresponding to the current working condition can be obtained by interpolation.

[0083] Finally, step two dynamically corrects the nominal allowable temperature and nominal allowable stress using the material residual strength reduction factor: This yields the maximum permissible temperature and maximum permissible stress that the current equipment can actually withstand.

[0084] As the cumulative damage to the equipment increases, the reduction factor for the remaining strength of the material continues to decrease, and the maximum allowable temperature and maximum allowable stress output by the dynamic constraint boundary function shrink synchronously. In subsequent strategy optimization, the reinforcement learning agent will no longer make decisions based on fixed safety boundaries, but will dynamically adjust the maintenance strategy based on the equipment's current actual remaining lifespan, thereby achieving safe maintenance decisions for high-temperature critical equipment in waste-to-energy plants under strong time-varying constraints.

[0085] Specifically, the core technology of step three lies in addressing the problem that the constrained Markov decision model established in step one cannot be tested online in practical applications, and further resolving the issue that the dynamic constraint boundary function constructed in step two is difficult to directly participate in high-dimensional policy optimization during the policy learning stage. By introducing offline historical operating data and a high-fidelity digital twin environment, real industrial operating experience is integrated with the physical mechanism simulation space, enabling the reinforcement learning agent to complete policy pre-training without contacting real equipment, while ensuring that the output policy remains safe under the constraints of the dynamic constraint boundary function defined in step two.

[0086] Step 3 consists of three parts: offline data-driven experience modeling, construction of a high-fidelity digital twin environment, and joint training process of offline policy evaluation and context-constrained policy optimization.

[0087] At the data level, offline historical operation data comes from the full-process operation records of at least one complete maintenance cycle of the waste-to-energy plant. This data covers the complete evolution process of the equipment from healthy state, mild degradation, moderate degradation to maintenance recovery in the time dimension, thus providing a sample basis with complete state transition trajectories for reinforcement learning.

[0088] The offline historical operation data specifically includes: time-series data of flue gas temperature at the furnace outlet, which is used to characterize combustion stability and heat release status; time-series data of current of the incinerator grate drive motor, which reflects changes in grate operating resistance and the degree of ash accumulation and jamming; inspection records of wall thickness reduction of the waste heat boiler heating surface, which is used to characterize the material loss process caused by high-temperature corrosion and erosion; and multi-level maintenance action tags executed at corresponding times, which are used to build a mapping relationship between status and maintenance decisions.

[0089] Based on the above data, a standard reinforcement learning training sample quadruple can be constructed: in Corresponding to the state space vector in step one, This is a multi-level maintenance procedure. The difference between the power generation revenue and maintenance costs defined in step one. This represents the device state at the next moment after the state transition.

[0090] At the physical modeling level, a high-fidelity digital twin environment is used to compensate for the sparsity of real data and insufficient coverage of extreme operating conditions. This digital twin environment is jointly constructed based on a finite element thermal-mechanical coupling model and a computational fluid dynamics flue gas flow model, enabling it to simultaneously characterize the stress distribution of the equipment structure and the evolution of the flue gas flow field.

[0091] Among them, the finite element thermal-mechanical coupling model is used to simulate the stress, strain and creep behavior of the grate, heating surface and superheater tube under the combined action of high temperature load and mechanical load; the computational fluid dynamics model is used to simulate the flue gas velocity distribution, temperature field distribution and pollutant diffusion characteristics during waste combustion.

[0092] This high-fidelity digital twin environment uses real-time operating parameters from offline historical operating data as boundary input conditions, including furnace outlet flue gas temperature, flue gas oxygen content, and grate drive motor current, thereby driving the simulation system to reconstruct the real operating trajectory in virtual space.

[0093] At the output level, the digital twin environment not only outputs the device degradation characteristics in the simulation state space, but also, in conjunction with the dynamic constraint boundary function constructed in step two, further outputs the maximum allowable temperature and maximum allowable stress at the corresponding time. This enables the simulation environment to dynamically shrink the constraint boundary as the device ages, thus achieving a safe evolution mechanism consistent with the real degradation process.

[0094] At the policy learning level, step three introduces an offline policy evaluation mechanism, the core of which is the fitted Q-evaluation algorithm. This algorithm trains a state-action-value network using a four-tuple of state, action, reward, and next state from offline historical running data, which is used to approximate the policy's performance in long-term cumulative returns.

[0095] The state and action value function is defined as follows: in Indicates the state Next action The expected cumulative reward after that, This is the discount factor.

[0096] This value network is not only used to evaluate the merits of historical actions, but also to conduct an offline joint evaluation of the security and profitability of the initial maintenance strategy network output actions, so that policy updates do not depend on real-world environment interactions.

[0097] At the strategy optimization level, a context-constrained strategy optimization algorithm is introduced. This algorithm embeds the dynamic constraint boundary function defined in step two on the basis of the standard strategy optimization framework, so that the constraint information can dynamically participate in the strategy update as the device state changes.

[0098] The key to this method lies in the constraint violation probability estimation mechanism. For each policy output action, the system performs simulation prediction in a high-fidelity digital twin environment, and progressively compares the simulation state with the maximum allowable temperature and maximum allowable stress output by the dynamic constraint boundary function.

[0099] The probability of constraint violation is defined as: in Indicates the simulated temperature state. This indicates the simulated stress state.

[0100] During the policy optimization process, a constraint threshold mechanism is introduced. When the probability of constraint violation exceeds the preset offline training constraint threshold, the policy update direction will be suppressed or penalized, thereby ensuring that the policy learning always converges to a safe and feasible region.

[0101] Meanwhile, this context-constrained policy optimization algorithm uses the dynamic constraint boundary function from step two as the context input, enabling the policy optimization process to adaptively adjust constraint weights as the device degradation state changes, thereby avoiding the policy failure problem caused by static constraints in traditional methods.

[0102] At the strategy network structure level, the initial maintenance strategy network adopts a parameterized Gaussian strategy model. Its input is the state space vector defined in step one, including equipment degradation features and real-time operating parameters. The output is the probability distribution parameters of multi-level maintenance actions, namely the mean vector and variance vector.

[0103] The structural form is as follows: in This represents the average action output by the policy network. This represents the action variance matrix.

[0104] This design enables the maintenance strategy to not only output deterministic maintenance actions, but also to express the confidence distribution of maintenance decisions under uncertain operating conditions, thereby enhancing its generalization ability in complex and degraded environments.

[0105] Finally, by continuously performing the joint training process of offline policy evaluation and context-constrained policy optimization in a high-fidelity digital twin environment, the policy network loss function reaches a stable convergence state between maximizing the expected cumulative reward and minimizing the constraint violation probability, thus obtaining the pre-trained initial maintenance policy network.

[0106] The pre-trained policy network has embedded the constrained Markov decision model information from step one, the dynamic constraint boundary function information from step two, and the equipment degradation and evolution law in its structure. This makes its output maintenance actions naturally satisfy the dynamic constraint boundary function constraints, providing a policy foundation with initial security for the online safety correction layer in the subsequent step four.

[0107] Specifically, step four plays a role in the safety closed-loop correction within the overall technical system. Its inputs include the state definition of the constrained Markov decision model from step one, the dynamic constraint boundary function from step two, and the output of the initial maintenance strategy network obtained from pre-training in step three. Its core objective is to correct actions that may exceed the dynamic constraint boundary in real time during the strategy execution phase, so that the system state always remains within the safe area defined by step two that dynamically shrinks as the equipment degrades. This ensures the executable safety of the reinforcement learning strategy in high-temperature, high-corrosion, and highly constrained industrial scenarios.

[0108] In step one, system operation is abstracted as a constrained Markov decision process. The state includes equipment degradation characteristics and real-time operating parameters, and actions consist of multi-level checks. However, this model itself does not guarantee the stepwise safety of the policy execution process; it only provides a constraint description. Step two further transforms the safety boundary from a static form into a time-varying boundary that evolves with equipment damage through a dynamic constraint boundary function. However, this boundary still belongs to the definition of external constraints and does not directly affect the policy execution process. Step three provides the policy with preliminary safety through offline policy evaluation and context constraint optimization. However, due to model errors and distribution offsets between the digital twin and the real system, there is still a risk of constraint boundary approximation or even short-term boundary overflow in local state regions. Therefore, an online safety correction mechanism needs to be introduced at the policy execution end.

[0109] The core idea of ​​the online safety correction layer based on Lyapunov candidate functions is to transform the system safety problem into a stability problem. By constructing a positive definite function about the safety boundary, the trend of the current system state converging towards the safe region is characterized. By constraining the time derivative of this function, real-time correction control of policy actions is achieved.

[0110] The correction layer is based on the constrained Markov decision model in step one, and uses the dynamic constraint boundary function in step two as the safety reference boundary, defining a positive definite Lyapunov candidate function in the state space.

[0111] This function is constructed as a weighted Euclidean distance squared between the current system state and the safe region defined by the dynamic constraint boundary function. Essentially, it measures the degree to which the current operating state deviates from the safe boundary. Its expression is as follows: in This represents the system state vector defined in step one. This represents the projected state of the safe boundary obtained from the derivation of the dynamic constraint boundary function in step two. This represents the weight matrix, used to characterize the degree of influence of different state variables on safety risks. For example, furnace temperature deviation has a higher weight than motor current deviation.

[0112] The Lyapunov candidate function has two key properties: first, positive definiteness, meaning that the function value is zero when the system state is within the safe boundary, and strictly greater than zero when it deviates from the safe region; second, differentiability, which allows it to be used to construct continuous-time derivatives.

[0113] Within each decision step, the online security correction layer executes a real-time correction process, which directly affects the initial maintenance strategy network results output in step three.

[0114] First, the initial maintenance strategy network outputs the original maintenance actions based on the parameterized Gaussian strategy model defined in step three: The original action may be derived from Gaussian distribution sampling or mean output, which may pose potential constraint risks under complex degenerate states.

[0115] Subsequently, the system adjusts according to the current state. and primitive actions By combining the state transition mechanism of the constrained Markov decision model, the dynamic evolution direction of the system under action-induced behavior is calculated, and the time derivative of the Lyapunov candidate function along this dynamic direction is further obtained.

[0116] The derivative is defined as: in This represents the state transition dynamics function implicitly defined in step one, reflecting the changing trends of the equipment's degradation state and operating parameters under the current maintenance action.

[0117] The Lyapunov derivative is used to characterize whether the current policy action will push the system away from the safe zone. When the derivative is less than or equal to zero, it indicates that the system state shows a convergence or stabilization trend under the action of the current action, that is, the state gradually moves towards the interior of the safe zone or remains stable at the boundary. At this time, the online safety correction layer does not intervene in the action and directly outputs the original maintenance action as the final execution action.

[0118] When the Lyapunov derivative is greater than zero, it indicates that the current strategy action will cause the system state to diverge outside the safety boundary, that is, there is a risk of violating the maximum allowable temperature or maximum allowable stress defined by the dynamic constraint boundary function in step two. At this time, the safety correction mechanism is triggered.

[0119] In this case, the online safety correction layer initiates a quadratic programming optimizer to make a minimum correction to the original action, so that the corrected action can both maintain the output characteristics of the policy network as much as possible and meet the safety convergence condition.

[0120] The quadratic programming problem is defined as follows: The objective function is to minimize the deviation between the corrected action and the original action: The physical meaning of this objective function is to maintain the consistency of the policy behavior obtained from training in step three as much as possible, and to avoid policy distortion caused by excessive intervention.

[0121] The constraint condition is that the Lyapunov derivative must not be greater than a negative safety contraction factor: in This represents the safety contraction coefficient, used to ensure that the system state not only does not diverge, but also returns to the safe region at a certain rate, thereby avoiding oscillating behavior near the constraint boundary.

[0122] By solving this quadratic programming problem, the minimum correction action can be obtained: This corrective action is mathematically the safe and feasible solution closest to the original policy output, thus ensuring policy stability while satisfying strict safety constraints.

[0123] Ultimately, the online security correction layer will correct the actions. The output is the final maintenance action, which is applied to the high-temperature critical equipment of the waste-to-energy plant, causing the system state to evolve according to the modified dynamic trajectory.

[0124] At the system level, this process ensures that the state space trajectory is no longer directly driven by the policy in step three, but by a safety control closed loop jointly constrained by the policy output, Lyapunov constraint projection, and dynamic constraint boundary function, thus forming a three-layer safety decision-making mechanism: step three provides an approximately optimal policy, step two provides a dynamic safety boundary, and step four corrects the policy in real time at the execution level through Lyapunov stability constraints.

[0125] The ultimate result is that even under extreme conditions such as increased equipment corrosion due to high temperature, accelerated creep, or sudden changes in flue gas parameters, the system can still force the operating state back to the dynamic safety zone defined in step two through this online safety correction layer, thus realizing the executable safety closed-loop control of the reinforcement learning maintenance strategy in the waste incineration power plant scenario.

[0126] Specifically, step five addresses a key issue built upon the constraint modeling, dynamic boundary construction, offline policy learning, and online safety correction already established in steps one through four. Specifically, the allowable stress and temperature coupling relationship upon which the dynamic constraint boundary function in step two relies still suffers from initial model bias and material aging drift, potentially leading to systemic mismatch in the safety boundary during long-term operation. By introducing an online monitoring-driven Bayesian update mechanism, the dynamic constraint boundary function evolves from a static dynamic model based on prior mechanisms into a closed-loop safety boundary system based on continuous data feedback correction. This forms a two-layer safety assurance structure with the online safety correction layer in step four.

[0127] In step one, the system has been abstracted into a constrained Markov decision model, with safety constraints existing in the form of emission safety constraints and equipment thermal safety constraints. However, in step two, these constraints are further structured into dynamic constraint boundary functions that evolve with equipment degradation. In step four, this dynamic boundary has participated in Lyapunov safety correction to constrain the feasibility of actions. However, with the long-term operation of the waste-to-energy plant, material properties will continuously evolve due to creep, corrosion, and thermal fatigue, causing the allowable stress-temperature coupling relationship originally established based on first principles of thermodynamics and material test data to gradually deviate from the actual state. Therefore, without continuous updates, the safety correction in step four, although formally satisfying the constraints, will have drift errors in the constraints themselves, thereby reducing overall safety and reliability.

[0128] Step 5 addresses the aforementioned issues through the adaptive constraint boundary update module. Its core is to use real online monitoring data to perform parameter-level corrections on the dynamic constraint boundary function in Step 2, enabling it to continuously evolve as the equipment ages, and forming a consistent closed loop with the strategy learning in Step 3 and the safety correction in Step 4.

[0129] The adaptive constraint boundary update module first collects real-time online monitoring data of key high-temperature equipment in the waste-to-energy plant. This data comes from a multi-source heterogeneous sensing system, including infrared temperature sensors on the outer wall of the incinerator grate, thermocouples on the outer wall of the superheater tubes of the waste heat boiler, and a continuous flue gas emission monitoring system.

[0130] Among them, the infrared temperature measurement data of the grate outer wall is used to reflect the changes in local heat load distribution, the temperature of the superheater tube outer wall is used to reflect the actual thermal stress state of the tube wall, and the carbon monoxide concentration and dioxin precursor concentration output by the continuous flue gas monitoring system are used to indirectly characterize the degree of incomplete combustion and the degree of abnormal chemical reaction in the high-temperature zone. These data together constitute the observation input that is consistent with the state-space variables in step one, but at a higher frequency and closer to the actual equipment state.

[0131] Regarding the triggering mechanism, this module does not continuously update all data, but rather uses safety boundary deviation events as the trigger condition. When the measured temperature in the online monitoring data exceeds the maximum allowable temperature output by the dynamic constraint boundary function in step two, the system considers that the current constraint model can no longer fully cover the actual working condition deviation, and at this time, the parameter update mechanism is triggered.

[0132] At the trigger moment, the system simultaneously collects three key pieces of information: cumulative equipment damage, real-time degradation characteristics, and stress-temperature paired observation data points. The cumulative equipment damage is consistent with that in step two and is used to characterize the long-term degradation level. The real-time degradation characteristics include corrosion thinning and creep elongation, used to characterize the current state. The stress-temperature paired data originates from the correspondence between high-temperature loads and structural responses during actual operation, serving as crucial observational evidence for revising allowable boundaries.

[0133] At the model update level, the adaptive constraint boundary update module uses a Bayesian linear regression framework to update the allowable stress and temperature coupling relationship online. The core of this method is that it does not directly replace the original mechanism model, but introduces data-driven posterior corrections based on the original material physics priors, thereby achieving mechanism and data fusion modeling.

[0134] In step two, the coupling relationship between allowable stress and allowable temperature is constructed using first principles of thermodynamics and material creep fracture test data, essentially constituting a priori model. Step five, building upon this, uses the material's endurance strength limit as the prior mean function of the Gaussian process, making this physical prior the foundational structure of Bayesian inference, thus ensuring that the model does not deviate from the physically interpretable range.

[0135] Let the stress-temperature coupling relationship be expressed in functional form: Within the Bayesian linear regression framework, this relationship is parameterized into a linear or weakly nonlinear form and updated based on observed data points: in This represents the observed stress response. The input matrix represents the temperature and degradation characteristics. This represents the coupling parameter to be estimated.

[0136] The posterior distribution can be obtained through Bayesian inference: Then, the posterior mean function and the posterior covariance function are calculated.

[0137] The posterior mean function is used to update the most likely physical model of the allowable stress-temperature coupling relationship, and the posterior covariance function is used to describe the degree of uncertainty of the relationship, thus providing risk margin information for the safety correction in step four.

[0138] During the update result write-back phase, the adaptive constraint boundary update module replaces the original relationship model in step two with the posterior mean function as the new allowable stress-temperature coupling relationship, and writes it back to the dynamic constraint boundary function, making it the basis for the new constraint generation.

[0139] The updated dynamic constraint boundary function recalculates the maximum allowable temperature and maximum allowable stress. Its expression mechanism remains unchanged in the structure of step two, but the source of parameters changes from a static mechanism model to a hybrid model of mechanism prior and online a posteriori correction, thereby realizing the adaptive evolution of the dynamic constraint boundary function.

[0140] In the overall system coordination mechanism, step five forms a strict closed loop relationship with the preceding steps: step one defines the constraint decision structure, step two constructs the dynamic safety boundary, step three learns the initial strategy under the constraint of the boundary, step four performs real-time safety projection at the execution level, and step five continuously corrects the boundary itself in the time dimension, so that the constraint model evolves synchronously with the actual degradation trajectory of the equipment.

[0141] Ultimately, this mechanism transforms the safety constraints of high-temperature critical equipment in waste-to-energy plants from static settings or one-time modeling results into a dynamic learning system that is continuously updated with operational data and gradually shrinks or corrects as materials age. This significantly improves the safety consistency and physical reliability of the overall reinforcement learning maintenance strategy during long-term operation.

[0142] Specifically, step six in the entire technical system is responsible for long-term constraint consistency control and policy convergence stability constraints. Its role is to further solve a key problem based on the constrained Markov decision model, dynamic constraint boundary function, offline pre-trained policy network, Lyapunov online safety correction and Bayesian adaptive boundary update already formed in steps one to five. That is, even if a single-step decision is corrected for safety in step four and updated for boundary in step five, there may still be an accumulation of constraint statistical deviations during the long-term policy iteration process, which makes it impossible to strictly quantify and control the long-term operational safety of the overall system.

[0143] In other words, step four guarantees instantaneous safety, and step five guarantees that boundary correctness is updated over time, but still lacks a global closed-loop control mechanism for the probability of long-term constraint violations. Step six, by introducing dual optimization theory and the Lagrange multiplier dynamic adjustment mechanism, extends the constraint reinforcement learning problem from stepwise safety control to a long-term probabilistic constraint control problem, thereby ensuring that the waste-to-energy plant meets the statistically strict upper bound control of emission safety constraints and equipment thermal safety constraints throughout its entire life cycle operation.

[0144] In step one, the system defines a constrained Markov decision model, where emission safety constraints and equipment thermal safety constraints exist as hard constraints. Step two transforms this into a dynamic constraint boundary function, allowing the constraints to change with equipment degradation. Step three completes offline policy learning under this dynamic boundary. Step four achieves stepwise safety correction using Lyapunov candidate functions. Step five further refines the dynamic boundary function using Bayesian online correction, enabling the constraints themselves to be adaptive. However, the above mechanisms are all local or state-dependent control mechanisms and do not yet provide global optimization constraints for the long-term cumulative behavior of constraint violation frequency.

[0145] Step six involves constructing a dual optimization framework to transform the constraint violation problem into a Lagrange dual problem, thereby explicitly introducing constraint penalty variables during policy updates and achieving closed-loop control over the long-term constraint violation probability.

[0146] During the strategy iteration process, before each update of the strategy network parameters, the system first uses the initial maintenance strategy network output in step three and the final action trajectory after correction in step four to count the constraint violations in the most recent batch of offline historical running data or online interaction trajectories.

[0147] The constraint violation judgment is based on the coupling relationship between the dynamic constraint boundary function in step two and the updated allowable stress and temperature in step five, and also combines the final execution action after Lyapunov correction in step four to determine whether the system state has exceeded the following conditions: Whether emission safety constraints are met, i.e. whether the flue gas temperature and residence time at the furnace outlet meet the dioxin suppression conditions; whether equipment thermal safety constraints are met, i.e. whether the grate wall temperature and superheater creep stress exceed the dynamic allowable boundary.

[0148] Based on this, define the cumulative constraint violation frequency: in This indicates the frequency of constraint violations by the current policy within the statistical window. Indicates the number of samples. This indicates an indicator function that takes a value of 1 when a constraint is violated and 0 otherwise. This represents the constraint detection function defined by steps two through five.

[0149] This metric essentially characterizes the level of security of a strategy over the long term, rather than its security in a single step.

[0150] At the optimization level, a dual gradient descent mechanism is introduced to maintain the Lagrange multiplier variable, which is used to dynamically penalize constraint violations. The Lagrange multiplier is denoted as... Its update process is based on the deviation between the constraint violation frequency and the preset safety threshold.

[0151] The specific update logic is as follows: the system first calculates the difference between the current constraint violation frequency and the preset safety threshold. This difference reflects the degree of deviation of the current strategy from the safety requirements. Then, the deviation is multiplied by the dual step size to control the update rate. Finally, it is accumulated to the current Lagrange multiplier variable and subjected to non-negative projection processing to ensure that its physical meaning is non-negative.

[0152] This process can be represented as: in Indicates the dual step size, This indicates a preset safety threshold.

[0153] The Lagrange multiplier acts as a dynamic safety penalty intensity regulator in the system. When the system violates constraints for an extended period... An increase indicates a stronger security penalty; when the system is operating stably in a safe region, Decrease indicates that higher-yield exploration is allowed.

[0154] Furthermore, the Lagrange multipliers are mapped to risk tolerance coefficients, which characterize the system's tolerance for constraint violations. These coefficients are defined as a monotonically decreasing function of the Lagrange multipliers: The function satisfies the following property: when When the coefficient approaches zero, the risk tolerance coefficient approaches 1, indicating that the system allows for a higher degree of freedom in strategy exploration; when When the value approaches infinity, the risk tolerance coefficient approaches 0, indicating that the system strictly suppresses all potentially high-risk actions.

[0155] The risk tolerance coefficient is further embedded in the reinforcement learning reward structure defined in step one, and the original reward function is Lagrange-corrected.

[0156] The original reward function is defined in step one as power generation revenue minus maintenance costs, and is expanded in step six to: The first item maintains an economic benefit orientation, while the second item introduces a penalty for violating constraints, so that the strategy optimization process can take into account both economy and security.

[0157] The Lagrange reward function serves as the unified optimization objective for the initial maintenance of the policy network and subsequent policy updates in step three. This transforms policy learning from solely relying on maximizing expected returns into a dual optimization problem of maximizing returns and minimizing constraint violations.

[0158] During the policy update process, the system utilizes dual gradient descent to achieve joint convergence on two levels: on the one hand, the initial maintenance policy network parameters are updated through policy gradient, so that they gradually approach the optimal policy under the current Lagrange reward function; on the other hand, the dynamic update mechanism of Lagrange multipliers is used to make the exponential moving average of the constraint violation frequency converge to below the preset safety threshold, thereby ensuring long-term operational safety.

[0159] The difference between step four and step six is ​​that step four controls the safety of a single-step state, ensuring instantaneous non-divergence through the Lyapunov function; step five controls the accuracy of the constraint boundaries, ensuring model consistency through Bayesian updates; while step six controls long-term statistical safety, ensuring the global boundedness of the constraint violation probability over time through dual optimization.

[0160] Ultimately, step six, together with steps one through five, constitutes a three-layer safety control system: the bottom layer provides structural constraints for the constrained Markov decision model; the middle layer ensures the authenticity of the physical boundary through dynamic constraint boundary functions and Bayesian updates; the upper layer ensures instantaneous safe execution through Lyapunov correction; and the outermost layer ensures that the probability of long-term constraint violation is controlled through the dual gradient descent mechanism. This achieves safe, stable, and controllable optimization of the maintenance strategy for waste incineration power plant equipment throughout its entire life cycle.

[0161] Specifically, step seven builds upon the complete reinforcement learning security decision-making system established in steps one through six, constructing an executable strategy output and a credible security interpretation mechanism for the engineering deployment phase. Its core objective is no longer strategy optimization itself, but rather to solidify the strategy, which has undergone multiple security constraints and optimizations, into an online, industrial-grade decision-making system, and simultaneously provide an interpretable constraint compliance confidence assessment, thereby meeting the unified requirements of executability, interpretability, and auditability for waste incineration power plants in high-risk scenarios.

[0162] Logically, the process involves several steps: Step 1 involves constructing a constrained Markov decision model to provide an abstract decision structure for the system; Step 2 involves defining a dynamic constraint boundary function to ensure the safety boundary evolves with equipment degradation; Step 3 involves pre-training the initial maintenance strategy network using offline historical operating data and a high-fidelity digital twin environment; Step 4 involves implementing online safety correction using Lyapunov candidate functions to ensure safe single-step execution; Step 5 involves updating the dynamic constraint boundary function using Bayesian linear regression to enable the safety boundary to have data-driven adaptive capabilities; and Step 6 involves ensuring the bounded probability of long-term constraint violations using dual gradient descent and Lagrange multiplier mechanisms. Step 7, building upon the above six steps, transforms the strategy system into an engineering-deployable form and introduces a probabilistic safety and reliability output mechanism.

[0163] The deployable maintenance strategy consists of two core sub-modules: a target strategy network with frozen parameters and a lightweight constraint confidence assessment network, which respectively undertake the functions of decision execution and security reliability assessment.

[0164] The target policy network with frozen parameters originates from the final output of the initial maintenance policy network in step three, after being corrected by the online security correction layer in step four. Parameter freezing is performed during the deployment phase, meaning the network parameters are no longer updated with gradients or adjusted in actual operation. The essence of this design is to solidify the complete learning results formed in steps one through six into a deterministic policy mapping function, enabling it to perform stable and repeatable execution in real-world industrial environments.

[0165] The target policy network still inherits the parameterized Gaussian policy structure in step three in form. Its input is the state space variables defined in step one, including equipment degradation characteristics and real-time operating parameters. Its output is the deterministic mapping or distributed parameters of multi-level maintenance actions. However, after the constraint mechanisms in steps four and six are applied, its output is already within the safe and feasible policy space.

[0166] Therefore, the target freezing policy network can be represented as: in This represents the final set of policy parameters after step three (learning), step four (correction), and step six (long-term constraint optimization).

[0167] The second module, running parallel to the target policy network, is a lightweight constrained confidence evaluation network. Its role is to evaluate the safety probability of each policy output maintenance action, thereby solving the problem that traditional reinforcement learning systems only output actions but cannot quantify whether the actions are safe and reliable.

[0168] The constraint compliance confidence assessment network adopts a Bayesian neural network structure. Its essence is to introduce probability distribution modeling at the neural network parameter level, so that the model output is no longer a single definite value, but a posterior distribution estimate of the probability of constraint satisfaction.

[0169] The network is in its current system state. As input, it also receives maintenance actions from the network output of the target freezing policy. It outputs the probability value that the action satisfies both emission safety constraints and equipment thermal safety constraints under the current dynamic constraint boundary function conditions.

[0170] This probability is defined as: in and These correspond to the emission safety constraints and equipment thermal safety constraints detection functions defined in step one, respectively. The constraint determination boundary is provided by the dynamic constraint boundary function in step two, and this boundary has been dynamically corrected in step five through a Bayesian update mechanism.

[0171] During the network training phase, the lightweight constrained confidence evaluation network is pre-trained using a posterior probability distillation mechanism, with its training data sourced from the high-fidelity digital twin environment in step three and offline historical running data.

[0172] In this process, each decision step in the simulation trajectory includes the state, action, and constraint detection results. The constraint detection results are represented in binary form to indicate whether the current action violates the dynamic constraint system defined in steps two to five.

[0173] The essence of this training process is to use the constraint results in the physical simulation environment as a supervision signal, so that the Bayesian neural network learns the mapping relationship from the state space to the constraint satisfaction probability space, thereby obtaining the probabilistic expression ability of security under complex degenerate states.

[0174] During the deployment phase, the lightweight constrained confidence assessment network no longer relies on offline data, but continuously receives online monitoring data defined in step one, including information such as temperature, stress, corrosion and emission concentration, and updates the internal Bayesian parameters online through a sliding window mechanism.

[0175] The purpose of this mechanism is to enable the confidence assessment model to adjust its probability distribution synchronously with the equipment aging process, so that the output constraint compliance probability not only reflects the current state, but also implies the long-term degradation trend of the equipment, thus maintaining consistency with the dynamic constraint boundary function in step five and the long-term constraint probability control mechanism in step six.

[0176] At the system output level, the deployable maintenance strategy outputs two results at each decision moment: one is the maintenance action given by the target strategy network, and the other is the constraint compliance confidence report generated by the lightweight constraint confidence assessment network.

[0177] Among them, the constraint compliance confidence report not only provides a single probability value, but also implies the comprehensive confidence level that the action satisfies all safety constraints under the current dynamic constraint boundary function conditions. This confidence level is simultaneously affected by the dynamic boundary in step two, the online update in step five, and the long-term constraint statistical control in step six.

[0178] When the probability of meeting the constraint is higher than the preset confidence threshold, the system considers the current strategy output to have reached an acceptable balance between safety and economy, and the maintenance action can be executed directly. When the probability is lower than the preset threshold, the system will not execute the strategy directly, but will trigger a manual review and warning mechanism to mark the action as a high-uncertainty and high-risk decision, which will be further confirmed by the operation and maintenance personnel.

[0179] Through this mechanism, step seven transforms the system from a policy output system to a security decision-making and credible interpretation system, enabling the entire reinforcement learning framework to possess three capabilities in the high-risk industrial scenario of waste-to-energy plants: first, the executable security decision-making capability formed by steps one through four; second, the dynamic constraint adaptation and long-term risk control capability formed by steps five and six; and third, the probabilistic credibility interpretation and human auditable interface capability provided by step seven. This completes a closed-loop reinforcement learning maintenance strategy system for actual engineering deployment.

[0180] Example 2 The present invention also discloses a system for optimizing maintenance strategies of waste incineration power plant equipment based on reinforcement learning, comprising: a system modeling module, establishing a constrained Markov decision model, a state space including equipment degradation features and real-time operating parameters, an action space including multi-level maintenance actions, a reward function including power generation revenue minus maintenance costs, and a constraint set including emission safety constraints and equipment thermal safety constraints; The dynamic constraint module constructs dynamic constraint boundary functions. The functions take the cumulative damage degree of the equipment and the real-time degradation characteristics as inputs and output the maximum allowable temperature and the maximum allowable stress. The strategy pre-training module uses offline historical running data and a digital twin environment to pre-train the initial maintenance strategy network. The network outputs maintenance actions that satisfy the dynamic constraint boundary function. The Lyapunov correction module calculates the Lyapunov derivative of the current policy action in real time. If the derivative is greater than 0, it solves the minimum correction action through quadratic programming and forces the state to drift back to the safe area. The Bayesian update module dynamically updates the allowable stress-temperature coupling relationship in the dynamic constraint boundary function; The dual optimization module employs dual gradient descent and Lagrange multiplier adaptive mechanism to adjust the risk tolerance coefficient of constraint violation during policy iteration, so that the long-term constraint violation probability is bounded to a preset safety threshold. The deployment decision module outputs deployable maintenance strategies. At each decision moment, the strategy receives the current status and outputs maintenance actions, while generating a constraint compliance confidence report for the corresponding actions.

[0181] The modules in Embodiment 2 are used to implement the functions in Embodiment 1. This embodiment can be implemented by a system including a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement the reinforcement learning-based optimization method for equipment maintenance and repair in waste incineration power plants according to Embodiment 1 of this application. The system also includes other components well-known to those skilled in the art, such as a communication bus and a communication interface. Their settings and functions are known in the art and will not be described in detail here.

[0182] In this application, the aforementioned memory can be any tangible medium containing or storing a program that can be used or combined with an instruction execution system, system, or device. For example, a computer-readable storage medium can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc., or any other medium that can be used to store required information and can be accessed by an application, module, or both. Any such computer storage medium can be part of a device or accessible to or connected to a device. Any application or module described in this application can be implemented using computer-readable / executable instructions that can be stored or otherwise retained by such a computer-readable medium.

[0183] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.

Claims

1. A method for optimizing maintenance strategies for waste-to-energy incineration plant equipment based on reinforcement learning, characterized in that, Includes the following steps: Step 1: Establish a constrained Markov decision model. The state space includes equipment degradation characteristics and real-time operating parameters, the action space includes multi-level maintenance actions, the reward function includes power generation revenue minus maintenance costs, and the constraint set includes emission safety constraints and equipment thermal safety constraints. Step 2: Construct a dynamic constraint boundary function. The function takes the cumulative damage degree of the equipment and the real-time degradation characteristics as inputs and outputs the maximum allowable temperature and the maximum allowable stress. Step 3: Using offline historical operation data and a digital twin environment, pre-train the initial maintenance strategy network. The network outputs maintenance actions that satisfy the dynamic constraint boundary function. Step 4: Calculate the Lyapunov derivative of the current policy action in real time. If the derivative is greater than 0, solve the minimum correction action through quadratic programming and force the state to drift back to the safe area. Step 5: Dynamically update the allowable stress-temperature coupling relationship in the dynamic constraint boundary function; Step 6: Using dual gradient descent and Lagrange multiplier adaptive mechanism, the risk tolerance coefficient of constraint violation is adjusted in policy iteration so that the long-term constraint violation probability is bounded to the preset safety threshold. Step 7: Output a deployable maintenance strategy. The strategy receives the current status and outputs maintenance actions at each decision point, while generating a constraint compliance confidence report for the corresponding actions.

2. The method for optimizing maintenance strategies of waste-to-energy incineration plant equipment based on reinforcement learning according to claim 1, characterized in that, Step one includes: The equipment degradation characteristics in the state space of the constrained Markov decision model include the cumulative running time of the incinerator grate, the reduction in wall thickness of the waste heat boiler heating surface, and the high-temperature corrosion rate. Real-time operating parameters in the state space include furnace outlet flue gas temperature, flue gas oxygen content, grate drive motor current, and carbon monoxide concentration in the flue gas. The multi-level maintenance actions in the action space include non-maintenance state maintenance actions, local water flushing and descaling actions, high-temperature anti-corrosion coating spraying and repair actions, and system-wide shutdown and overhaul actions. The power generation revenue in the reward function is calculated based on the product of the real-time grid-connected electricity price of the waste-to-energy plant and the power generation capacity. The maintenance cost in the reward function is calculated based on the weighted sum of the spare parts cost and labor cost corresponding to different actions in the multi-level maintenance action. The emission safety constraints are centralized, namely, the flue gas temperature at the furnace outlet is not less than 850 degrees Celsius and the flue gas residence time is not less than 2 seconds. The equipment thermal safety constraints are centralized, including the maximum allowable wall temperature limit on the exhaust side of the incinerator and the maximum allowable creep stress limit on the superheater tubes of the waste heat boiler.

3. The method for optimizing maintenance strategies of waste-to-energy incineration plant equipment based on reinforcement learning according to claim 1, characterized in that, Step two includes: The dynamic constraint boundary function adopts a nonlinear mapping relationship in the form of cumulative piecewise segments. The first-level mapping maps the cumulative damage degree of the equipment and the real-time degradation characteristic quantity together into the material residual strength reduction coefficient. The cumulative damage degree of the equipment is calculated jointly according to the linear Palmgren-Miner fatigue accumulation law and the creep time fraction law. The real-time degradation characteristic quantity includes the high-temperature corrosion thinning amount on the exhaust side of the incinerator and the creep elongation amount of the superheater tube of the waste heat boiler. The second-order mapping in the dynamic constraint boundary function multiplies the material residual strength reduction factor with the nominal allowable temperature and nominal allowable stress calculated based on the first principle of thermodynamics, to obtain the maximum allowable temperature and maximum allowable stress at the current moment, respectively. The nominal allowable temperature is derived in reverse from the furnace adiabatic combustion temperature and the flue gas convective heat transfer coefficient, and the nominal allowable stress is determined by interpolation of the endurance strength limit from the Lamé formula and the material creep fracture test data.

4. The method for optimizing maintenance strategies of waste-to-energy incineration plant equipment based on reinforcement learning according to claim 1, characterized in that, Step three includes: Offline historical operation data includes time-series data of flue gas temperature at furnace outlet, time-series data of current of incinerator grate drive motor, inspection records of wall thickness reduction of waste heat boiler heating surface, and labels of multi-level maintenance actions performed at the corresponding time during at least one complete maintenance cycle of the waste incineration power plant. The digital twin environment is jointly constructed based on the finite element thermo-mechanical coupling model and the computational fluid dynamics flue gas flow model. The digital twin environment takes the real-time operating parameters in the offline historical operating data as the input boundary conditions and outputs the equipment degradation characteristics and dynamic constraint boundary function values ​​in the simulation state space. The offline strategy evaluation adopts the fitted Q evaluation algorithm. The fitted Q evaluation algorithm uses the state-action-reward-next state quadruple from the offline historical running data to train a state-action value network. The state-action value network is used to evaluate the expected cumulative reward of the maintenance action output by the initial maintenance strategy network in the digital twin environment. The context constraint policy optimization algorithm adopts the constraint policy optimization algorithm framework. In the constraint update step, the estimated value of the constraint violation probability is obtained by comparing the dynamic constraint boundary function with the simulation state output by the digital twin environment in real time. When optimizing the parameters of the initial maintenance policy network, the context constraint policy optimization algorithm forces the simulation constraint violation probability corresponding to the maintenance action output by the initial maintenance policy network to be lower than the preset offline training constraint threshold. The initial maintenance strategy network adopts a parameterized Gaussian strategy network structure. The initial maintenance strategy network takes the equipment degradation features and real-time operating parameters in the state space as inputs, and outputs the mean and variance of the multi-level maintenance action combination. The pre-trained initial maintenance strategy network is obtained by executing the constrained strategy optimization algorithm in the digital twin environment until the strategy network loss function converges.

5. The method for optimizing equipment maintenance strategies in waste-to-energy incineration plants based on reinforcement learning according to claim 1, characterized in that, Step four includes: The online safety correction layer based on Lyapunov candidate functions constructs a positive definite Lyapunov candidate function with emission safety constraints and equipment thermal safety constraints in the constrained Markov decision model as the boundary. The positive definite Lyapunov candidate function is defined as the square of the weighted Euclidean distance between the current system state and the safety region boundary defined by the dynamic constraint boundary function. The online safety correction layer based on Lyapunov candidate functions performs the following computational process within each decision step: First, the original maintenance action output by the initial maintenance strategy network is received. Then, based on the current system state, the time derivative of the positive definite Lyapunov candidate function induced by the original maintenance action is calculated to obtain the Lyapunov derivative. If the Lyapunov derivative is negative or zero, the online safety correction layer directly outputs the original maintenance action as the final maintenance action. If the Lyapunov derivative is positive, the online safety correction layer starts the quadratic programming solver. The quadratic programming solver takes minimizing the L2 norm square between the correction action and the original maintenance action as the objective function and the safety shrinkage coefficient that forces the Lyapunov derivative to be no greater than negative as the constraint condition to solve for the minimum correction action. The online safety correction layer outputs the minimum correction action as the final maintenance action and applies the final maintenance action to the high-temperature critical equipment of the waste incineration power plant, causing the system state to drift back to the safe region defined by the dynamic constraint boundary function along the direction of the decreasing Lyapunov function value.

6. The method for optimizing maintenance strategies for waste-to-energy incineration plant equipment based on reinforcement learning according to claim 1, characterized in that, Step five includes: The adaptive constraint boundary update module collects real-time online monitoring data of key high-temperature equipment in waste incineration power plants. The online monitoring data includes the actual temperature measured by infrared temperature sensors on the outer wall of the incinerator grate, the actual temperature measured by thermocouples on the outer wall of the superheater tubes of the waste heat boiler, and the carbon monoxide concentration and dioxin precursor concentration measured by the continuous flue gas emission monitoring system. The adaptive constraint boundary update module uses the event that the measured temperature in the online monitoring data exceeds the maximum allowable temperature output by the dynamic constraint boundary function as the trigger condition, and collects the cumulative damage degree of the equipment, real-time degradation characteristics, and measured stress-temperature paired data points at the trigger time; The adaptive constraint boundary update module uses a Bayesian linear regression framework to update the allowable stress-temperature coupling relationship online. This includes using the material's endurance strength limit as the prior mean function of the Gaussian process, using measured stress-temperature paired data points as observation samples, and obtaining the updated posterior mean function and posterior covariance function of the allowable stress-temperature coupling relationship by calculating the mean and variance of the posterior distribution. The adaptive constraint boundary update module uses the posterior mean function as the updated allowable stress-temperature coupling relationship and writes it back into the dynamic constraint boundary function. The dynamic constraint boundary function recalculates the maximum allowable temperature and maximum allowable stress at the current moment based on the updated allowable stress-temperature coupling relationship, so as to realize the closed-loop adaptive adjustment of the dynamic constraint boundary function as the equipment ages.

7. The method for optimizing maintenance strategies for waste-to-energy incineration plant equipment based on reinforcement learning according to claim 1, characterized in that, Step six includes: Before each update of the strategy network parameters during the strategy iteration process, the cumulative constraint violation frequency of the current initial maintenance strategy network on the most recent batch of offline historical running data or online interaction trajectory is calculated. The cumulative constraint violation frequency is the ratio of the total number of decision steps that violate emission safety constraints or equipment thermal safety constraints to the total number of decision steps. A dual gradient descent method is used to maintain a Lagrange multiplier variable. This Lagrange multiplier variable is used to penalize deviations where the cumulative constraint violation frequency exceeds a preset safety threshold. The update rule for the Lagrange multiplier variable is as follows: Subtract the preset safety threshold from the cumulative constraint violation frequency, multiply the difference by the dual step size and accumulate it to the Lagrange multiplier variable, and finally project the Lagrange multiplier variable to the non-negative real number field. The risk tolerance coefficient is defined as a monotonically decreasing function of the Lagrange multiplier variable. The monotonically decreasing function restricts the risk tolerance coefficient to the open interval between 0 and 1. When the Lagrange multiplier variable approaches infinity, the risk tolerance coefficient approaches 0, and when the Lagrange multiplier variable approaches 0, the risk tolerance coefficient approaches 1. The risk tolerance coefficient is embedded in the Lagrange reward function of the constrained Markov decision process. The Lagrange reward function is equal to the reward function minus the product of the Lagrange multiplier variable and the constraint violation indicator function. The Lagrange reward function is then used as the optimization objective for updating the network parameters of the initial maintenance strategy. By adaptively adjusting the Lagrange multiplier variable, the exponential moving average of the long-term cumulative constraint violation frequency is bounded to a preset safety threshold.

8. The method for optimizing maintenance strategies of waste-to-energy incineration plant equipment based on reinforcement learning according to claim 1, characterized in that, Step seven includes: The deployable maintenance strategy consists of a target strategy network with frozen parameters and a lightweight constraint confidence evaluation network. The target strategy network with frozen parameters is obtained by copying the final maintenance action output parameters of the initial maintenance strategy network after it has been corrected by the online safety correction layer. The target strategy network with frozen parameters is not updated with gradients during the deployment phase. The constraint compliance confidence report is generated in real time by the lightweight constraint confidence assessment network at each decision moment. The lightweight constraint confidence assessment network adopts a Bayesian neural network architecture, takes the current system state as input, and outputs the constraint compliance probability value corresponding to the maintenance action output by the target policy network with frozen parameters. The constraint compliance probability value is the predicted probability that the maintenance action simultaneously meets the emission safety constraint and the equipment thermal safety constraint under the current value of the dynamic constraint boundary function. Before deployment, the lightweight constraint confidence assessment network is trained using offline historical operating data and simulation trajectories generated by the digital twin environment for posterior probability distillation. The training label is a binary indicator vector indicating whether the maintenance action at each decision step in the simulation trajectory actually violates emission safety constraints and equipment thermal safety constraints. The lightweight constraint confidence assessment network continuously receives online monitoring data from the device during deployment and updates the posterior parameters of the internal Bayesian neural network online using a sliding window approach, so that the constraint compliance probability value is dynamically adjusted as the device ages. The deployable maintenance strategy displays the maintenance actions output by the target strategy network with frozen parameters and the constraint compliance confidence report output by the lightweight constraint confidence assessment network at each decision moment. When the constraint compliance probability value in the constraint compliance confidence report is lower than the preset confidence threshold, the deployable maintenance strategy simultaneously issues a manual review warning signal.

9. A method for optimizing maintenance strategies for waste-to-energy incineration plant equipment based on reinforcement learning, characterized in that, include: The system modeling module establishes a constrained Markov decision model. The state space includes equipment degradation characteristics and real-time operating parameters, the action space includes multi-level maintenance actions, the reward function includes power generation revenue minus maintenance costs, and the constraint set includes emission safety constraints and equipment thermal safety constraints. The dynamic constraint module constructs dynamic constraint boundary functions. The functions take the cumulative damage degree of the equipment and the real-time degradation characteristics as inputs and output the maximum allowable temperature and the maximum allowable stress. The strategy pre-training module uses offline historical running data and a digital twin environment to pre-train the initial maintenance strategy network. The network outputs maintenance actions that satisfy the dynamic constraint boundary function. The Lyapunov correction module calculates the Lyapunov derivative of the current policy action in real time. If the derivative is greater than 0, it solves the minimum correction action through quadratic programming and forces the state to drift back to the safe area. The Bayesian update module dynamically updates the allowable stress-temperature coupling relationship in the dynamic constraint boundary function; The dual optimization module employs dual gradient descent and Lagrange multiplier adaptive mechanism to adjust the risk tolerance coefficient of constraint violation during policy iteration, so that the long-term constraint violation probability is bounded to a preset safety threshold. The deployment decision module outputs deployable maintenance strategies. At each decision moment, the strategy receives the current status and outputs maintenance actions, while generating a constraint compliance confidence report for the corresponding actions.