A thermal power unit AGC instruction optimization method based on deep reinforcement learning
Patent Information
- Application Number
- CN202610621516.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-08
- Publication Date
- 2026-09-11
AI Technical Summary
[0004]然而,上述现有方法在面对深度调峰工况下高度动态且不确定的AGC指令序列时,存在由调控机理所引发的局限性;具体而言,固定限幅方法中的变负荷速率限值属于静态参数,无法根据转子当前的实时热力学状态进行动态适配:当转子温度场分布均匀、应力裕度充足时,固定限值过于保守,牺牲了本可实现的调节速率;当转子已处于高应力水平时,固定限值又可能因缺乏前瞻性而未能提前收紧,导致明显的应力越限风险
[0074]1. This invention embeds a deep reinforcement learning agent that integrates a thermal stress reduction model between the scheduling command and the coordination control system. This agent explicitly quantifies rotor fatigue life loss as a penalty term in the reward function, ensuring that the load command generation process simultaneously considers both minimizing tracking deviation and minimizing life loss. During training, the agent autonomously learns to proactively avoid load change patterns that may cause severe thermal stress while meeting the grid AGC assessment requirements, thereby effectively extending rotor service life without sacrificing regulation quality.
Smart Images

Figure CN122732289A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of thermal power generation technology, and in particular relates to an AGC instruction optimization method for thermal power units based on deep reinforcement learning. Background Technology
[0002] With the continuous increase in the proportion of new energy power generation, thermal power units, as the main power source for grid peak shaving and frequency regulation, need to frequently respond to automatic grid generation control commands to perform large-scale and high-speed load adjustments. Under deep peak shaving conditions, this operating mode causes the high-pressure rotor to experience drastic steam temperature changes, resulting in large fluctuations in thermal stress on the rotor surface and inside. According to the law of fatigue cumulative damage, each stress cycle causes irreversible life loss to the rotor. As the core high-temperature pressure-bearing component of the steam turbine, the rotor has extremely high replacement costs and long maintenance cycles.
[0003] Currently, solutions to the aforementioned contradictions mainly fall into two categories: one is the experience-based limiting method based on operating procedures, which sets fixed limits on the rate of change of load or the rate of change of main steam temperature in the unit's coordinated control system to prevent excessively rapid load changes from causing thermal stress to exceed limits; the tuning of the thermal stress limit is usually based on offline finite element analysis or simplified theoretical formulas and is solidified into the control logic during the unit commissioning phase. The other is the optimization method based on model predictive control, which establishes a simplified analytical model of rotor thermal stress, uses the thermal stress amplitude as a constraint condition for the optimization problem, and solves for the optimal load command sequence that satisfies the stress limits within a rolling optimization framework.
[0004] However, the aforementioned existing methods have limitations stemming from their control mechanisms when dealing with highly dynamic and uncertain AGC command sequences under deep peak-shaving conditions. Specifically, the variable load rate limit in the fixed-limit method is a static parameter and cannot be dynamically adapted to the rotor's current real-time thermodynamic state. When the rotor temperature field is uniform and the stress margin is sufficient, the fixed limit is too conservative, sacrificing the achievable regulation rate. When the rotor is already at a high stress level, the fixed limit may fail to tighten in advance due to a lack of foresight, leading to a significant risk of stress exceeding the limit. Although model predictive control methods can theoretically treat stress as a dynamic constraint, their practical application is limited by the accuracy of simplified stress models. These simplified models typically only consider the one-dimensional heat conduction process of the rotor's inner wall and cannot accurately describe the propagation lag effect of steam excitation in the rotor's axial and radial two-dimensional temperature fields, nor can they perceive two-dimensional distribution characteristics such as rotor inlet and outlet temperature differences and local stress concentrations.
[0005] More importantly, neither of the above two methods takes long-term cumulative life loss as the objective function of closed-loop control, and still falls within the limit protection range, that is, it allows rotor stress to fluctuate significantly below the alarm value. However, such frequent fluctuations within the allowable range will still generate a non-negligible cumulative life loss according to Miner's linear cumulative damage law. Therefore, the following solution is proposed to address the above problems. Summary of the Invention
[0006] To solve the above-mentioned technical problems, the present invention is achieved through the following technical solution:
[0007] This invention relates to an AGC instruction optimization method for thermal power units based on deep reinforcement learning, comprising the following steps:
[0008] Step S1: Construct a reduced-order thermal stress model of the high-pressure rotor of the unit. Using the rotor material and geometric parameters, through simulation and intrinsic orthogonal decomposition, establish a reduced-order state space model with steam temperature excitation as input and thermal stress as output, which is used to replace the finite element model to realize real-time thermal stress calculation.
[0009] Step S2: Real-time acquisition of AGC commands, actual unit power generation, steam temperature after regulating stage and rotor metal temperature measurements; construction of a reduced-order state observer using a discretized reduced-order model; estimation of the reduced-order state vector inside the rotor and calculation of the current thermal stress value by fusing the reduced-order model with the metal temperature measurements.
[0010] Step S3: Construct the state space of the deep reinforcement learning agent, and combine the power deviation, load change rate, steam temperature, the reduced state vector, the current thermal stress value and the historical AGC instruction sequence into a normalized state vector so that the agent can perceive the load adjustment requirements and the rotor thermodynamic state.
[0011] Step S4: Define the action space of the intelligent agent as the continuous correction increment of the load setpoint, generate an optimized load command based on the action output and the current actual power generation, and after processing the deviation limit link from the original AGC command, send it to the unit coordination control system.
[0012] Step S5: Establish a reward function that integrates lifetime loss perception. The reward function consists of four parts: a tracking quality penalty term, an equivalent lifetime loss rate penalty term based on the fatigue characteristics of rotor materials, a thermal stress penalty term for forward prediction using the reduced-order model, and an action smoothing penalty term.
[0013] Step S6: Construct the Actor network and Critic network using the deep deterministic strategy gradient algorithm. Embed the thermal stress features predicted by the reduced-order model for the current action into the input of the Critic network. Introduce a thermal stress limit violation penalty term into the value update objective of the Critic network. Through the dual embedding of the physical model and reinforcement learning, force the learned load adjustment strategy to meet the rotor thermal stress life constraint.
[0014] Step S7: Build a simulation environment using the unit's historical operating data and the reduced-order model, perform offline pre-training on the Actor network and Critic network, and deploy the Actor network online after training. Generate optimized instructions in each control cycle and configure over-limit safety monitoring logic. When thermal stress is detected to be continuously exceeding the limit, automatically restore the original AGC instruction path.
[0015] Further, step S1 includes the following steps: Step S11: Obtain the geometric dimensions and material properties of the high-pressure rotor, such as density, specific heat capacity, thermal conductivity, elastic modulus, and coefficient of thermal expansion. Establish an axisymmetric finite element model and use the step and ramp excitation of steam temperature as boundary conditions to perform simulation and obtain the time series data of transient temperature field and thermal stress field. Step S12: To transform the computationally complex finite element model into a simplified model that can run online in real time, the spatial modes and time coefficients with the highest energy proportions in the first r orders are extracted from the simulation data using intrinsic orthogonal decomposition. A reduced-order state-space model of the high-pressure rotor thermal stress is established. The reduced-order state-space model directly calculates the characterization value of the rotor inner wall thermal stress through state equations and output equations, as shown in the following formula: In the formula, This is an r-dimensional reduced-order state vector, which physically corresponds to the main modal coordinates of the temperature field; The equivalent steam temperature excitation acting on the rotor surface is characterized by the steam temperature after the regulating stage through a first-order inertial element. This represents the thermal stress characterization value of the inner wall of the high-pressure rotor; and These are the system matrix and the input matrix, respectively, obtained from modal analysis; The output matrix maps the reduced-order state to thermal stress.
[0021] Furthermore, in step S2, to adapt to the periodic operation characteristics of the digital control system, the reduced-order state-space model is configured according to the control cycle. Discretization yields the discrete state-space equations, which are used to recursively deduce the evolution of the reduced-order states and establish a mapping relationship with the measured rotor inner wall temperature, as shown in the following formula: In the formula, For the first Reduced-order state vectors in the reduced-order model of high-pressure rotor thermal stress for each control cycle and These are the discretized system matrix and input matrix; To regulate the steam temperature after the regulating stage; This is the measured value of the metal temperature inside the rotor wall; This is the mapping matrix between the temperature measurement points on the inner wall of the rotor and the reduced-order state vector; For measuring noise.
[0025] Furthermore, step S2 also includes: due to the reduced-order state vector Since direct measurement is not possible, a Kalman filter is constructed as a reduced-order state observer. The Kalman filter is based on the steam temperature at the previous moment. and the current metal temperature measurement value The optimal estimate of the reduced-order state vector is calculated recursively. And by the optimal estimate Calculate the current thermal stress The formula is as follows: In the formula, This is the optimal estimate of the reduced-order state vector.
[0028] Furthermore, in step S3, a state vector is constructed to comprehensively reflect the unit's operation and rotor safety status. It is composed of the following components after normalization and splicing:
[0029] Deviation between AGC instructions and actual power output , used to characterize the adjustment tracking error;
[0030] First-order difference of actual power This is used to characterize the current rate of load change;
[0031] Current regulating stage steam temperature , used to reflect thermal shock strength;
[0032] The reduced-order state estimation vector obtained in claim 4 It is used to provide complete information on the temperature field distribution;
[0033] Current thermal stress value It is used to directly indicate the stress state of the rotor;
[0034] and the AGC instruction sequence over the past m time steps This is used to provide recent trends in scheduling instructions.
[0035] Furthermore, step S4 includes the following steps: Step S41: Define the action output by the agent. The correction increment for the load setpoint of the current control cycle, and satisfying the bounded continuous interval. The optimized load command is generated by directly adding the correction increment to the actual generated power, giving the command adjustment a clear physical reference. The formula is as follows: Step S42: To ensure that the optimized command always meets the power grid's assessment requirements for AGC regulation deviation, a deviation limiting circuit is introduced. When the command differs from the original AGC command... The deviation exceeds the allowable limit When this happens, the instruction is forcibly corrected to the allowed boundary. The calculation formula for the limit operation is as follows: In the formula, The upper limit of allowable deviation is set according to the power grid AGC assessment rules; after the limit is applied. As a target load command, it is sent to the unit coordination and control system via analog output.
[0041] Furthermore, in step S5, to unify rotor life protection and AGC tracking performance into a single scalar optimization objective, the total reward function... It consists of a weighted sum of four penalties, as shown in the following formula: In the formula, To track quality penalty items; This is a penalty item for lifespan depletion; Penalty item for forward-looking thermal stress prediction; Penalty for smooth movement; Among them, tracking quality penalty items Used to suppress load deviation when the deviation is within the acceptable accuracy dead zone. The penalty is lighter when the target is within the acceptable range and heavier when it is exceeded, in order to incentivize the agent to prioritize meeting the AGC performance indicators. The formula is as follows: In the formula, For power deviation, , To adjust the quality weighting coefficient; Lifetime depletion penalty It directly reflects the fatigue damage caused to the rotor by each control action, follows Miner's linear cumulative damage law, and converts the thermal stress amplitude into the equivalent life loss rate. Furthermore, a penalty is imposed on the equivalent life loss rate, thereby explicitly incorporating equipment life costs into the control. The calculation formula is as follows: In the formula, This is the current thermal stress value. This represents the fatigue limit stress amplitude of the rotor material. and These are the material fatigue characteristic constants. This is the lifespan loss weighting coefficient.
[0051] Furthermore, the reward function in step S5 also includes: Forward thermal stress prediction penalty Using the reduced-order model, assuming the current action Based on the trend of steam temperature change, recursively extrapolate backwards. Step 1: Predict the maximum potential thermal stress in the future and detect when it exceeds the alarm limit. The penalty is applied to certain parts of the system to drive the agent to avoid dangerous stress conditions in advance. The calculation formula is as follows: In the formula, For thermal stress alarm limits, To predict penalty weights; Smooth motion penalty This method is used to suppress high-frequency and drastic fluctuations in load commands, at the cost of the absolute difference between the increments of two adjacent actions, to ensure stable unit operation. , For smoothing weights.
[0057] Furthermore, step S6 includes the following steps: Step S61: Construct the Actor Network Used according to state Generate Actions Constructing a Critic network Used to assess the state Take action below Expected cumulative reward; Step S62: To address the problem that the general Critic network struggles to perceive the consequences of long-term thermodynamic damage, physical model knowledge is injected into the input of the Critic network. The specific steps are as follows: At each decision, the current reduced-order state is used to estimate... Combined with actions The resulting change in equivalent steam temperature The thermal stress value at the next moment can be predicted in one step using the discrete order reduction model. The prediction formula is as follows: The predicted thermal stress Compared with the current state ,action Concatenate into augmented vectors As input to the Critic network, the value function... It can directly sense the instantaneous changes in thermal stress caused by the action; Step S63: To further enforce the avoidance of high-stress actions from the mathematical objective of value iteration, during the training of the Critic network, the target value in the loss function is adjusted. After modifying the physical constraints, the target value is calculated using the following formula: In the formula, Discount factor; This is an additional thermal stress limit penalty factor; and These are the target Critic network and the target Actor network, respectively. The thermal stress at the next moment is obtained from the prediction formula; This refers to the thermal stress alarm limit. The above modifications force the expected reward of any action that causes the predicted thermal stress value to exceed the limit in the next moment to be forcibly reduced, thereby guiding the Actor network to actively generate load corrections that keep the rotor stress safe during policy gradient updates.
[0068] Furthermore, step S7 includes the following steps:
[0069] Step S71: Collect historical data of the unit's long-term operation under deep peak shaving conditions, including AGC command curves, power response, steam temperature and metal temperature records. Using the reduced-order model as the core, combined with a simplified thermodynamic model that reflects the change of steam temperature with load, construct a reinforcement learning environment simulator. The reinforcement learning environment simulator can provide simulation feedback on the power, steam temperature and thermal stress state at the next moment based on the current action.
[0070] Step S72: In the simulation environment, the Actor-Critic network structure and loss function defined in step S6 are used for offline pre-training until the cumulative reward converges, so that the agent learns a shaping strategy that takes into account both AGC assessment and rotor life under different initial thermal stress and command modes.
[0071] Step S73: Solidify the trained Actor network parameters into an inference model, and deploy it in the upper-level optimization station of the unit's distributed control system via the OPC communication interface. During online operation, the state vector defined in step S3 is obtained in each control cycle, normalized, and then fed into the Actor network to output the action. In step S4, generate optimized load instructions and send them to the coordinated control system;
[0072] Step S74: Simultaneously set up a safety monitoring logic independent of the intelligent agent. If the thermal stress value calculated in step S2 is detected in multiple consecutive control cycles... All exceeded the alarm limit. If the intelligent agent outputs abnormally, it will automatically cut off the optimized instruction path and restore the original AGC scheduling instruction direct path to ensure the safety of the unit hardware.
[0073] The present invention has the following beneficial effects:
[0074] 1. This invention embeds a deep reinforcement learning agent that integrates a thermal stress reduction model between the scheduling command and the coordination control system. This agent explicitly quantifies rotor fatigue life loss as a penalty term in the reward function, ensuring that the load command generation process simultaneously considers both minimizing tracking deviation and minimizing life loss. During training, the agent autonomously learns to proactively avoid load change patterns that may cause severe thermal stress while meeting the grid AGC assessment requirements, thereby effectively extending rotor service life without sacrificing regulation quality.
[0075] 2. This invention directly injects the thermal stress characteristics predicted in one step by the reduced-order model into the input of the Critic network, and adds a thermal stress exceeding-limit penalty term to the value update objective, enabling the value function to have the ability to perceive thermodynamic consequences in real time. This structure, which combines a physical model with data-driven approaches, ensures that the learned control strategy inherently conforms to rotor thermal stress constraints, rather than relying entirely on trial and error. At the same time, the physical laws embodied in the reduced-order model enhance the strategy's generalization ability to different operating conditions, reducing the risk of a pure black-box strategy generating dangerous commands under unseen conditions.
[0076] 3. Only the inference model needs to be deployed in the existing distributed control system of the unit's supervisory optimization station. Existing operating data such as temperature and power are obtained through the OPC communication interface. The Kalman filter observer is used to estimate the internal thermal stress state of the rotor in real time. The load correction is calculated by the Actor network and then sent to the coordinated control system. The whole process does not require the addition of new sensors or modification of actuators, and does not change the underlying control loop structure. The engineering implementation difficulty and modification cost are both low.
[0077] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description
[0078] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0079] Figure 1 This is a flowchart illustrating an AGC instruction optimization method for thermal power units based on deep reinforcement learning, according to the present invention. Detailed Implementation
[0080] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0081] Please see Figure 1 As shown, this invention is a method for optimizing AGC instructions for thermal power units based on deep reinforcement learning, comprising the following steps:
[0082] Step S1: Construct a reduced-order thermal stress model of the high-pressure rotor of the unit. Using the rotor material and geometric parameters, through simulation and intrinsic orthogonal decomposition, establish a reduced-order state space model with steam temperature excitation as input and thermal stress as output, which is used to replace the finite element model to realize real-time thermal stress calculation.
[0083] Step S2: Real-time acquisition of AGC commands, actual unit power generation, steam temperature after regulating stage and rotor metal temperature measurements. A reduced-order state observer is constructed using a discretized reduced-order model. By fusing the reduced-order model and metal temperature measurements, the reduced-order state vector inside the rotor is estimated in real time and the current thermal stress value is calculated.
[0084] Step S3: Construct the state space of the deep reinforcement learning agent, and combine the power deviation, load change rate, steam temperature, reduced state vector, current thermal stress value and historical AGC instruction sequence into a normalized state vector so that the agent can perceive the load adjustment requirements and rotor thermodynamic state.
[0085] Step S4: Define the action space of the intelligent agent as the continuous correction increment of the load setpoint, generate an optimized load command based on the action output and the current actual power generation, and after processing the deviation limit link from the original AGC command, send it to the unit coordination control system.
[0086] Step S5: Establish a reward function that integrates lifetime loss perception. The reward function consists of four parts: a tracking quality penalty term, an equivalent lifetime loss rate penalty term based on the fatigue characteristics of rotor materials, a thermal stress penalty term using a reduced-order model for forward prediction, and a motion smoothing penalty term.
[0087] Step S6: Construct the Actor network and Critic network using the deep deterministic strategy gradient algorithm. Embed the thermal stress features predicted one step by the reduced-order model for the current action at the input of the Critic network, and introduce a thermal stress limit violation penalty term into the value update objective of the Critic network. Through the dual embedding of the physical model and reinforcement learning, the learned load adjustment strategy is forced to meet the rotor thermal stress life constraint.
[0088] Step S7: Build a simulation environment using historical unit operating data and a reduced-order model, perform offline pre-training of the Actor network and Critic network, and deploy the Actor network online after training. Generate optimized instructions in each control cycle and configure over-limit safety monitoring logic. When thermal stress is detected to be continuously exceeding the limit, automatically restore the original AGC instruction path.
[0089] Step S1 includes the following steps: Step S11: Obtain the geometric dimensions and material properties of the high-pressure rotor, such as density, specific heat capacity, thermal conductivity, elastic modulus, and coefficient of thermal expansion. Establish an axisymmetric finite element model and use the step and ramp excitation of steam temperature as boundary conditions to perform simulation and obtain the time series data of transient temperature field and thermal stress field. Step S12: To transform the computationally complex finite element model into a simplified model that can run online in real time, the spatial modes and time coefficients with the highest energy proportions in the first r orders are extracted from the simulation data using intrinsic orthogonal decomposition. A reduced-order state-space model of thermal stress in the high-pressure rotor is established. The reduced-order state-space model directly calculates the characteristic value of thermal stress on the inner wall of the rotor through state equations and output equations, as shown in the following formula: In the formula, This is an r-dimensional reduced-order state vector, which physically corresponds to the main modal coordinates of the temperature field; The equivalent steam temperature excitation acting on the rotor surface is characterized by the steam temperature after the regulating stage through a first-order inertial element. This represents the thermal stress characterization value of the inner wall of the high-pressure rotor; and These are the system matrix and the input matrix, respectively, obtained from modal analysis; The output matrix maps the reduced-order state to thermal stress.
[0095] In step S2, to adapt to the periodic operation characteristics of the digital control system, the reduced-order state-space model is modified according to the control cycle. Discretization yields the discrete state-space equations, which are used to recursively deduce the evolution of the reduced-order states and establish a mapping relationship with the measured rotor inner wall temperature, as shown in the following formula: In the formula, For the first Reduced-order state vectors in the reduced-order model of high-pressure rotor thermal stress for each control cycle and These are the discretized system matrix and input matrix; To regulate the steam temperature after the regulating stage; This is the measured value of the metal temperature inside the rotor wall; This is the mapping matrix between the temperature measurement points on the inner wall of the rotor and the reduced-order state vector; For measuring noise.
[0099] Step S2 also includes: due to the reduced-order state vector Since direct measurement is not possible, a Kalman filter is constructed as a reduced-order state observer. The Kalman filter is based on the steam temperature at the previous moment. and the current metal temperature measurement value The optimal estimate of the reduced-order state vector is calculated recursively. And by the optimal estimate Calculate the current thermal stress The formula is as follows: In the formula, This is the optimal estimate of the reduced-order state vector.
[0102] In step S3, a state vector is constructed to comprehensively reflect the unit's operation and rotor safety status. It is composed of the following components after normalization and splicing:
[0103] Deviation between AGC instructions and actual power output , used to characterize the adjustment tracking error;
[0104] First-order difference of actual power This is used to characterize the current rate of load change;
[0105] Current regulating stage steam temperature , used to reflect thermal shock strength;
[0106] The reduced-order state estimation vector obtained in claim 4 It is used to provide complete information on the temperature field distribution;
[0107] Current thermal stress value It is used to directly indicate the stress state of the rotor;
[0108] and the AGC instruction sequence over the past m time steps This is used to provide recent trends in scheduling instructions.
[0109] Step S4 includes the following steps: Step S41: Define the action output by the agent. The correction increment for the load setpoint of the current control cycle, and satisfying the bounded continuous interval. The optimized load command is generated by directly adding the correction increment to the actual generated power, giving the command adjustment a clear physical reference. The formula is as follows: Step S42: To ensure that the optimized command always meets the power grid's assessment requirements for AGC regulation deviation, a deviation limiting circuit is introduced. When the command differs from the original AGC command... The deviation exceeds the allowable limit When this happens, the instruction is forcibly corrected to the allowed boundary. The calculation formula for the limit operation is as follows: In the formula, The upper limit of allowable deviation is set according to the power grid AGC assessment rules; after the limit is applied. As a target load command, it is sent to the unit coordination and control system via analog output.
[0115] In step S5, to unify rotor life protection and AGC tracking performance into a single scalar optimization objective, the total reward function is... It consists of a weighted sum of four penalties, as shown in the following formula: In the formula, To track quality penalty items; This is a penalty item for lifespan depletion; Penalty item for forward-looking thermal stress prediction; Penalty for smooth movement; Among them, tracking quality penalty items Used to suppress load deviation when the deviation is within the acceptable accuracy dead zone. The penalty is lighter when the target is within the acceptable range and heavier when it is exceeded, in order to incentivize the agent to prioritize meeting the AGC performance indicators. The formula is as follows: In the formula, For power deviation, , To adjust the quality weighting coefficient; Lifetime depletion penalty It directly reflects the fatigue damage caused to the rotor by each control action, follows Miner's linear cumulative damage law, and converts the thermal stress amplitude into the equivalent life loss rate. Furthermore, a penalty is imposed on the equivalent life loss rate, thereby explicitly incorporating equipment life costs into the control. The calculation formula is as follows: In the formula, This is the current thermal stress value. This represents the fatigue limit stress amplitude of the rotor material. and These are the material fatigue characteristic constants. This is the lifespan loss weighting coefficient.
[0125] The reward function in step S5 also includes: Forward thermal stress prediction penalty Using a reduced-order model, assuming the current action Based on the trend of steam temperature change, recursively extrapolate backwards. Step 1: Predict the maximum potential thermal stress in the future and detect when it exceeds the alarm limit. The penalty is applied to certain parts of the system to drive the agent to avoid dangerous stress conditions in advance. The calculation formula is as follows: In the formula, For thermal stress alarm limits, To predict penalty weights; Smooth motion penalty This method is used to suppress high-frequency and drastic fluctuations in load commands, at the cost of the absolute difference between the increments of two adjacent actions, to ensure stable unit operation. , For smoothing weights.
[0131] Step S6 includes the following steps: Step S61: Construct the Actor Network Used according to state Generate Actions Constructing a Critic network Used to assess the state Take action below Expected cumulative reward; Step S62: To address the problem that the general Critic network struggles to perceive the consequences of long-term thermodynamic damage, physical model knowledge is injected into the input of the Critic network. The specific steps are as follows: At each decision, the current reduced-order state is used to estimate... Combined with actions The resulting change in equivalent steam temperature The thermal stress value at the next moment can be predicted in one step using a discrete order reduction model. The prediction formula is as follows: The predicted thermal stress Compared with the current state ,action Concatenate into augmented vectors As input to the Critic network, the value function... It can directly sense the instantaneous changes in thermal stress caused by the action; Step S63: To further enforce the avoidance of high-stress actions from the mathematical objective of value iteration, during the training of the Critic network, the target value in the loss function is adjusted. After modifying the physical constraints, the target value is calculated using the following formula: In the formula, Discount factor; This is an additional thermal stress limit penalty factor; and These are the target Critic network and the target Actor network, respectively. The thermal stress at the next moment is obtained from the prediction formula; This refers to the thermal stress alarm limit. The above modifications force the expected reward of any action that causes the predicted thermal stress value to exceed the limit in the next moment to be forcibly reduced, thereby guiding the Actor network to actively generate load corrections that keep the rotor stress safe during policy gradient updates.
[0142] Step S7 includes the following steps:
[0143] Step S71: Collect historical data of the unit's long-term operation under deep peak shaving conditions, including AGC command curves, power response, steam temperature and metal temperature records. Using the reduced-order model as the core, combined with a simplified thermodynamic model that reflects the change of steam temperature with load, construct a reinforcement learning environment simulator. The reinforcement learning environment simulator can provide simulation feedback on the power, steam temperature and thermal stress state at the next moment based on the current action.
[0144] Step S72: In the simulation environment, the Actor-Critic network structure and loss function defined in step S6 are used for offline pre-training until the cumulative reward converges, so that the agent learns a shaping strategy that takes into account both AGC assessment and rotor life under different initial thermal stress and command modes.
[0145] Step S73: Solidify the trained Actor network parameters into an inference model, and deploy it in the upper-level optimization station of the unit's distributed control system via the OPC communication interface. During online operation, the state vector defined in step S3 is obtained in each control cycle, normalized, and then fed into the Actor network to output the action. In step S4, generate optimized load instructions and send them to the coordinated control system;
[0146] Step S74: Simultaneously set up a safety monitoring logic independent of the intelligent agent. If the thermal stress value calculated in step S2 is detected in multiple consecutive control cycles... All exceeded the alarm limit. If the intelligent agent outputs abnormally, it will automatically cut off the optimized instruction path and restore the original AGC scheduling instruction direct path to ensure the safety of the unit hardware.
[0147] One specific application of this embodiment is:
[0148] This embodiment uses a 600MW supercritical thermal power unit as the implementation object. The high-pressure rotor material of this unit is 12Cr1MoV alloy steel, the rated main steam pressure is 24.2MPa, the rated main steam temperature is 566℃, the grid AGC dead zone is 6MW, and the allowable range of load change rate is 0~12MW / min. The specific implementation steps are as follows:
[0149] Step S1: Construct a reduced-order model of thermal stress on the high-pressure rotor of the unit.
[0150] This step establishes a reduced-order thermal stress model that can be run online in real time through simulation and intrinsic orthogonal decomposition, replacing the highly complex finite element model to achieve rapid calculation of thermal stress. The specific steps are as follows:
[0151] Step S11: Rotor Finite Element Modeling and Simulation
[0152] Obtain the geometric dimensions and material properties of the high-pressure rotor: Geometric dimensions include key structural parameters such as rotor shaft diameter, shaft shoulder fillet radius, and impeller hub thickness; material properties include material density. Specific heat capacity Thermal conductivity Elastic modulus Coefficient of thermal expansion Poisson's ratio .
[0153] Based on the above parameters, an axisymmetric finite element model was established. The step excitation (amplitude 20~100℃) and ramp excitation (change rate 1~5℃ / min) of steam temperature were used as boundary conditions to carry out transient thermo-mechanical coupling simulation. The transient temperature field and thermal stress field time series data with a duration of 120h were obtained, with a sampling interval of 1s.
[0154] Step S12: Construction of a reduced-order model based on eigenorthogonal decomposition
[0155] Extracting the frontier from simulation data using intrinsic orthogonal decomposition The spatial mode and time coefficient with the highest energy percentage are selected in this embodiment. The top 6 modes with a cumulative energy percentage of 99.9% are chosen. A reduced-order state-space model of thermal stress in a high-pressure rotor is established. The state equations and output equations of the model are as follows:
[0156] ;
[0157] ;
[0158] In the formula, for A reduced-order state vector, which physically corresponds to the main modal coordinates of the temperature field; The equivalent steam temperature excitation acting on the rotor surface is characterized by the steam temperature after the regulating stage through a first-order inertial element, with an inertial time constant of 30s. This represents the thermal stress characterization value of the inner wall of the high-pressure rotor; and These are the system matrix and the input matrix, respectively, obtained from POD modal analysis; This is the output matrix, used to map the reduced-order state to thermal stress.
[0159] Step S2: Construct a reduced-order state observer to calculate rotor thermal stress in real time.
[0160] Step S21: Discretization of the reduced-order model
[0161] Control cycle adapted to the unit's distributed control system (DCS) The continuous reduced-order state-space model is discretized according to the control cycle to obtain the discrete state-space equation, which is used to deduce the evolution of the reduced-order state. At the same time, a mapping relationship with the rotor inner wall temperature measurement value is established, as shown in the following formula:
[0162] ;
[0163] ;
[0164] In the formula, For the first Reduced-order state vector in the reduced-order model of high-pressure rotor thermal stress for each control cycle; and The discretized system matrix and input matrix are obtained by discretizing using the zero-order preservation method. For the first The steam temperature after the regulating stage is collected in each control cycle; For the first The measured value of the rotor inner wall metal temperature collected in each control cycle; The mapping matrix between the temperature measurement points on the inner wall of the rotor and the reduced-order state vector is determined by the spatial distribution characteristics of the POD mode. The noise is measured and follows a Gaussian distribution with a mean of 0 and a variance of 0.5.
[0165] Step S22: Design of Reduced-Order State Observer and Real-Time Calculation of Thermal Stress
[0166] Due to the reduced-order state vector Since direct measurement is not possible, a Kalman filter is constructed as a reduced-order state observer. The Kalman filter is based on the steam temperature at the previous time step. and the current metal temperature measurement value The optimal estimate of the reduced-order state vector is calculated recursively. And calculate the current thermal stress using the optimal estimate. The calculation formula is as follows:
[0167] ;
[0168] In the formula, For the first The optimal estimate of the reduced-order state vector for each control cycle.
[0169] In this embodiment, the Kalman filter is recursively executed according to the standard linear Kalman filter algorithm, and the process noise covariance matrix is taken as... Measure the noise covariance The initial state covariance matrix is taken as , for An identity matrix of order 1.
[0170] Step S3: Construct the state space of the deep reinforcement learning agent.
[0171] To comprehensively reflect the unit's operation and rotor safety status, a state vector is constructed. The following components are normalized using min-max normalization (normalization interval). It was pieced together later:
[0172] Deviation between AGC instructions and actual power output , used to characterize the adjustment tracking error, where For the first Power output of grid AGC commands per control cycle For the first The actual power output of the unit in each control cycle;
[0173] First-order difference of actual power This is used to characterize the current rate of load change;
[0174] Current regulating stage steam temperature , used to reflect thermal shock strength;
[0175] The reduced-order state estimation vector obtained in step S2 It is used to provide complete information on the temperature field distribution;
[0176] Current thermal stress value It is used to directly indicate the stress state of the rotor;
[0177] past AGC instruction sequence at each moment This is used to provide recent trends in scheduling instructions; in this embodiment, it is taken as... .
[0178] The final state vector Dimensions Dimensions serve as inputs to deep reinforcement learning agents.
[0179] Step S4: Define the agent's action space and generate optimized load instructions.
[0180] Step S41, Defining the Action Space
[0181] Define the action output by the intelligent agent. The correction increment for the load setpoint of the current control cycle, and satisfying the bounded continuous interval. In this embodiment, we take That is, the maximum correction range per cycle is 2MW; the optimized load command is generated by directly adding the correction increment to the actual generated power, so that the command adjustment has a clear physical reference, as shown in the following formula:
[0182] ;
[0183] Step S42, Deviation Limiting Design
[0184] To ensure that the optimized commands consistently meet the power grid's performance requirements for AGC (Automatic Gauge Control) deviation, a deviation limiting mechanism is introduced. This mechanism limits the deviation when the command differs from the original AGC command. The deviation exceeds the allowable limit When necessary, the instruction is forcibly corrected to the allowable boundary. In this embodiment, according to the power grid AGC assessment rules, the following is taken: The formula for calculating the limit operation is as follows:
[0185] ;
[0186] In the formula, For sign function; after limiting As a target load command, it is sent to the unit coordination and control system via a 4-20mA analog output.
[0187] Step S5: Establish a reward function that integrates lifetime loss perception.
[0188] To unify rotor life protection and AGC tracking performance into a scalar optimization objective, the total reward function... It consists of a weighted sum of four penalties, as shown in the following formula:
[0189] ;
[0190] In the formula, To track quality penalty items; This is a penalty item for lifespan depletion; Penalty item for forward-looking thermal stress prediction; This is a penalty for smoothing out actions.
[0191] The specific calculation methods for each item are as follows:
[0192] Track quality penalty items Used to suppress load deviation when the deviation is within the acceptable accuracy dead zone. The penalty is lighter when the target is within the acceptable range and heavier when it is exceeded, in order to incentivize the agent to prioritize meeting the AGC performance indicators; in this embodiment, we take... (Power grid AGC regulation dead zone), weighting coefficient , The formula is as follows:
[0193] ;
[0194] In the formula, The power deviation is defined in step S3.
[0195] Lifetime depletion penalty It directly reflects the fatigue damage caused to the rotor by each control action, follows Miner's linear cumulative damage law, and converts the thermal stress amplitude into the equivalent life loss rate. Furthermore, a penalty is imposed on the equivalent life loss rate, explicitly incorporating the equipment life cost into the control. In this embodiment, the fatigue limit stress amplitude of the rotor material 12Cr1MoV is... Fatigue characteristic constants , Weighting coefficient The calculation formula is as follows:
[0196] ;
[0197] ;
[0198] In the formula, This is the current thermal stress value calculated in step S2.
[0199] Forward thermal stress prediction penalty Using the reduced-order model established in step S1, assuming the current action Based on the trend of steam temperature change, recursively extrapolate backwards. Step 1: Predict the maximum potential thermal stress in the future and detect when it exceeds the alarm limit. The portion of the penalty is used to drive the agent to avoid dangerous stress conditions in advance; in this embodiment, the following is taken: (60s forward prediction) Thermal stress alarm limit Weighting coefficient The calculation formula is as follows:
[0200] ;
[0201] ;
[0202] Smooth motion penalty To suppress high-frequency and drastic fluctuations in load commands, the absolute difference between two adjacent action increments is used as a trade-off to ensure stable unit operation; in this embodiment, the weighting coefficient... The formula is as follows:
[0203] ;
[0204] Step S6: Construct a DDPG algorithm network architecture with embedded physical constraints.
[0205] A deep deterministic policy gradient algorithm is used to construct the Actor network and the Critic network. Through the dual embedding of the physical model and reinforcement learning, the learned load regulation strategy is forced to meet the rotor thermal stress life constraint.
[0206] Step S61: Building the basic network architecture
[0207] Constructing an Actor Network Used to determine the state Generate Actions ; Constructing a Critic network Used to evaluate in state Take action below The expected cumulative reward. Simultaneously, construct the corresponding target Actor network. and target Critic network The network structure is consistent with the main network and is used to stabilize the training process.
[0208] In this embodiment, the Actor network adopts a 3-layer fully connected neural network structure: the input layer has a dimension of 30 (consistent with the dimension of the state vector), both hidden layer 1 and hidden layer 2 have 128 neurons, the activation function is ReLU, the output layer has a dimension of 1, the activation function is Tanh, and the output is mapped to... Interval.
[0209] The Critic network employs a 3-layer fully connected neural network structure: the input layer is an augmented vector, both hidden layers 1 and 2 have 128 neurons and ReLU is used as the activation function, and the output layer has a dimension of 1, no activation function, and outputs the value of the state-action pair.
[0210] Step S62: Embedding thermal stress features at the input end of the Critic network
[0211] To address the issue that general-purpose Critic networks struggle to perceive the consequences of long-term thermodynamic damage, physical model knowledge is injected into the input of the Critic network. During each decision-making process, the current reduced-order state is used to estimate... Combined with actions The resulting change in equivalent steam temperature The thermal stress value at the next moment is predicted in one step using a discrete reduced-order model; in this embodiment, the mapping relationship between the steam temperature change and the load correction increment is taken as follows: (℃ / MW), the prediction formula is as follows:
[0212] ;
[0213] ;
[0214] The predicted thermal stress Compared with the current state ,action Concatenate into augmented vectors As input to the Critic network, the value function can directly sense the instantaneous changes in thermal stress caused by the action. The value function expression is as follows:
[0215] ;
[0216] Step S63: Critic network value update with thermal stress penalty
[0217] To further enforce the avoidance of high-stress actions from the mathematical objective of value iteration, during the training of the Critic network, the target value in the loss function is adjusted. After modifying the physical constraints, the target value is calculated using the following formula:
[0218] ;
[0219] In the formula, As the discount factor, in this embodiment, we take... ; As an additional thermal stress limit penalty coefficient, in this embodiment, we take... ; The thermal stress at the next moment is obtained from the prediction formula; The value is the thermal stress alarm limit, and it is consistent with the value in step S5.
[0220] Based on the above target value, the loss function of the Critic network adopts the mean squared error loss, as shown in the following formula:
[0221] ;
[0222] In the formula, For the training batch size, in this embodiment, we take... .
[0223] The policy gradient update formula for the Actor network is:
[0224] ;
[0225] During training, the target network updates its parameters using a soft update method, and the soft update coefficients... The updated formula is as follows:
[0226] ;
[0227] ;
[0228] Step S7: Offline pre-training and online deployment of the model
[0229] Step S71: Setting up a reinforcement learning environment simulator
[0230] Historical operating data of the unit under deep peak shaving conditions for 6 consecutive months were collected, with a sampling interval of 1 second. The data included AGC command curves, power response, steam temperature after the regulating stage, and rotor metal temperature records. Based on the reduced-order model established in step S1, and combined with a simplified thermodynamic model reflecting the change of steam temperature with load, a reinforcement learning environment simulator was constructed. This simulator can provide simulation feedback on the power, steam temperature, and thermal stress state at the next moment based on the current action. The simulation step size is consistent with the control cycle, which is 1 second.
[0231] Step S72: Offline pre-training of the model
[0232] In the simulation environment described above, the Actor-Critic network structure and loss function defined in step S6 are used for offline pre-training. The optimizer employs the Adam algorithm, and the learning rate is set to: Actor network Critic Network The experience replay pool capacity is set to... The total number of training steps is The process continues until the cumulative reward converges, enabling the agent to learn the optimal strategy that balances AGC assessment and rotor life under different initial thermal stress and instruction modes.
[0233] Step S73: Online Deployment of the Model
[0234] The trained Actor network parameters are solidified into an ONNX format inference model and deployed in the host computer's DCS upper-level optimization station via the OPCUA communication interface. The communication cycle between the upper-level optimization station and the DCS system is 1 second. During online operation, each control cycle acquires the components of the state vector defined in step S3, processes them using the min-max normalization parameters determined during offline training, and then feeds them into the Actor network to infer the output action. Step S4 generates an optimized load command and sends it to the unit coordination and control system.
[0235] Step S74, Security Monitoring Logic Configuration
[0236] Set up a security monitoring logic independent of the intelligent agent. If the thermal stress value calculated in step S2 is detected for 5 consecutive control cycles... All exceeded the alarm limit. If the intelligent agent outputs abnormally, it will immediately cut off the optimization command path through the relay output contact and restore the original AGC scheduling command direct path to ensure the safety of the unit hardware.
[0237] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0238] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A method for optimizing AGC instructions for thermal power units based on deep reinforcement learning, characterized in that, Includes the following steps: Step S1: Construct a reduced-order thermal stress model of the high-pressure rotor of the unit. Using the rotor material and geometric parameters, through simulation and intrinsic orthogonal decomposition, establish a reduced-order state space model with steam temperature excitation as input and thermal stress as output, which is used to replace the finite element model to realize real-time thermal stress calculation. Step S2: Real-time acquisition of AGC commands, actual unit power generation, steam temperature after regulating stage and rotor metal temperature measurements; construction of a reduced-order state observer using a discretized reduced-order model; estimation of the reduced-order state vector inside the rotor and calculation of the current thermal stress value by fusing the reduced-order model with the metal temperature measurements. Step S3: Construct the state space of the deep reinforcement learning agent, and combine the power deviation, load change rate, steam temperature, the reduced state vector, the current thermal stress value and the historical AGC instruction sequence into a normalized state vector so that the agent can perceive the load adjustment requirements and the rotor thermodynamic state. Step S4: Define the action space of the intelligent agent as the continuous correction increment of the load setpoint, generate an optimized load command based on the action output and the current actual power generation, and after processing the deviation limit link from the original AGC command, send it to the unit coordination control system. Step S5: Establish a reward function that integrates lifetime loss perception. The reward function consists of four parts: a tracking quality penalty term, an equivalent lifetime loss rate penalty term based on the fatigue characteristics of rotor materials, a thermal stress penalty term for forward prediction using the reduced-order model, and an action smoothing penalty term. Step S6: Construct the Actor network and Critic network using the deep deterministic strategy gradient algorithm. Embed the thermal stress features predicted by the reduced-order model for the current action into the input of the Critic network. Introduce a thermal stress limit violation penalty term into the value update objective of the Critic network. Through the dual embedding of the physical model and reinforcement learning, force the learned load adjustment strategy to meet the rotor thermal stress life constraint. Step S7: Build a simulation environment using the unit's historical operating data and the reduced-order model, perform offline pre-training on the Actor network and Critic network, and deploy the Actor network online after training. Generate optimized instructions in each control cycle and configure over-limit safety monitoring logic. When thermal stress is detected to be continuously exceeding the limit, automatically restore the original AGC instruction path.
2. The method for optimizing AGC instructions for thermal power units based on deep reinforcement learning according to claim 1, characterized in that, Step S1 includes the following steps: Step S11: Obtain the geometric dimensions and material properties of the high-pressure rotor, such as density, specific heat capacity, thermal conductivity, elastic modulus, and coefficient of thermal expansion. Establish an axisymmetric finite element model and use the step and ramp excitation of steam temperature as boundary conditions to perform simulation and obtain the time series data of transient temperature field and thermal stress field. Step S12: To transform the computationally complex finite element model into a simplified model that can run online in real time, the spatial modes and time coefficients with the highest energy proportions in the first r orders are extracted from the simulation data using intrinsic orthogonal decomposition. A reduced-order state-space model of the high-pressure rotor thermal stress is established. The reduced-order state-space model directly calculates the characterization value of the rotor inner wall thermal stress through state equations and output equations, as shown in the following formula: ; ; In the formula, This is an r-dimensional reduced-order state vector, which physically corresponds to the main modal coordinates of the temperature field; The equivalent steam temperature excitation acting on the rotor surface is characterized by the steam temperature after the regulating stage through a first-order inertial element. This represents the thermal stress characterization value of the inner wall of the high-pressure rotor; and These are the system matrix and the input matrix, respectively, obtained from modal analysis; The output matrix maps the reduced-order state to thermal stress.
3. The method for optimizing AGC instructions for thermal power units based on deep reinforcement learning according to claim 1, characterized in that, In step S2, to adapt to the periodic operation characteristics of the digital control system, the reduced-order state-space model is modified according to the control cycle. Discretization yields the discrete state-space equations, which are used to recursively deduce the evolution of the reduced-order states and establish a mapping relationship with the measured rotor inner wall temperature, as shown in the following formula: ; ; In the formula, For the first Reduced-order state vectors in the reduced-order model of high-pressure rotor thermal stress for each control cycle and These are the discretized system matrix and input matrix; To regulate the steam temperature after the regulating stage; This is the measured value of the metal temperature inside the rotor wall; This is the mapping matrix between the temperature measurement points on the inner wall of the rotor and the reduced-order state vector; For measuring noise.
4. The method for optimizing AGC instructions for thermal power units based on deep reinforcement learning according to claim 1, characterized in that, Step S2 further includes: due to the reduced state vector Since direct measurement is not possible, a Kalman filter is constructed as a reduced-order state observer. The Kalman filter is based on the steam temperature at the previous moment. and the current metal temperature measurement value The optimal estimate of the reduced-order state vector is calculated recursively. And by the optimal estimate Calculate the current thermal stress The formula is as follows: ; In the formula, This is the optimal estimate of the reduced-order state vector.
5. The method for optimizing AGC instructions for thermal power units based on deep reinforcement learning according to claim 1, characterized in that, In step S3, a state vector is constructed to comprehensively reflect the unit's operation and rotor safety status. It is composed of the following components after normalization and splicing: Deviation between AGC instructions and actual power output , used to characterize the adjustment tracking error; First-order difference of actual power This is used to characterize the current rate of load variation; Current regulating stage steam temperature , used to reflect thermal shock strength; The reduced-order state estimation vector obtained in claim 4 It is used to provide complete information on the temperature field distribution; Current thermal stress value It is used to directly indicate the stress state of the rotor; and the AGC instruction sequence over the past m time steps This is used to provide recent trends in scheduling instructions.
6. The method for optimizing AGC instructions for thermal power units based on deep reinforcement learning according to claim 1, characterized in that, Step S4 includes the following steps: Step S41: Define the action output by the agent. The correction increment for the load setpoint of the current control cycle, and satisfying the bounded continuous interval. The optimized load command is generated by directly adding the correction increment to the actual generated power, giving the command adjustment a clear physical reference. The formula is as follows: ; Step S42: To ensure that the optimized command always meets the power grid's assessment requirements for AGC regulation deviation, a deviation limiting circuit is introduced. When the command differs from the original AGC command... The deviation exceeds the allowable limit When this happens, the instruction is forcibly corrected to the allowed boundary. The calculation formula for the limit operation is as follows: ; In the formula, The upper limit of allowable deviation is set according to the power grid AGC assessment rules; after the limit is applied. As a target load command, it is sent to the unit coordination and control system via analog output.
7. The method for optimizing AGC instructions for thermal power units based on deep reinforcement learning according to claim 1, characterized in that, In step S5, to unify rotor life protection and AGC tracking performance into a single scalar optimization objective, the total reward function... It consists of a weighted sum of four penalties, as shown in the following formula: ; In the formula, To track quality penalty items; This is a penalty item for lifespan depletion; Penalty item for forward-looking thermal stress prediction; Penalty for smooth movement; Among them, tracking quality penalty items Used to suppress load deviation when the deviation is within the acceptable accuracy dead zone. The penalty is lighter when the target is within the acceptable range and heavier when it is exceeded, in order to incentivize the agent to prioritize meeting the AGC performance indicators. The formula is as follows: ; In the formula, For power deviation, , To adjust the quality weighting coefficient; Lifetime depletion penalty It directly reflects the fatigue damage caused to the rotor by each control action, follows Miner's linear cumulative damage law, and converts the thermal stress amplitude into the equivalent life loss rate. Furthermore, a penalty is imposed on the equivalent life loss rate, thereby explicitly incorporating equipment life costs into the control. The calculation formula is as follows: ; ; In the formula, This is the current thermal stress value. This represents the fatigue limit stress amplitude of the rotor material. and These are the material fatigue characteristic constants. This is the lifespan loss weighting coefficient.
8. The method for optimizing AGC instructions for thermal power units based on deep reinforcement learning according to claim 1, characterized in that, The reward function in step S5 further includes: Forward thermal stress prediction penalty Using the reduced-order model, assuming the current action Based on the trend of steam temperature change, recursively extrapolate backwards. Step 1: Predict the maximum potential thermal stress in the future and detect when it exceeds the alarm limit. The penalty is applied to certain parts of the system to drive the agent to avoid dangerous stress conditions in advance. The calculation formula is as follows: ; ; In the formula, For thermal stress alarm limits, To predict penalty weights; Smooth motion penalty This method is used to suppress high-frequency and drastic fluctuations in load commands, at the cost of the absolute difference between the increments of two adjacent actions, to ensure smooth unit operation. , For smoothing weights.
9. The method for optimizing AGC instructions for thermal power units based on deep reinforcement learning according to claim 1, characterized in that, Step S6 includes the following steps: Step S61: Construct the Actor Network Used according to state Generate Actions Constructing a Critic network Used to assess the state Take action below Expected cumulative reward; Step S62: To address the problem that the general Critic network struggles to perceive the consequences of long-term thermodynamic damage, physical model knowledge is injected into the input of the Critic network. The specific steps are as follows: At each decision, the current reduced-order state is used to estimate... Combined with actions The resulting equivalent steam temperature change The thermal stress value at the next moment can be predicted in one step using the discrete order reduction model. The prediction formula is as follows: ; ; The predicted thermal stress Compared with the current state ,action Concatenate into augmented vectors As input to the Critic network, the value function... It can directly sense the instantaneous changes in thermal stress caused by the action; Step S63: To further enforce the avoidance of high-stress actions from the mathematical objective of value iteration, during the training of the Critic network, the target value in the loss function is adjusted. After modifying the physical constraints, the target value is calculated using the following formula: ; In the formula, Discount factor; This is an additional thermal stress limit penalty factor; and These are the target Critic network and the target Actor network, respectively. The thermal stress at the next moment is obtained from the prediction formula; This is the limit for thermal stress alarm.
10. The method for optimizing AGC instructions for thermal power units based on deep reinforcement learning according to claim 1, characterized in that, Step S7 includes the following steps: Step S71: Collect historical data of the unit's long-term operation under deep peak shaving conditions, including AGC command curves, power response, steam temperature and metal temperature records. Using the reduced-order model as the core, combined with a simplified thermodynamic model that reflects the change of steam temperature with load, construct a reinforcement learning environment simulator. The reinforcement learning environment simulator can provide simulation feedback on the power, steam temperature and thermal stress state at the next moment based on the current action. Step S72: In the simulation environment, the Actor-Critic network structure and loss function defined in step S6 are used for offline pre-training until the cumulative reward converges, so that the agent learns a shaping strategy that takes into account both AGC assessment and rotor life under different initial thermal stress and command modes. Step S73: Solidify the trained Actor network parameters into an inference model, and deploy it in the upper-level optimization station of the unit's distributed control system via the OPC communication interface. During online operation, the state vector defined in step S3 is obtained in each control cycle, normalized, and then fed into the Actor network to output the action. In step S4, generate optimized load instructions and send them to the coordinated control system; Step S74: Simultaneously set up a safety monitoring logic independent of the intelligent agent. If the thermal stress value calculated in step S2 is detected in multiple consecutive control cycles... All exceeded the alarm limit. If the intelligent agent outputs abnormally, it will automatically cut off the optimized instruction path and restore the original AGC scheduling instruction direct path to ensure the safety of the unit hardware.