Coordinated control method for slow dynamic unknown coal-fired power generation system based on TS fuzzy and TD3

By decomposing the coal-fired power generation system into fast and slow subsystems and combining TS fuzzy and TD3 algorithms, the problems of fast and slow dynamic differences and insufficient adaptability in coordinated control of the coal-fired power generation system are solved, and efficient and stable control of complex working conditions is achieved.

CN119536167BActive Publication Date: 2025-09-23CHINA UNIV OF MINING & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411691750.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-25
Publication Date
2025-09-23
Estimated Expiration
2044-11-25

AI Technical Summary

Technical Problem

Coal-fired power generation systems have problems in coordinated control, such as obvious differences between fast and slow dynamics, incomplete dynamic information, and insufficient adaptability. Existing methods are not effective under complex working conditions.

Method used

The coal-fired power generation system is decomposed into a fast subsystem and a slow subsystem by combining TS fuzzy and TD3 algorithms. The slow subsystem is reinforced by TD3 algorithm, and the fast subsystem is processed by TS fuzzy model. Loss clipping, dynamic learning rate and batch size adjustment mechanism are introduced to optimize the training process and generate a coordinated control strategy.

Benefits of technology

It significantly improves the dynamic adaptability and robustness of the coal-fired power generation system under complex operating conditions, enhances the stability and efficiency of the control strategy, and ensures the stable operation of boilers and turbines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119536167B_ABST
    Figure CN119536167B_ABST
Patent Text Reader

Abstract

The present invention discloses a coordinated control method for a slow-dynamic unknown coal-fired power generation system based on T-S fuzzy and TD3, comprising the following steps: first, the coal-fired power generation system is decomposed into a fast subsystem and a slow subsystem using singular perturbation theory, and the original control task is decomposed into a stabilization task of the fast subsystem and a tracking task of the slow subsystem. For the fast subsystem, the slow-varying characteristics of steam pressure are utilized to select fuzzy sets on its definition domain to construct a T-S fuzzy model. The fast subsystem controller is obtained by solving the algebraic Riccati equations corresponding to multiple linear systems. For the slow subsystem, a dual-Q network is used to reduce the overestimation of the Q value, and a dynamic learning rate and batch size adjustment mechanism are introduced to accelerate training convergence. The control input of the slow subsystem is obtained by learning under the TD3 algorithm framework. The control inputs of the fast and slow subsystems are combined and acted on the original system to obtain the state information at the next moment. The intelligent agent interacts with the original system to complete collaborative optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data-driven control of coal-fired power generation systems, and mainly to a coordinated control method for slow-dynamic unknown dual-time-scale coal-fired power generation systems based on TS fuzzy and TD3 algorithms. Background Art

[0002] With the continued growth of global energy demand, coal-fired power generation, a vital component of modern power systems, ensures safe and efficient operation, crucial for grid stability. Coal-fired power generation systems exhibit highly nonlinear characteristics and strongly coupled fast and slow dynamics, making optimizing their operational strategies a key research focus. Coordinated control systems (CCSs) receive load commands from the power grid and control the boiler and steam turbine as an integrated unit, sending separate commands to coordinate the slow dynamics of the boiler with the fast dynamics of the steam turbine. Boilers have a slow response due to their large thermal inertia, while steam turbines react quickly. The primary goal of a CCS is to adjust output power to meet grid load demand while maintaining turbine pressure and boiler water level at stable levels to ensure stable unit operation. In recent years, researchers have utilized singular perturbation theory to decompose coal-fired power generation systems into fast and slow subsystems. The fast subsystem utilizes a TS fuzzy controller for rapid response and stable adjustment, while the slow subsystem utilizes the Lyapunov method to ensure asymptotic stability. However, this approach relies heavily on precise system modeling and lacks adaptability to environmental and parameter changes, limiting its application in complex operating conditions. In addition, although existing data-driven methods can adaptively adjust the output of boilers and turbines to optimize power generation efficiency and respond to load changes, these methods usually fail to fully consider the dual-time scale characteristics and may cause numerical ill-conditioning problems, thereby affecting the stability and reliability of the control system.

[0003] The integration of singular perturbation theory and reinforcement learning is a viable approach for solving the coordinated control of coal-fired power generation systems with coupled fast and slow dynamics, where precise models are difficult to construct. Singular perturbation theory decomposes a coupled fast and slow dynamic system into fast and slow subsystems, reducing the original control problem to a stabilization task for the fast subsystem and a tracking task for the slow subsystem. TD3, an advanced reinforcement learning algorithm, significantly improves training stability and convergence by introducing a dual critic network, delayed policy updates, and soft updates of the target network, thereby enhancing its ability to handle complex system dynamics. TS fuzzy modeling uses fuzzy rules to approximate nonlinear systems into multiple linear subsystems, simplifying the design of control strategies for the fast subsystem.

[0004] Therefore, there is an urgent need to develop a coordinated optimization control method for coal-fired power generation systems with self-learning ability by combining the stability of TD3 and the modeling capability of TS fuzzy. Summary of the Invention

[0005] This invention provides a coordinated control method for a slow-dynamic, unknown, dual-time-scale coal-fired power generation system based on the TS fuzzy and TD3 algorithms. This method solves the problems of significant differences in fast and slow dynamics, incomplete dynamic information, and insufficient adaptability in the coordinated control of coal-fired power generation systems. See the following description for details:

[0006] A coordinated control method for a slow-dynamic unknown dual-time-scale coal-fired power generation system based on TS fuzzy and TD3 algorithms includes the following steps:

[0007] Step 1: Decompose the original coal-fired power generation system into a fast subsystem and a slow subsystem, and decompose the original control task into a stabilization task of the fast subsystem and a tracking task of the slow subsystem;

[0008] Step 2: Model the tracking task of the slow subsystem as a Markov decision process and define a four-tuple (S, A, R, γ), where A is the action space, representing the set of operations performed by the reinforcement learning agent;

[0009] S is the state space, and the slow component of the original coal-fired power generation system state is selected to construct the slow subsystem state data;

[0010] R is the reward function; γ is the discount factor; the TD3 algorithm framework is used, and loss clipping, dynamic learning rate and batch size adjustment mechanism are introduced to optimize the training process, and the control input a of the slow subsystem is obtained. k ;

[0011] Step 3: For the stabilization task of the fast subsystem, a TS fuzzy model is established, the controller gain is obtained by solving the algebraic Riccati equation, and the fast subsystem control input u is obtained by the parallel compensation method. 2f ;

[0012] Step 4, fast system control input a k and the slow subsystem control input u 2f After the combination, the original coal-fired power generation system is acted on to obtain the state at the next moment. Return to step 2, calculate the reward information according to the reward function R, and the intelligent agent optimizes the control strategy according to the state and reward information at the next moment until the set maximum number of steps is reached.

[0013] Furthermore, step 1 specifically includes the following contents:

[0014] Step 101: Consider the following dual-time-scale coal-fired power generation system:

[0015]

[0016] Where x1 represents the slow state variables of the coal-fired power generation system, namely, steam pressure and steam water density; x2 represents the fast state variables of the coal-fired power generation system, namely, output power; u is the control input, namely, the valve opening of fuel flow, steam flow, and feed water flow; 0<ε<1 is the time scale parameter, which characterizes the degree of difference between the fast and slow dynamic time scales; f and g are vector or matrix equations of appropriate dimensions.

[0017] Step 102, let ε = 0, and get x 2s =φ s (x 1s ,u s ,t), and substituting into (35a) we can obtain the state equation of the slow subsystem

[0018]

[0019] The subscript "s" in the variable indicates its slow part.

[0020] Step 103, state x2 contains slow component x 2s and fast component x 2f , that is, x2=x 2s +x 2f , based on (35b), we can get the state equation of the fast subsystem

[0021]

[0022] The subscript "f" in the variable represents its fast part.

[0023] Furthermore, the control goal of the original system is to achieve the tracking task, that is, to ensure that the drum pressure, coal flow and steam flow can accurately track their preset target values ​​while minimizing energy consumption. Its performance indicators are defined as

[0024]

[0025] Among them, Q and R are the weight matrices of error state and control input respectively, T s is the target state, u is the control input of the system, and γ is the discount factor.

[0026] Then the performance indicators of the decomposed fast and slow subsystems are as follows:

[0027]

[0028] Among them, J f is the performance index of the fast subsystem, Q f and R f are the weight matrices corresponding to the state and control input of the fast subsystem, J s is the performance index of the slow subsystem, Q s and Rs are the weight matrices corresponding to the error state and control input of the slow subsystem, x e is the deviation between the slow component of the current state and the slow component of the target state.

[0029] Furthermore, step 2 specifically includes the following contents:

[0030] Step 201: Formulate the tracking problem of the coal-fired power generation slow subsystem under the reinforcement learning framework. It is necessary to define a four-tuple (S, A, R, γ), specifically:

[0031] (1) A is the action space, which represents the set of operations that the reinforcement learning agent can perform. For the coordination control problem, the control input can be the action of the agent, that is, a = [u 1s ,u 2s ,u 3s ] T , where u 1s ,u 2s ,u 3s Represents the slow component of the fuel flow, steam flow, and drum feedwater flow control valve opening. Each valve can rotate in the range of (0,1), and each control variable is a continuous action space, where 0 represents fully closed and 1 represents fully open.

[0032] (2) S is the state space, which represents the set of information observed by the agent. Select the slow component of the system state, the boiler drum steam pressure x 1s (kg / cm 2 ), turbine power x 2s (MW) and the fluid density x in the boiler drum 3s (kg / cm 3 ), the slow component of the current state and the target state T s The deviation between the slow components is e=[e1,e2,e3] T The complete observation information can be described as

[0033] s=[x 1s ,x 2s ,x 3s ,e1,e2,e3] T (39)

[0034] (3) r is the reward function used to evaluate the target and effect of the agent when performing actions. When the system state is far away from the target state T s The reward becomes smaller when the system state is close to the target state T s When the slow component is large, the reward becomes larger, so the reward r is set to

[0035] r=-ω1||e||-ω2a (40)

[0036] Among them, ω1 and ω2 are the tracking error penalty weight and the control consumption penalty weight, ||e|| represents the difference between the current system state and the target state T s The tracking error norm between .

[0037] (4)γ is a discount factor, which is used to measure the importance of future rewards. Specifically, the discount factor determines the degree to which the agent reduces future rewards;

[0038] Step 202: Initialize two critic networks and and an actor network π φ , and initialize with random parameters θ1, θ2, φ, and assign the initial values ​​to the target network:

[0039] θ′1←θ1,θ′2←θ2,φ′←φ (41)

[0040] Where θ′1 and θ′2 are the parameters of the target critic network, and φ′ is the parameter of the target actor network.

[0041] Step 203: When the agent is training, the current state at the current time k is represented as s k ∈S. When taking action a t After π(s)+∈, ε~N(0,σ), it transfers to the next state s k+1 and obtain rewards r from the environment k And store the data sequence (s,a,r,s′) of each time step into the experience pool;

[0042] Step 204, define Used to describe the coal-fired power generation system in state s at a specific time k. k When , execute action a k , and then execute the long-term expected control cost function of the coordinated control strategy π(s), that is,

[0043]

[0044] Step 205: Extract N (s, a, r, s′) data sequences from the experience pool for neural network training. When the state is s, select an action according to the target actor network. Using two target critic networks and Evaluate state s′ and target action The value of Q The minimum value among the values ​​to reduce the overestimation bias

[0045]

[0046] Step 206, the loss of the critic network is a measure of the difference between the predicted Q value and the target Q value

[0047]

[0048] Where B is the batch size of the extraction.

[0049] Step 207: Randomly initialize the running mean μ of the critic network loss c1 and μ c2 , decay rate β c1 and β c2 , standard deviation multiple n c Calculate the standard deviation based on the initial loss value Then, we can get the upper and lower bounds of loss based on the operation:

[0050]

[0051] Step 208: Based on the relationship between the running loss and the upper and lower limits, a new loss L is updated. c :

[0052]

[0053] Step 209: Dynamically update the network loss mean μ c1 and μ c2 , used to solve the upper and lower limits of the next loss:

[0054]

[0055] Step 210, if L critic <μ c1 , dynamically adjust the Critic network learning rate:

[0056] α c ←(1-γ c )α c (48)

[0057] Among them, γ c Tune parameters for the critic network learning rate.

[0058] If L critic >μ c1 , dynamically adjust the Critic network learning rate:

[0059] α c ←(1+γ c )α c (49)

[0060] Step 211, update critic network parameters:

[0061]

[0062] Step 212, continuously looping steps 206-211 until the training is completed and the optimal control strategy is obtained.

[0063] In step 213, the role of the actor network is to generate the optimal action for a given state, that is, to maximize the Q value output by the critic network. The loss of the actor network is defined as:

[0064]

[0065] Step 214: Randomly initialize the running mean μ of the actor network loss a1 and μ a2 , decay rate β a1 and β a2 , standard deviation multiple n c , calculate the standard deviation based on the initial loss value Then we can get the upper and lower bounds of loss:

[0066]

[0067] Step 215: Update the new loss L according to the relationship between the running loss and the upper and lower limits. a :

[0068]

[0069] Step 216: Dynamically update the network loss mean μ a1 and μ a2 , used to solve the upper and lower limits of the next loss:

[0070]

[0071] Step 217, if L actor <μ a1 , dynamically adjust the Actor network learning rate:

[0072] α a ←(1-γ a )α a (55)

[0073] Among them, γ a Tuning parameters for the Actor network learning rate.

[0074] If L actor >μ a1 , dynamically adjust the Actor network learning rate:

[0075] α a ←(1+γ a )α a (56)

[0076] Step 218, update actor network parameters:

[0077]

[0078] Step 219, continuously looping steps 213-218 until the training is completed and the optimal control strategy is obtained.

[0079] Step 220, dynamically adjust the batch size based on the current weighted loss. The weighted loss is L ac =tL actor +(1-t)L critic , where t is the weight factor, μ ac =tμ a1 +(1-t)μ c1 is the weighted average of actor and critic losses. If L ac <μ ac , then increase the batch size

[0080] B←(1+ρ)B (58)

[0081] Here, ρ is the batch size adjustment parameter.

[0082] If L ac >μ ac , then reduce the batch size

[0083] B←(1-ρ)B (59)

[0084] Furthermore, step 3 specifically includes the following contents:

[0085] Step 301: Divide the domain of the slow state variable x1 into multiple fuzzy sets and establish corresponding fuzzy rules. Each fuzzy rule is expressed as follows:

[0086] Fast subsystem rule i:

[0087] If x1 belongs to the fuzzy set

[0088] but

[0089]

[0090] Where i = 1, 2, ..., s, s is the number of fuzzy rules, A fi and B fi is the corresponding constant matrix.

[0091] Step 302: Select an appropriate fuzzy set based on the slow variable x1 and calculate its membership degree μ i (x1)

[0092]

[0093] Among them, a i is the left endpoint of the fuzzy set, b i is the midpoint of the fuzzy set, c i is the right endpoint of the fuzzy set.

[0094] Solving for membership

[0095]

[0096] Step 303: The overall TS fuzzy fast subsystem model is as follows:

[0097]

[0098] Step 304, with T f The fast subsystem is discretized for the sampling period. The discretized fast subsystem is as follows:

[0099]

[0100] in, h i (k) = h i (x1(k)).

[0101] Solve the following algebraic Riccati equation

[0102]

[0103] Among them, P i is the solution of the Riccati equation, and the feedback gain matrix K can be obtained i :

[0104]

[0105] Further, we get the fuzzy fast controller:

[0106]

[0107] Furthermore, step 4 specifically includes the following contents:

[0108] Step 401: Add the slow subsystem control input selected in step 204 and the fast subsystem control input obtained in step 305 to obtain a combined control input:

[0109] u=a k +u 2f (k) (67)

[0110] In step 402, the combined control input interacts with the original environment to obtain the next state, reward and other information, providing data support for the subsequent training of the intelligent agent. Through this interaction process, the intelligent agent continuously updates the network parameters and optimizes the control strategy.

[0111] The beneficial effects of the technical solution provided by the present invention are:

[0112] 1) By combining the TD3 algorithm with TS fuzzy modeling, the coordinated control problem of slow dynamic unknown coal-fired power generation systems is effectively addressed, and the adaptability and robustness of the control strategy to the dynamic changes of complex systems are significantly improved.

[0113] 2) By introducing a dynamic learning rate and batch size adjustment mechanism into the TD3 algorithm, the neural network training process is optimized, enabling the algorithm to automatically adjust training parameters at different stages, avoiding gradient explosion or vanishing problems, improving training stability and efficiency, and accelerating the convergence process. BRIEF DESCRIPTION OF THE DRAWINGS

[0114] Figure 1 This is a coordinated control flow chart of a slow dynamic unknown dual time scale coal-fired power generation system based on TS fuzzy and TD3 algorithm;

[0115] Figure 2 This is a diagram of the slow subsystem training;

[0116] Figure 3 is the trajectory diagram of steam pressure x1 and error e1;

[0117] Figure 4 It is the trajectory diagram of steam turbine power x2 and error e2;

[0118] Figure 5 is the trajectory diagram of the fluid density x3 and error e3 in the boiler drum;

[0119] Figure 6 It is a trajectory diagram of the valve opening that controls the input fuel flow u1, steam flow u2 and drum feed water flow u3. DETAILED DESCRIPTION

[0120] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention are described in further detail below.

[0121] The present invention provides a coordinated control method for a slow dynamic unknown dual time scale coal-fired power generation system based on TS fuzzy and TD3 algorithm. Figure 1 As shown, the method includes the following steps:

[0122] Step 1: Use singular perturbation theory to decompose the coal-fired power generation system into a fast subsystem and a slow subsystem, and decompose the original control task into the stabilization task of the fast subsystem and the tracking task of the slow subsystem. The details are as follows:

[0123] Step 101: Consider the following dual-time-scale coal-fired power generation system:

[0124]

[0125] Where x1 represents the slow state variables of the coal-fired power generation system, namely, steam pressure and steam water density; x2 represents the fast state variables of the coal-fired power generation system, namely, output power; u is the control input, namely, the valve opening of fuel flow, steam flow, and feed water flow; 0<ε<1 is the time scale parameter, which characterizes the degree of difference between the fast and slow dynamic time scales; f and g are vector or matrix equations of appropriate dimensions.

[0126] The control goal of the original system is to achieve the tracking task, that is, to ensure that the drum pressure, coal flow rate and steam flow rate can accurately track their preset target values ​​while minimizing energy consumption. Its performance indicators are defined as

[0127]

[0128] Among them, Q and R are the weight matrices of error state and control input respectively, T s is the target state, u is the control input of the system, and γ is the discount factor.

[0129] Step 102, let ε = 0, and get x 2s =φ s (x 1s ,u s ,t), and substituting into (69a) we can obtain the state equation of the slow subsystem

[0130]

[0131] The subscript "s" in the variable indicates its slow part.

[0132] Step 103, state x2 contains slow component x 2s and fast component x 2f , that is, x2=x 2s +x 2f , based on (69b), we can get the state equation of the fast subsystem

[0133]

[0134] The subscript "f" in the variable represents its fast part.

[0135] Step 104, the performance indicators of the decomposed fast and slow subsystems are as follows

[0136]

[0137] Among them, J f is the performance index of the fast subsystem, Q f and R f are the weight matrices corresponding to the state and control input of the fast subsystem, J s is the performance index of the slow subsystem, Q s and R s are the weight matrices corresponding to the error state and control input of the slow subsystem, x e is the deviation between the slow component of the current state and the slow component of the target state.

[0138] Step 2: For the tracking control problem of the slow subsystem, it is modeled as a Markov decision process. The TD3 algorithm framework is used to collect slow state data from the original system to reconstruct the slow subsystem state data. Loss clipping, dynamic learning rate, and batch size adjustment mechanisms are introduced to optimize the training process to obtain the slow subsystem control input. The details are as follows:

[0139] Step 201: Formulate the tracking problem of the coal-fired power generation slow subsystem under the reinforcement learning framework. It is necessary to define a four-tuple (S, A, R, γ), specifically:

[0140] (1) A is the action space, which represents the set of operations that the reinforcement learning agent can perform. For the coordination control problem, the control input can be the action of the agent, that is, a = [u 1s ,u 2s ,u 3s ] T , where u 1s ,u 2s ,u 3s Represents the slow component of the fuel flow, steam flow, and drum feedwater flow control valve opening. Each valve can rotate in the range of (0,1), and each control variable is a continuous action space, where 0 represents fully closed and 1 represents fully open.

[0141] (2) S is the state space, which represents the set of information observed by the agent. Select the slow component of the system state, the boiler drum steam pressure x 1s (kg / cm 2 ), turbine power x 2s (MW) and the fluid density x in the boiler drum 3s (kg / cm 3 ), the slow component of the current state and the target state T s The deviation between the slow components is e=[e1,e2,e3] T The complete observation information can be described as

[0142] s=[x 1s ,x 2s ,x 3s ,e1,e2,e3] T (73)

[0143] (3) r is the reward function used to evaluate the target and effect of the agent when performing actions. When the system state is far away from the target state T s The reward becomes smaller when the system state is close to the target state T s When the slow component is large, the reward becomes larger, so the reward r is set to

[0144] r=-ω1||e||-ω2a (74)

[0145] Among them, ω1 and ω2 are the tracking error penalty weight and the control consumption penalty weight, ||e|| represents the difference between the current system state and the target state T s The tracking error norm between .

[0146] (4)γ is a discount factor, which is used to measure the importance of future rewards. Specifically, the discount factor determines the degree to which the agent reduces future rewards;

[0147] Step 202: Initialize two critic networks and and an actor network π φ , and initialize with random parameters θ1, θ2, φ, and assign the initial values ​​to the target network:

[0148] θ1′←θ1,θ2′←θ2,φ′←φ(75)

[0149] Where θ1′ and θ2′ are the parameters of the target critic network, and φ′ is the parameter of the target actor network.

[0150] Step 203: When the agent is training, the current state at the current time k is represented as s k ∈S. When taking action a t After π(s)+∈, ε~N(0,σ), it transfers to the next state s k+1 and obtain rewards r from the environment k And store the data sequence (s,a,r,s′) of each time step into the experience pool;

[0151] Step 204, define Used to describe the coal-fired power generation system in state s at a specific time k. k When , execute action a k , and then execute the long-term expected control cost function of the coordinated control strategy π(s), that is,

[0152]

[0153] Step 205: Extract N (s, a, r, s′) data sequences from the experience pool for neural network training. When the state is s, select an action according to the target actor network. Using two target critic networks and Evaluate state s′ and target action The value of Q The minimum value among the values ​​to reduce the overestimation bias

[0154]

[0155] Step 206, the loss of the critic network is a measure of the difference between the predicted Q value and the target Q value

[0156]

[0157] Where B is the batch size of the extraction.

[0158] Step 207: Randomly initialize the running mean μ of the critic network loss c1 and μ c2 , decay rate β c1 and β c2 , standard deviation multiple n c The decay rate β c1 and β c2 A regulation factor used to dynamically adjust variable updates. By using a smooth weighted approach, the decay rate controls the degree of integration between new and historical values.

[0159] Calculate the standard deviation based on the initial loss value Then, we can get the upper and lower bounds of loss based on the operation:

[0160]

[0161] Step 208: Based on the relationship between the running loss and the upper and lower limits, a new loss L is updated. c :

[0162]

[0163] Step 209: Dynamically update the network loss mean μ c1 and μ c2 , used to solve the upper and lower limits of the next loss:

[0164]

[0165] Step 210, if Lcritic <μ c1 , dynamically adjust the Critic network learning rate:

[0166] α c ←(1-γ c )α c (82)

[0167] Among them, γ c Tune parameters for the critic network learning rate.

[0168] If L critic >μ c1 , dynamically adjust the Critic network learning rate:

[0169] α c ←(1+γ c )α c (83)

[0170] Step 211, update critic network parameters:

[0171]

[0172] Step 212, continuously looping steps 206-211 until the training is completed and the optimal control strategy is obtained.

[0173] In step 213, the role of the actor network is to generate the optimal action for a given state, that is, to maximize the Q value output by the critic network. The loss of the actor network is defined as:

[0174]

[0175] Step 214: Randomly initialize the running mean μ of the actor network loss a1 and μ a2 , decay rate β a1 and β a2 , standard deviation multiple n c ; where the decay rate β a1 and β a2 , used to dynamically adjust the adjustment factor of variable updates, and through a smooth weighted approach, the decay rate controls the degree of fusion between new values ​​and historical values;

[0176] Calculate the standard deviation based on the initial loss value Then we can get the upper and lower bounds of loss:

[0177]

[0178] Step 215: Update the new loss L according to the relationship between the running loss and the upper and lower limits. a :

[0179]

[0180] Step 216: Dynamically update the network loss mean μ a1 and μ a2 , used to solve the upper and lower limits of the next loss:

[0181]

[0182] Step 217, if L actor <μ a1 , dynamically adjust the Actor network learning rate:

[0183] α a ←(1-γ a )α a (89)

[0184] Among them, γ a Tuning parameters for the Actor network learning rate.

[0185] If L actor >μ a1 , dynamically adjust the Actor network learning rate:

[0186] α a ←(1+γ a )α a (90)

[0187] Step 218, update actor network parameters:

[0188]

[0189] Step 219, continuously looping steps 213-218 until the training is completed and the optimal control strategy is obtained.

[0190] Step 220, dynamically adjust the batch size based on the current weighted loss. The weighted loss is L ac =tL actor +(1-t)L critic , where t is the weight factor, μ ac =tμ a1 +(1-t)μ c1 is the weighted average of actor and critic losses. If L ac <μ ac , then increase the batch size

[0191] B←(1+ρ)B (92)where ρ is the batch size adjustment parameter.

[0192] If L ac >μ ac , then reduce the batch size

[0193] B←(1-ρ)B (93)

[0194] Step 3: For the stabilization problem of the fast subsystem, considering the slow variation characteristics of the slow steam pressure, select appropriate fuzzy sets within its definition domain and calculate the membership degree, establish the TS fuzzy model, solve the algebraic Riccati equation to generate the controller gain, and obtain the control input of the fast subsystem through the parallel compensation method; the details include the following:

[0195] Step 301: Divide the domain of the slow state variable x1 into multiple fuzzy sets and establish corresponding fuzzy rules. Each fuzzy rule is expressed as follows:

[0196] Fast subsystem rule i:

[0197] If x1 belongs to the fuzzy set

[0198] but

[0199]

[0200] Where i = 1, 2, ..., s, s is the number of fuzzy rules, A fi and B fi is the corresponding constant matrix.

[0201] Step 302: Select an appropriate fuzzy set based on the slow variable x1 and calculate its membership degree μ i (x1)

[0202]

[0203] Among them, a i is the left endpoint of the fuzzy set, b i is the midpoint of the fuzzy set, c i is the right endpoint of the fuzzy set.

[0204] Solving for membership

[0205]

[0206] Step 303: The overall TS fuzzy fast subsystem model is as follows:

[0207]

[0208] Step 304, with T f The fast subsystem is discretized for the sampling period. The discretized fast subsystem is as follows:

[0209]

[0210] in, hi (k) = h i (x1(k)).

[0211] Solve the following algebraic Riccati equation

[0212]

[0213] Among them, P i is the solution of the Riccati equation, and the feedback gain matrix K can be obtained i :

[0214]

[0215] Further, we get the fuzzy fast controller:

[0216]

[0217] Step 4: The fast and slow subsystem control inputs are combined and applied to the original system to obtain the next state and reward information. The agent updates the network parameters and optimizes the control strategy based on the feedback from the environment. This includes the following:

[0218] Step 401: Add the slow subsystem control input selected in step 204 and the fast subsystem control input obtained in step 305 to obtain a combined control input:

[0219] u=a k +u 2f (k) (101)

[0220] In step 402, the combined control input interacts with the original environment (69) to obtain the next state, reward and other information to provide data support for the subsequent training of the intelligent agent. Through this interaction process, the intelligent agent continuously updates the network parameters and optimizes the control strategy.

[0221] To enable those skilled in the art to better understand the present invention, the coordinated control method for a slow dynamic unknown dual-time-scale coal-fired power generation system based on TS fuzzy and TD3 algorithm is described in detail below in conjunction with specific embodiments.

[0222] When designing a coordinated controller for a coal-fired power generation system based on the combination of TD3 algorithm and TS fuzzy algorithm, a 160MW coal-fired power generation system is taken as an example. For the fast subsystem, the TS fuzzy model is used to deal with the problem of slow variables in the system model. Specifically, by setting three fuzzy sets of steam pressure - low pressure (100 to 110), medium pressure (105 to 115), high pressure (110 to 120), the triangular membership a i 、b i and c iare the left endpoint, midpoint, and right endpoint of each fuzzy set, respectively. For the slow subsystem, the TD3 algorithm framework was used. The actor and critic networks each used ReLU as the activation function, and the actor network's output layer used the Tanh activation function. The Adam optimizer was selected, with a learning rate of 1e-4, a minimum batch size of 128, and a soft update rate ξ set to 5e-3.

[0223] Set the initial state and target state to x(0) = [102, 60, 438.93] T and T s =[h1,h2,h3]=[121,90,389.92], the initial control input is a0=[0.3102,0.6711,0.3967]. The control cost function is

[0224] r=-ω1||e||-ω2a (102)

[0225] By interacting with the actual coal-fired power generation environment, system operation data was collected and used for network training. In order to verify the control effect, this study conducted multiple operation tests in this environment. The average reward curve of the slow subsystem training process is shown in Figure 2. Figure 2 As shown in , the horizontal axis represents the number of training rounds and the vertical axis represents the corresponding reward value. The solid line is the average reward of multiple runs, and the shaded area represents the variance of the return. Figure 2 It can be seen that the reward curve gradually converges after about 600 training rounds. After the training, the TS fuzzy and TD3 algorithms are applied to the coal-fired power generation system to test their ability to track the target state T s The effect of the original system state variables x1, x2, x3 and the corresponding tracking targets h1, h2, h3 and the error trajectory diagrams are as follows: Figure 3 、 Figure 4 and Figure 5 As shown, the trajectory of the system state variables under combined control is as follows Figure 6 As shown, it can be seen that the system can eventually track the given target value.

[0226] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for coordinated control of a coal-fired power generation system, characterized in that: The steps include: Step 1: Decompose the original coal-fired power generation system into a fast subsystem and a slow subsystem, and decompose the original control task into a stabilization task of the fast subsystem and a tracking task of the slow subsystem; Step 2: Model the tracking task of the slow subsystem as a Markov decision process and define a four-tuple (S, A, R, γ), where A is the action space, representing the set of operations performed by the reinforcement learning agent; S is the state space, and the slow component of the original coal-fired power generation system state is selected to construct the slow subsystem state data; R is the reward function; γ is the discount factor; the TD3 algorithm framework is used, and loss clipping, dynamic learning rate and batch size adjustment mechanism are introduced to optimize the training process, and the control input a of the slow subsystem is obtained. k ; Step 3: For the stabilization task of the fast subsystem, a TS fuzzy model is established, the controller gain is obtained by solving the algebraic Riccati equation, and the fast subsystem control input u is obtained by the parallel compensation method. 2f ; Step 4, fast system control input a k and the slow subsystem control input u 2f After the combination, the original coal-fired power generation system is acted on to obtain the state at the next moment. Return to step 2, calculate the reward information according to the reward function R, and the intelligent agent optimizes the control strategy according to the state and reward information at the next moment until the set maximum number of steps is reached.

2. A method for coordinated control of a coal-fired power generation system according to claim 1, characterized in that: The steps include: Step 101: The coal-fired power generation system is written as follows: Among them, x1 represents the slow state variables of the coal-fired power generation system, namely steam pressure and steam water density, x2 represents the fast state variables of the coal-fired power generation system, namely output power, and u is the control input, namely the valve opening of fuel flow, steam flow, and feed water flow. represents the derivative of x1, represents the derivative of x2, 0<ε<1 is the time scale parameter, which characterizes the degree of difference in the time scale between fast and slow dynamics, f, g are vector or matrix equations; Step 102, let ε = 0, and get x 2s =φ s (x 1s ,u s ,t), and put it into formula (1a) to get the state equation of the slow subsystem The subscript "s" in the variable represents the slow component in the variable, x 1s represents the slow component in variable x1, Representing variables The slow component in s represents the slow component in the variable φ, u s represents the slow component in the variable u; Step 103: The fast state variable x2 contains the slow component x 2s and fast component x 2f , that is, x2=x 2s +x 2f , based on formula (1b), the state equation of the fast subsystem is obtained The subscript "f" in the variable represents the fast component of the variable.

3. A method for coordinated control of a coal-fired power generation system according to claim 1, characterized in that: The performance index of the coal-fired power generation system is defined as Among them, Q and R are the weight matrices of error state and control input respectively, T s is the target state, u is the control input of the coal-fired power generation system, γ is the discount factor, x represents the current state, i represents the time step index, and k represents the current moment.

4. A method for coordinated control of a coal-fired power generation system according to claim 3, characterized in that: Performance index of fast subsystem J f and performance index of slow subsystem J s They are represented as follows: Among them, J f is the performance index of the fast subsystem, that is, the performance index of the system is optimized while ensuring the stability of the system; Q f and R f are the weight matrices corresponding to the state and control input of the fast subsystem respectively; J s is the performance index of the slow subsystem, i.e., optimizing the performance index while completing the tracking task; Q s and R s are the weight matrices corresponding to the error state and control input of the slow subsystem, x e is the deviation between the slow component of the current state and the slow component of the target state; Expresses the fast component x 2f The transpose of .

5. A method for coordinated control of a coal-fired power generation system according to claim 3, characterized in that: Step 2 specifically includes the following: Step 201: Model the tracking task of the slow subsystem as a Markov decision process. First, define the quaternion (S, A, R, γ), specifically: (1) A is the action space. For the coordinated control problem, the control input is the action of the agent, that is, a = [u 1s ,u 2s ,u 3s ] T , where u 1s Indicates the slow component of the fuel flow control valve opening, u 2s Indicates the slow component of the steam flow control valve opening, u 3s Represents the slow component of the drum feedwater flow control valve opening; the range of each valve rotation is (0,1), and each control variable is a continuous action space, 0 represents fully closed, and 1 represents fully open; (2) S is the state space, which represents the set of information observed by the agent; the slow component of the original system state is selected to construct the slow subsystem state data, the boiler drum steam pressure x 1s , turbine power x 2s and the fluid density x in the boiler drum 3s , the slow component x of the current state 1s 、x 2s 、x 3s With the target state T s The deviation between the slow components is e=[e1,e2,e3] T ; The complete observation information is described as s=[x 1s ,x 2s ,x 3s ,e1,e2,e3] T (6) (3) r is the immediate function, set to: r=-ω1||e||-ω2a (7) Among them, ω1 and ω2 are the tracking error penalty weight and the control consumption penalty weight, ||e|| represents the difference between the current system state and the target state T s The tracking error norm between ; (4)γ is a discount factor used to measure the importance of future rewards; Step 202: Initialize two critic networks using the TD3 algorithm framework. and and an actor network π φ , and initialize with random parameters θ1, θ2, φ, and assign the initial values ​​to the target network: θ′1←θ1,θ′2←θ2,φ′←φ (8) Among them, θ1, θ2 are critic networks and The network parameters of φ are the network parameters of π φ , θ1′ and θ2′ are the parameters of the target critic network, and φ′ is the parameter of the target actor network; Step 203: When the agent is training, the current state at time k is represented as s k ∈S; when taking action a t After π(s)+∈, ε~N(0,σ), the agent moves to the next state s k+1 And obtain the reward r at the current k moment from the environment k And store the data sequence (s,a,r,s′) of each time step into the experience pool; Step 204, define Q π (s k ,a k ): Used to describe the coal-fired power generation system in state s at a specific time k. k When , execute action a k , and then execute the long-term expected control cost function of the coordinated control strategy π(s), that is, E stands for expectation Step 205: Extract B (s, a, r, s′) data sequences from the experience pool for training the actor and critic networks. When the state is s, select an action according to the target actor network. Using two target critic networks and Evaluate state s′ and target action The value of Minimum value among values Step 206, the loss of the critic network is a measure of the difference between the predicted Q value and the target Q value Where B is the batch size of the extraction; Step 207 introduces loss clipping and randomly initializes the running mean μ of the critic network loss. c1 and μ c2 , decay rate β c1 and β c2 , standard deviation multiple n c ; Calculate the standard deviation based on the initial loss value Then, we can get the upper and lower bounds of the critic network loss based on the operation: Step 208: Based on the relationship between the running loss and the upper and lower limits, the new loss L of the critic network is updated. c : Step 209: Dynamically update the running mean μ of the network loss c1 and μ c2 , used to solve the upper and lower limits of the next loss: Step 210, if L critic <μ c1 , dynamically adjust the critic network learning rate: a c ←(1-c c )a c (15) Among them, γ c Adjust parameters for critic network learning rate; If L critic >μ c1 , dynamically adjust the critic network learning rate: a c ←(1+c c )a c (16) Step 211, update critic network parameters: Step 212, continuously looping steps 206-211 until the training is completed and the optimal control strategy is obtained; In step 213, the actor network is used to generate the optimal action for a given state, that is, to maximize the Q value output by the critic network. The loss of the actor network is defined as: Step 214: Randomly initialize the running mean μ of the actor network loss a1 and μ a2 , decay rate β a1 and β a2 , standard deviation multiple n c , calculate the standard deviation based on the initial loss value Then we can find the upper and lower bounds of the loss of the actor network: Step 215: Based on the relationship between the actor network's running loss and the upper and lower limits, the new actor network loss L is updated. a : Step 216: Dynamically update the running mean μ of the actor network loss a1 and μ a2 , used to solve the upper and lower limits of the next loss: Step 217, if L actor <μ a1 , dynamically adjust the Actor network learning rate: a a ←(1-c a )a a (22) Among them, γ a Adjust parameters for the Actor network learning rate; If L actor >μ a1 , dynamically adjust the Actor network learning rate: a a ←(1+c a )a a (23) Step 218, update actor network parameters: Step 219, continuously looping steps 213-218 until the training is completed and the optimal control strategy is obtained; Step 220, dynamically adjust the batch size according to the current weighted loss, update the batch size B, return to step 205 to re-extract the data sequence, and train the actor and critic networks to obtain the optimal execution action a k ; The weighted loss is: <h2 style=";text-align:left;direction:ltr">L<h2 style=";text-align:left;direction:ltr"> ac <h2 style=";text-align:left;direction:ltr"> =tL<h2 style=";text-align:left;direction:ltr"> actor <h2 style=";text-align:left;direction:ltr"> +(1-t)L<h2 style=";text-align:left;direction:ltr"> critic <h2 style=";text-align:left;direction:ltr"> ; Where t is the weight factor; μ ac is the weighted average of actor and critic losses μ ac =tμ a1 +(1-t)μ c1 ; If L ac <μ ac , then increase the batch size B←(1+ρ)B (25) Where ρ is the batch size adjustment parameter; If L ac >μ ac , then reduce the batch size B←(1-ρ)B (26).

6. A method for coordinated control of a coal-fired power generation system according to claim 1, characterized in that: The process includes the following steps: Step 3 is as follows: Step 301: Divide the domain of the slow state variable x1 into multiple fuzzy sets and establish corresponding fuzzy rules; Each fuzzy rule is expressed as follows: Fast subsystem rule i: If the slow state variable x1 belongs to the fuzzy set but Where i = 1, 2, ..., n, n is the number of fuzzy rules, A fi and B fi is a constant matrix; Step 302: Select an appropriate fuzzy set based on the slow state variable x1 and calculate the membership degree μ of the slow state variable x1. i (x1) Among them, a i Fuzzy set The left endpoint, b i Fuzzy set The midpoint of i Fuzzy set The right endpoint of Solving for membership j represents the index of the fuzzy set, and n represents the total number of fuzzy sets; Step 303: The overall TS fuzzy fast subsystem model is as follows: Step 304, with T f The fast subsystem is discretized for the sampling period. The discretized fast subsystem is as follows: in, h i (k) = h i (x1(k)); T f Represents the sampling period, s t The variable representing the time integral; Solve the following algebraic Riccati equation Among them, P i is the solution of the Riccati equation, and the feedback gain matrix K is obtained i , that is, the controller gain: Furthermore, the fuzzy fast controller is obtained through the parallel compensation method:

7. A method for coordinated control of a coal-fired power generation system according to claim 1, characterized in that: The method comprises the following steps: the combination of the fast system control input and the slow subsystem control input in step 4 is specifically: u=a k +u 2f (k)。