Main steam temperature control method for guaranteeing safe operation of steam turbine rotor

Through deep reinforcement learning agent interaction with the steam turbine, the main steam temperature is adjusted in real time, the thermal stress problem during the rapid start-up process is solved, a safe and efficient start-up process is achieved, and the operation efficiency of the turbine and the stability of the power system are optimized.

CN120159549APending Publication Date: 2025-06-17XI AN JIAOTONG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510567549.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

During the rapid start-up of the turbine, the sharp changes in the main steam temperature lead to a huge temperature gradient inside and outside the rotor, causing high-intensity thermal stress, threatening the safety and service life of the rotor.

Method used

The interaction between the agent based on deep reinforcement learning and the steam turbine is adopted, and the reward function is defined through the stress field reconstruction model, and the main steam temperature is adjusted in real time to ensure the stress state of the rotor.

Benefits of technology

It has achieved technical support to significantly shorten the start-up time, optimize the unit operation efficiency, improve the stability of the power system and green development while ensuring the safety of the rotor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120159549A_ABST
    Figure CN120159549A_ABST
Patent Text Reader

Abstract

The invention discloses a main steam temperature control method for guaranteeing safe operation of a steam turbine rotor, which relates to the technical field of steam turbines, and comprises the following steps: defining a system state, a temperature control action and a reward function, taking a neural network as an intelligent agent, and interacting with a quick starting process of a steam turbine; in the interaction process, the intelligent agent receives the system state of the steam turbine and outputs a temperature control action; the steam turbine performs temperature adjustment according to the temperature control action and feeds back rewards to the intelligent agent based on a reward function; the intelligent agent evaluates whether the temperature control action is good or bad through rewards until the accumulated rewards converge; and after the accumulated rewards converge, the intelligent agent outputs an optimal temperature control action, and the temperature of main steam in the quick starting process of the steam turbine is controlled through the optimal temperature control action. The main steam temperature can be dynamically adjusted, and the starting time can be remarkably shortened on the premise that the safety of the rotor is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of steam turbines, and particularly to a main steam temperature control method for ensuring the safe operation of a steam turbine rotor. Background Art

[0002] With the global energy structure transforming towards low-carbon, efficient and clean energy forms play an increasingly important role in the power system. Due to problems such as high carbon emissions and limited resources, traditional fossil energies are gradually being replaced by renewable energies and nuclear energy. Among them, nuclear energy, with its advantages of high energy density, low carbon emissions and strong operation stability, has become an important choice for many countries to achieve energy transformation and sustainable development. However, as the proportion of renewable energies (such as wind energy and solar energy) in the power system continues to increase, the volatility and instability of the power grid also intensify. To maintain the stable operation of the power grid, the demand for traditional generating units (such as thermal power and nuclear power) to participate in peak shaving is becoming increasingly urgent. Peak shaving can not only effectively balance the grid load, but also reduce the dependence on fossil energies, lower carbon emissions, and contribute to the realization of the "dual carbon" goal.

[0003] As the core equipment of thermal power generation and nuclear power units, steam turbines face complex and harsh working conditions during the peak shaving process. Different from long-term stable operation, during the peak shaving process, steam turbines need to start and stop frequently, and the load changes drastically, resulting in a significant increase in the alternating mechanical load and thermal load borne by the rotor. Shortening the start-up time of the steam turbine has significant advantages: on the one hand, rapid start-up can improve the peak shaving response speed of the unit and enhance the flexibility of the power grid; on the other hand, shortening the start-up time can also reduce the energy loss and pollutant emissions during the start-up process, improving the economy and environmental benefits of the unit. However, rapid start-up also brings higher technical challenges. How to achieve rapid start-up while ensuring the safety of the rotor has become a key issue in the development of steam turbine technology. During the rapid start-up process, the sharp change in the main steam temperature will generate a huge temperature gradient inside and outside the rotor, which will in turn cause high-intensity thermal stress, seriously threatening the safety and service life of the rotor. Therefore, studying the main steam temperature control method of the steam turbine rotor during the rapid start-up process is of great significance for improving the peak shaving performance, operation safety and economy of the unit.

[0004] In the prior art, the regulation of the main steam temperature usually adopts fixed heating rate control or empirical curve control. Although these methods are simple and easy to implement, they lack the ability of dynamic adjustment and cannot adjust the main steam temperature in real time according to the stress state of the rotor, which easily leads to excessive thermal stress; at the same time, the fixed heating rate results in a long start-up time, affecting the peak shaving performance of the unit. Summary of the Invention

[0005] Based on the defects existing in the above-mentioned prior art, the present invention provides a main steam temperature control method for ensuring the safe operation of a steam turbine rotor, which solves the existing problems.

[0006] The present invention adopts the following technical solutions:

[0007] The present invention provides a main steam temperature control method for ensuring the safe operation of a steam turbine rotor, including the following steps:

[0008] Collect the rotor surface temperature, temperature change rate and corresponding equivalent stress data of the steam turbine under different rapid start-up conditions. Taking the rotor surface temperature and temperature change rate of the steam turbine under different start-up conditions as inputs and the corresponding equivalent stress data as outputs, train a deep fully convolutional neural network to obtain a stress field reconstruction model;

[0009] Define the system state, temperature control action and reward function. Among them, the system state includes the main steam temperature and rotor surface temperature at different time steps, and the temperature control action includes the heating rate and corresponding heating time corresponding to the current time step;

[0010] Input the temperature control action of the previous time step and the system state of the current time step into the stress field reconstruction model to obtain the equivalent stress of the current time step, and obtain the reward function based on the equivalent stress of the current time step;

[0011] Take the neural network as an agent and interact with the rapid start-up process of the steam turbine; during the interaction process, the agent receives the system state of the steam turbine and outputs a temperature control action; the steam turbine adjusts the temperature according to the temperature control action and feedbacks a reward to the agent based on the reward function; the agent evaluates the quality of the temperature control action through the reward until the accumulated reward converges;

[0012] When the accumulated reward converges, the agent outputs the optimal temperature control action, and controls the main steam temperature of the steam turbine during the rapid start-up process through the optimal temperature control action.

[0013] Preferably, the reward function is specifically as follows:

[0014]

[0015] In the formula, r is the reward, σ p is the allowable stress, σ max is the maximum equivalent stress.

[0016] Preferably, the agent includes an actor network and a double critic network. The actor network is used to receive the system state and output the temperature control action. The double critic network is used to receive the system state and the main steam temperature control action and output the Q value, and the network parameters of the actor network and the double critic network are updated through the Q value.

[0017] Preferably, the actor network includes a first input layer, a first hidden layer, and a first output layer; the first input layer receives the system state S at the current time step t t :

[0018] h μ (0) =S t ;

[0019] In the formula, h is the output value, and μ is the actor network;

[0020] The first hidden layer performs a non-linear transformation on the system state at the current time step, specifically as follows:

[0021] h μ (l) =f(W μ (l) h μ (l-1) +b μ (l ));

[0022] In the formula, W μ (l) is the weight matrix of the l-th first hidden layer, b is the bias vector, and f(·) is the activation function;

[0023] The first output layer generates the main steam temperature control action at the current time step, specifically as follows:

[0024]

[0025] In the formula, A is the temperature control action at the current time step, α is the heating rate, and Δt is the heating time.

[0026] Preferably, the double critic network includes two critic networks, and each critic includes a second input layer, a second hidden layer, and a second output layer;

[0027] The second input layer receives the system state S at the current time step t and the main steam temperature control action, specifically as follows:

[0028]

[0029] In the formula, X t is the comprehensive input vector;

[0030] The second hidden layer performs a non - linear transformation on the comprehensive input vector, as shown below:

[0031]

[0032] In the formula, Q i is the i - th critic network;

[0033] The second output layer inputs the corresponding Q value, as shown below:

[0034]

[0035] In the formula, are the network parameters of the i - th critic network, and L is the total number of layers of the critic network.

[0036] Preferably, using the neural network as an agent and interacting with the fast start - up process of the steam turbine specifically includes the following steps:

[0037] Specifically includes the following steps:

[0038] Collect the system state at the current time step, input the system state at the current time step into the actor network, and generate the temperature control action at the current time step;

[0039] Adjust the main steam temperature through the temperature control action, and collect the reward after the main steam temperature adjustment and the system state at the next time step;

[0040] Take the system state, temperature control action, reward at the current time step, and the system state at the next time step as a single sample;

[0041] Repeat the acquisition process of a single sample to obtain multiple samples;

[0042] Input the system state and temperature control action of each sample into the double - critic network to obtain the Q value of each sample; and obtain the target Q value based on the reward of each sample;

[0043] Iteratively update the parameters of the double - critic network and the actor network through the Q value of each sample and the corresponding target Q value to obtain the optimal control parameters.

[0044] Preferably, obtaining the target Q value based on the reward of each sample, the target Q value is specifically as follows:

[0045]

[0046] In the formula, y t is the t - th target Q value, γ is the discount factor, Q' i is the i - th target critic network, A′t+1 The action generated by the target actor network in state S t+1 Preferably, the parameters of the double critic network and the actor network are iteratively updated by the Q-value of each sample and the corresponding target Q-value, specifically including the following steps:

[0047] Preferably, the parameters of the double critic network and the actor network are iteratively updated by the Q-value of each sample and the corresponding target Q-value, specifically including the following steps:

[0048] Obtain the first loss function of the double critic network, and the first loss function is as follows:

[0049]

[0050] In the formula, is the first loss function, N is the batch size, S j is the j-th state sampled from the experience replay buffer, A j is the action of the j-th sample, y j is the target Q-value of the j-th sample;

[0051] Use the gradient descent method to minimize the loss function and update the critic network parameters, specifically as follows:

[0052]

[0053] In the formula, is the gradient symbol;

[0054] Obtain the second loss function of the actor network, and the second loss function is as follows:

[0055]

[0056] In the formula, L μ (θ μ ) is the second loss function;

[0057] Use the gradient descent method to minimize the loss function and update the critic network parameters, specifically as follows:

[0058]

[0059] Preferably, the constraint range of the heating rate is α ∈ [α min , α max , α min is the lowest heating rate, α max is the highest heating rate, and the constraint range of the heating time is Δt ∈ [Δt min , Δt max , Δt min is the lowest heating time, Δt max is the highest heating time.

[0060] Compared with the prior art, the above at least one technical solution adopted by the present invention can achieve the following beneficial effects:

[0061] Based on reinforcement learning, the present invention constructs the interaction between the intelligent agent and the fast start-up process of the steam turbine, and realizes the control of the main steam temperature through the intelligent agent. Among them, the system state includes the main steam temperature and the rotor surface temperature, and the control actions include the heating rate and the corresponding heating time. Based on the stress field reconstruction model, a reward function is defined to establish the connection between the stress state of the rotor and the reward, so that the main steam temperature can be adjusted in real time according to the stress state of the rotor in the subsequent interaction process.

[0062] During the interaction process, the intelligent agent receives the system state and outputs the temperature control action. The steam turbine receives the temperature control action and feedbacks the reward to the intelligent agent according to the reward function; the intelligent agent evaluates the quality of the temperature control action through the feedback reward until the accumulated reward converges. By making the accumulated reward converge through the interaction, the intelligent agent will output the optimal temperature control action, and control the main steam temperature of the steam turbine during the fast start-up process through the optimal temperature control action. The present invention can dynamically adjust the main steam temperature, significantly shorten the start-up time on the premise of ensuring the safety of the rotor, optimize the operation efficiency of the unit, and provide technical support for the stability and green development of the power system. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0064] Figure 1 It is an algorithm framework diagram of a main steam temperature control method for ensuring the safe operation of a steam turbine rotor of the present invention;

[0065] Figure 2 It is a schematic diagram of the high-pressure and middle-pressure rotor model of the steam turbine of the present invention;

[0066] Figure 3 It is a schematic diagram of the position of the measuring points on the surface temperature of a certain section of the steam turbine rotor of the present invention;

[0067] Figure 4 It is a schematic diagram of the division of the stress field reconstruction area of the rotor of the present invention;

[0068] Figure 5 It is a schematic diagram of the process for establishing the stress field reconstruction network data set of the present invention;

[0069] Figure 6Schematic diagram for the design of the rotor stress field reconstruction model of the present invention;

[0070] Figure 7 Schematic diagram for the training process of TD3-MSTC of the present invention;

[0071] Among them, Figure 7 (a) of : Schematic diagram of the cumulative reward change during the training process, Figure 7 (b) of : Schematic diagram of the heating-up curve, Figure 7 (c) of : Schematic diagram of the maximum equivalent stress during the heating-up process;

[0072] Figure 8 Schematic diagram for the optimization design process of the heating-up curve of the present invention;

[0073] Figure 9 Schematic diagram of the parameterization of the heating-up curve in the optimization design of the present invention;

[0074] Among them, Figure 9 (a) of : Schematic diagram of the design variables, Figure 9 (b) of : Schematic diagram of the interpolated heating-up curve;

[0075] Figure 10 Schematic diagram of the optimization design result of the heating-up curve of the present invention;

[0076] Among them, Figure 10 (a) of : Schematic diagram of cold start, Figure 10 (b) of : Schematic diagram of steady-state start;

[0077] Figure 11 Schematic diagram of the cold-start heating-up curve of the present invention;

[0078] Figure 12 Schematic diagram of the curve of the maximum equivalent stress change during the cold-start heating-up process of the present invention;

[0079] Figure 13 Schematic diagram of the steady-state start heating-up curve of the present invention;

[0080] Figure 14 Schematic diagram of the curve of the maximum equivalent stress change during the warm-start heating-up process of the present invention;

[0081] Figure 15 Schematic diagram of the stability evaluation result of TD3-MSTC under cold start of the present invention;

[0082] Among them, Figure 15 (a) of : Statistical violin plot of the heating-up time, Figure 15 (b) of : Schematic diagram of the maximum equivalent stress band; Detailed implementation manners

[0083] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0084] I. Explanation of the embodiment. In order to enable those skilled in the art to fully understand how the present invention is specifically implemented, this part is an explanatory embodiment that expands and explains the technical solutions of the claims.

[0085] To solve the technical problems existing in the prior art, the present invention provides a main steam temperature control method for ensuring the safe operation of a steam turbine rotor, which specifically includes the following steps:

[0086] S1: Construct a stress field reconstruction model.

[0087] The stress field reconstruction area of the rotor is as Figure 4 shown. The input data of the model in the training set consists of the temperature and temperature change rate of the monitoring points, and the output data is the equivalent stress of the nodes in the reconstruction area. The establishment process of the training set is as Figure 5 shown.

[0088] First, it is necessary to perform parametric modeling on the main steam temperature curve during the rotor startup process, which is described as follows:

[0089] T = α i (t - t i ) + T i,b , t i ≤ t < t i+1 , i = 1, 2, K, N s ;

[0090] In the formula, T is the main steam temperature, t is the time, N s is the number of segments; α i and T i,b are the heating rate and the initial temperature of the stage within the i-th time period respectively. Therefore, the main steam temperature T is a piecewise function of time t, and within each time period, the main steam temperature and time have a linear relationship.

[0091] Then, according to the value range of the parameters, Latin hypercube sampling is used to complete the sampling of different startup conditions; finally, under different startup conditions, the temperature field and stress field of the rotor are calculated through finite element calculation, and the temperature data and stress data are extracted respectively to establish a data set. The sampling parameters are: the initial main steam temperature T0, that is, the main steam temperature after the rotor speed increase ends; the main steam temperature T f; The heating rate α within a single-segment curve and its corresponding heating time Δt; The number of segments of the main steam temperature curve.

[0092] The expressions of Latin hypercube sampling for each sampling parameter are as follows:

[0093] Let the generated sample points be {x (1) , x (2) ,..., x (N)}, and each point x (k) = {x1 (k) , x2 (k) ,..., x d (k)}. Ensure that the projection {x i (1) , x i (2) ,..., x i (N)} for each dimension i satisfies:

[0094] x i (k) ∈ [x i min + (p i (k) - 1)Δx i , x i min + p i (k) Δx i ), 1 ≤ p i (k) ≤ N;

[0095] Among them, p i (k) is the randomly shuffled interval index, and p i (k) are non-repetitive for different k. Combine the N groups of intervals for all dimensions to generate N sample points.

[0096] After sampling, use the sampling parameters as boundary conditions for aerodynamic and strength finite element calculations. After the calculations are completed, extract the measured point temperatures, heating rates, and the stresses of the nodes in the reconstruction area, and write them into the dataset. The input tensor is composed of the temperature and the temperature change rate of the monitoring points on the rotor surface at a certain moment, and the output tensor is composed of the equivalent stresses of the structured grid nodes in the reconstruction area. The min-max normalization method is used to map the data in the input tensor and the output tensor to the interval [0, 1], that is:

[0097]

[0098] Due to the large dataset, a rotor stress field reconstruction model based on a one-dimensional fully convolutional neural network is adopted. The width of the features is reduced and the depth of the features is increased through the downsampling layer, transforming the low-dimensional features of the input data into high-dimensional features. Then, the upsampling layer is used to transform the extracted high-dimensional features into the low-dimensional features of the output data to achieve the reconstruction of the stress field. The schematic diagram of the model design is shown in Figure 6 as follows.

[0099] Assume that the number of measurement points on the rotor surface is N, and the temperature and temperature change rate of each monitoring point can be expressed as:

[0100] T = [T1, T2,..., T N T ∈ R N ;

[0101] ΔT = [ΔT1, ΔT2,..., ΔT N T ∈ R N ;

[0102]

[0103] The temperature and temperature change rate are combined into an input matrix:

[0104]

[0105] The downsampling layer includes multiple convolutional layers and max pooling layers. The l-th layer convolutional operation:

[0106] Z field (l) = W field (l) * X (l-1) + b field (l) ;

[0107] where is the output of the (l - 1)-th layer, is the convolutional kernel of the l-th layer, k l is the convolutional kernel size, C l-1 and C l are the number of input and output channels respectively, is the bias vector, * represents the convolutional operation, and the activation function is ReLU:

[0108] X (l) = ReLU(Z field (l) );

[0109] The max pooling operation is performed after the convolutional layer to process the output X (l) of the convolutional layer to reduce the feature length. ​​

[0110] For each channel \(c = 1, 2, \ldots, C\) l :[[]]

[0111]

[0112] where \(k\) p is the pooling window size, \(s\) p is the pooling stride, and \(X\) (l) (c, s p ·n + m) is the value at the \((s p ·n + m)\)-th position on channel \(c\) of the output of the convolutional layer.

[0113] After connecting the convolutional layer and the max-pooling layer, the output of the \(l\)-th downsampling layer can be expressed as:[[]]

[0114]

[0115] where is the output of the \(l\)-th downsampling layer, where \(N l ' is the length of the feature after convolution and max-pooling.

[0116] The upsampling layer consists of multiple interpolation layers and a convolutional layer, and the output is a single convolutional layer.

[0117] Interpolation layer:[[]]

[0118]

[0119] where is the upsampled feature tensor,[[]] is the upsampling scale factor, and \(f upsample (·) denotes the interpolation function. For one-dimensional linear interpolation, the value at the \(n\)-th position on each channel \(c\) of the feature is calculated as:[[]]

[0120]

[0121] where

[0122]

[0123]

[0124] The convolution operation of the convolutional layer is the same as that of the downsampling layer, and the mathematical description is:[[]]

[0125] Z field (l) = W field (l) * X up (l-1) + b field (l) ;

[0126] The one-dimensional convolution calculation formula is as follows:

[0127]

[0128] where c′ = 1, 2,..., C l denotes the output channel index, n = 0, 1,..., N′ l -1 denotes the feature position index, p is the padding size, usually to keep the feature length unchanged.

[0129] The activation function is:

[0130] X (l) = σ(Z field (l) );

[0131] In the last layer of upsampling, the output is the final prediction result. The corresponding convolutional layer and activation function are:

[0132] Z field (L) = W field (L) * X (L-1) + b field (L) ;

[0133]

[0134] The expression of the loss function is as follows:

[0135]

[0136] where N is the total number of samples, D is the dimension of the output physical field parameters, M is the number of grid nodes where the physical field is discretized in space, represents the predicted value of the i-th sample, with output dimension j and spatial position k, and Y i,j,k represents the corresponding true value.

[0137] To prevent the model from overfitting, a regularization term is added here to obtain the network parameters of the stress field reconstruction network:

[0138]

[0139] where λ is the regularization coefficient and θ (l) is the parameter of the l-th layer of the network.

[0140] The gradient descent method is used to minimize the loss function and update the network parameters of the stress field prediction network:

[0141]

[0142] Among them, η field is the learning rate, which controls the step size of parameter update.

[0143] The stress distribution of the rotor can be directly predicted based on the temperature results of a small number of measuring points during the startup process.

[0144] S2: Construct a temperature control model based on deep reinforcement learning.

[0145] Referring to Figure 1 , the temperature control model (agent) includes an actor network and a double critic network.

[0146] The input of the actor network is the state vector S t . At time step t, the state vector S t contains the normalized temperature results of the previous time step, expressed as:

[0147]

[0148] Among them, T steam,t is the main steam temperature at time step t, T rotor,n,t is the temperature of the measuring point on the rotor surface at time step t, n is the measuring point number, and T f is the rated main steam temperature. Figure 2 is the high-pressure and intermediate-pressure rotor model of the steam turbine, and its measuring point layout is as Figure 3 shown.

[0149] The actor network μ maps the current state S t to the action vector A, that is, the main steam temperature adjustment:

[0150] A = μ(S t |θ μ );

[0151] Among them, A is the main steam temperature action vector at time step t, with a length of 2, including the heating rate α and the heating time Δt of each heating action. The given constraint range of the heating rate α is α ∈ [α min , α max (°C / min), and the constraint range of the heating time Δt is Δt ∈ [Δt min , Δt max (min), and θ μ is the network parameter (weights and biases) of the actor network.

[0152] The actor network μ consists of a multi-layer neural network (fully connected layer), including an input layer, a hidden layer, and an output layer.

[0153] The input layer receives the state vector S t :

[0154] h μ (0)= S t ;

[0155] For each hidden layer l = 1, 2, ..., L-1, compute:

[0156] h μ (l) = f(W μ (l) h μ (l-1) + b μ (l) );

[0157] where W μ (l) is the weight matrix of the l-th layer of the actor network, b μ (l) is the bias vector of the l-th layer of the actor network, and f(·) is an activation function (such as ReLU).

[0158] The output layer generates the control action A. Since the control action A consists of the heating rate α and the heating time Δt of each heating action, the output of the last layer passes through an activation function (usually tanh) to ensure that the output of the action is within a reasonable range. The mathematical expression is as follows:

[0159]

[0160] Perform linear transformations on the heating rate α and the heating time Δt respectively to restore them to the given constraint range:

[0161]

[0162] The double critic network is used to evaluate the Q value of the given state and action, that is, the expected cumulative reward. The two critic networks independently estimate the Q value to reduce the estimation bias. The input of each critic network includes the state vector S t and the action vector A. The two are concatenated into the combined input vector X t , that is:

[0163]

[0164] For the critic network Q i (S t , A|θ Qi ), where i = 1, 2, both use a multi-layer fully connected layer neural network. From the input layer to the first hidden layer, there are:

[0165]

[0166] where is the weight matrix of the first layer of the i-th critic network, is the bias vector of the i-th critic network, and f(·) is the activation function, usually ReLU, with the expression:

[0167] f(z) = ReLU(z) = max(0, z);

[0168] For the hidden layers l = 2, 3,..., L - 1:

[0169]

[0170] The output layer has:

[0171]

[0172] S3: Collect the system state S at the current time step t t , including the main steam temperature T steam,t and the rotor surface sensor temperature T rotor,n,t .

[0173] The actor network generates the main steam temperature control action according to the state S t

[0174] Apply the control action A to the system to adjust the main steam temperature, and the execution time is Δt.

[0175] After the system executes the action, it enters a new state S t+1 and obtains an immediate reward r.

[0176] The stress field reconstruction model receives the new temperature data and calculates the maximum equivalent stress σ inside the rotor max,t+1

[0177] For the immediate reward r, the following settings are made:

[0178]

[0179] σ p is the allowable stress of the control model, and the value is is the yield strength of the rotor material at the rated main steam temperature, and β = 0.95 is a reduction coefficient given after considering the underestimation of the maximum equivalent stress by the rotor stress field reconstruction model. When the maximum equivalent stress of the rotor exceeds the allowable stress after the action is executed, a large penalty is given to the current action (i.e., assigned by the absolute value of the stress difference), and a large negative reward is given, and the more it exceeds, the heavier the penalty; otherwise, a positive reward is given. The purpose of this design is to guide the intelligent agent to choose the fastest heating strategy as much as possible under the condition of meeting the allowable stress.

[0180] ​In addition, to accelerate the training speed and improve the performance of the control model, long-term correction of the rewards in the replay memory is introduced in the TD3 algorithm, that is

[0181] r′ = 100r if t > t max ;

[0182] where t max is the set maximum heating time. When the agent adopts a conservative heating strategy passively to avoid the huge penalty caused by excessive stress, resulting in the heating time exceeding t max , a larger penalty coefficient is given to all immediate rewards in the Markov decision process in the replay memory. Of course, the penalty coefficient is significantly smaller than the yield strength of the material, which is equivalent to making the agent realize that actions exceeding the allowable stress should still be avoided first.

[0183] The system state at the current time step, the main steam temperature control action, the immediate reward, and the system state at the next time step (S t , A, r, S t+1 ) are stored as a single sample in the experience replay buffer; the process of obtaining a single sample is repeated to obtain multiple samples.

[0184] S4: Input the system state and the main steam temperature control action of each sample into the dual critic network to obtain the Q value of each sample; and obtain the target Q value based on the immediate reward of each sample.

[0185] S5: Iteratively update the parameters of the dual critic network and the actor network through the Q value and the corresponding target Q value of each sample to obtain the optimal temperature control model.

[0186] To obtain the loss function of the critic network, the target Q value needs to be calculated, and the next action needs to be obtained from the target actor network:

[0187] A' t+1 = μ'(S t+1 |θ μ' ) + ε;

[0188] where μ' is the target actor network, which has the same structure as the current actor network, but different parameters and parameter update methods. ε is the noise for target policy smoothing. The parameter update method of the target actor network is:

[0189] θ μ′ ← τθ μ + (1 - τ)θ μ′ ;

[0190] where τ is the soft update coefficient, usually taking a small value. This makes the parameters of the target actor network a weighted average of the parameters of the current actor network, ensuring the smoothness of the update.

[0191] Calculate the target Q value:

[0192]

[0193] where r is the immediate reward, γ is the discount factor, and Q' i is the i-th target critic network.

[0194] The following settings are made for the immediate reward r:

[0195]

[0196] For each critic network Q i , there is the following loss function:

[0197]

[0198] where N is the batch size, and (S j , A j , r j , S j+1 ) is a transition sample sampled from the experience replay buffer.

[0199] The gradient of the loss function with respect to the parameters is:

[0200]

[0201] Use gradient descent to minimize the loss function and update the critic network parameters:

[0202]

[0203] The parameters θ μ = {W μ (l) , b μ (l)} of the actor network are updated by the policy gradient method of reinforcement learning, by maximizing the Q-value estimate provided by the dual critic network, as follows:

[0204]

[0205] where Q(S i , μ(S i |θ μ )|θ Q ) refers to the Q value estimated by the critic network for the state S i and the action μ(S i |θ μ ), and refers to maximizing the expected cumulative reward of the action selected by the actor network under the evaluation of the critic network.

[0206] Define the loss function of the actor network as:

[0207]

[0208] Calculate the gradient of the loss function with respect to the actor network parameters θ μ :

[0209]

[0210] Further expand it to:

[0211]

[0212] where, is the gradient of the critic network with respect to the action, is the gradient of the actor network output with respect to its parameters.

[0213] Update the parameters of the actor network using gradient descent:

[0214]

[0215] where, λ μ denotes the learning rate.

[0216] Iterate in a loop: Repeat the above steps until the policy converges and the agent learns the optimal temperature control policy.

[0217] S6: Control the main steam temperature through the optimal temperature control model.

[0218] The above is the construction process of the three types of networks,

[0219] II. Evidence of related effects of the embodiments. Some positive effects have been achieved during the research and development or use of the embodiments of the present invention, and there are indeed great advantages compared with the prior art. The following content is described in combination with the data and charts in the test process.

[0220] The design results of the parameters of the Markov decision process for the main steam temperature control task during the rotor startup process are shown in Table 1.

[0221] Table 1 Parameter settings of the TD3-MATC Markov decision process

[0222]

[0223]

[0224] State vector S t includes 5 dimensionless quantities, which are the normalized current main steam temperature and Figure 3 the normalized temperatures of the first 4 measuring points on the rotor surface in where T f is the rated main steam temperature; the action vector A has a length of 2, including the heating rate α and the heating time Δt of each heating action; the immediate reward r ≤ 0, σ p is the allowable stress of the control model, and the value is is the yield strength of the rotor material 30Cr1Mo1V at the rated main steam temperature, and β = 0.95 is a reduction factor given after considering the underestimation of the maximum equivalent stress by the rotor stress field reconstruction model.

[0225] A deep reinforcement learning model was established using the open-source deep learning framework Pytorch. The main hyperparameters of the TD3-MSTC model are shown in Table 2.

[0226] Table 2 Selection of main hyperparameters during TD3-MSTC training

[0227]

[0228] The results of the TD3-MSTC training process are as Figure 7 shown. Among them, Figure 7 (a) is the line graph of the cumulative reward change during the training process. As the training process progresses, the cumulative reward of TD3-MSTC gradually increases until it stabilizes. Due to the continuous correction of the rewards in the replay memory during the training of the improved model, the cumulative reward shows a pulsed change with a large amplitude in the first 410 iterations; when the number of iterations exceeds 410, the cumulative reward of the control model reaches a stable state and converges. The average cumulative reward for 100 rounds is -0.0295. Figure 7 (b) and (c) of Figure 7 respectively give the heating curves of 5 examples and the corresponding maximum equivalent stress change curves during the training process. Examples 1 to 5 are the heating strategies output by the control model in the 305th, 330th, 355th, 620th, and 665th iterations during the training process. It can be seen from the figure that as the training process progresses, the agent is constantly looking for shorter heating curves that meet the strength requirements. From Example 1 to Example 3, the time of the heating curve gradually decreases, but the maximum equivalent stress of Example 3 exceeds the allowable stress. Therefore, the agent further adjusts the heating strategy, increasing the heating rate at the initial moment while shortening the heating time and alleviating the increase in the maximum equivalent stress on the rotor surface after heating at subsequent moments, and finally obtaining the optimal control strategy under Example 5,

[0229] The TD3-MSTC control model can give the optimal heating actions according to the real-time temperature measurement results during the heating-up stage of the rotor starting process, forming the optimal heating curve. Therefore, in order to further verify the effectiveness of TD3-MSTC, it is necessary to compare it with the results of the optimized heating curve. Still based on the rotor stress field reconstruction model, the simulated annealing algorithm is used to carry out the optimization design of the heating curve during the rotor starting process, and the process is as Figure 8 shown.

[0230] In order to reduce the dimension of the design variables, the parameterization of the heating curve is slightly different from that of the sampling process in the stress field reconstruction network. The heating curve in the optimization design is not directly composed of key points, but is established by cubic spline interpolation based on the key points, as Figure 9 shown.

[0231] Select Figure 9 seven variables in

[0232] x = (t1, T1, t2, T2, t3, T3, t4);

[0233]

[0234] The optimization control pseudocode of the simulated annealing algorithm used in the optimization design of the heating curve is shown in Table 3.

[0235] Table 3 Optimization control pseudocode

[0236]

[0237]

[0238] After completing the optimization design, the optimal heating curves for cold start and warm start are obtained, as Figure 10 shown. Next, performance evaluation and stability analysis are carried out for the algorithm.

[0239] Performance evaluation:

[0240] The computational overheads of TD3-MSTC and the optimization design are shown in Table 4.

[0241] Table 4 Comparison of computational overheads between TD3-MSTC and the optimization design

[0242]

[0243] In terms of computational overhead, TD3-MSTC occupies more than seven times the memory of other optimization algorithms, and model training also takes more time, but these are not unacceptable. Compared with the optimized design, TD3-MSTC based on deep reinforcement learning can already initially achieve the function of online deployment. For the actual operating steam turbine rotor, after the given temperature measurement results, the time for the control model to output a single-step temperature increase action is only 0.103 s, and it has a certain stability in the face of measurement and control errors generated by mechanical equipment. This advantage is sufficient to make up for the defects brought by the huge computational overhead.

[0244] 1) Cold start

[0245] First, analyze the optimal temperature increase curves obtained by the TD3-MSTC control model and the optimized design in the cold start state. Figure 11 The preset curve, the temperature increase curves formed by TD3-MSTC and the optimized design are given. The process represented by the preset curve is as follows: after the impulse rotation ends, the steam turbine rotor warms up for 30 minutes and then enters the main steam temperature increase stage, and the sum of the warm-up time and the temperature increase time is equal to the temperature increase time under TD3-MSTC. As shown in the figure, different from the "warm-up first and then temperature increase" mode of the preset curve, in the results of TD3-MSTC and the optimized design, the main steam temperature maintains a relatively high temperature increase rate (TD3-MSTC: 3.89 °C / min, optimized design: 3.25 °C / min) in the initial period of time, and then the temperature increase rate decreases until the rated main steam temperature is reached. The temperature increase time in the optimized design result is 133.8 min, and the temperature increase time of TD3-MSTC is 3.7% shorter than that of the optimized design result, which is 128.8 min. This shows that TD3-MSTC can often select the optimal temperature increase action during the temperature increase process, reflecting the excellent control performance of the control model in the rapid temperature increase task.

[0246] Figure 12 The curve of the maximum equivalent stress on the surface of the steam turbine rotor changing with time during the start-up process corresponding to each curve under cold start is given. It can be seen that under the temperature increase curves formed by TD3-MSTC and the optimized design, the maximum equivalent stress of the steam turbine rotor during the temperature increase process remains near the allowable stress and does not exceed the allowable stress; while for the preset curve with the same temperature increase time, the maximum equivalent stress of the rotor first decreases and then increases, and finally exceeds the allowable stress. Thus, it can be seen that the traditional "warm-up - linear temperature increase" mode cannot effectively meet the rapid start-up requirements of steam turbine generator sets.

[0247] 2) Warm start

[0248] Next, conduct the performance evaluation of TD3-MSTC in the test environment of warm start. Figure 13The preset curve under warm start-up, the temperature rise curve formed by TD3-MSTC and the optimized design are given. Among them, the warm-up time of the steam turbine rotor in the preset curve under warm start-up is 20 minutes, and the sum of the warm-up time and the temperature rise time is equal to the temperature rise time of TD3-MSTC. As shown in the figure, similar to cold start-up, in the results of TD3-MSTC and the optimized design, the main steam temperature also maintains a relatively high temperature rise rate (TD3-MSTC: 4.32 °C / min, optimized design: 4.11 °C / min) for an initial period of time, and then the temperature rise rate decreases until the rated main steam temperature is reached. The temperature rise time in the optimized design result is 80.7 min, and the temperature rise time under TD3-MSTC control is 78.4 min, which is 2.9% shorter than the optimized design result. This result not only reflects the excellent performance of the TD3-MSTC control model in the task of rapid temperature rise, but also shows that TD3-MSTC still has excellent control performance when the working conditions change.

[0249] Figure 14 The curve of the maximum equivalent stress on the surface of the steam turbine rotor changing with time during the start-up process corresponding to each curve under warm start-up is given. Different from the stress change curve of cold start-up, during the temperature rise process formed by TD3-MSTC and the optimized design in warm start-up, the degree of coincidence between the maximum equivalent stress curve of the steam turbine rotor and the allowable stress decreases. For TD3-MSTC, this is due to the slight decrease in the strategy performance caused by the change of the initial conditions in the test environment; for the optimized design, this is caused by various factors such as the value range of design variables, the setting of initial values, and the direction of random perturbations. The phenomenon of falling into local optimal solutions is also an inevitable problem for greedy algorithms such as the simulated annealing algorithm. Similarly, in the preset curve with the same temperature rise time, the maximum equivalent stress of the rotor first decreases and then increases, and exceeds the allowable stress at about 73 min. This further demonstrates the defects of the "warm-up - linear temperature rise" start-up mode during the rapid start-up process of the unit.

[0250] Stability analysis:

[0251] Considering that after the model is deployed in the real environment, there will be certain deviations between the measurement results of temperature measurement points and the temperature rise control instructions of the main steam and the input and output of the control model during operation. It is necessary to consider the impact of these deviations on the model performance. Therefore, when the noises of the input and output of the actor network are 5%, 10%, and 15% respectively, an evaluation study on the stability of the control model is carried out. Figure 15 This is the stability evaluation result of TD3-MSTC under cold start-up when the noise sample size is 256, where Figure 15 (a) is the statistical violin plot of the temperature rise time satisfying the allowable stress condition under different noise levels, Figure 15 (b) is the maximum equivalent stress band during the temperature rise process.

[0252] It can be seen from Figure 15 that as the input-output noise increases, under the condition of meeting the allowable stress, the heating-up time controlled by TD3-MSTC generally shows an increasing trend. However, the median and quartile lines of the output under different noises are both within the range of 130 - 135 min, and the centralization of the data is good, indicating that when there are deviations in the measurement and control results, the control model still maintains a stable output mode and can still meet the requirements of the rapid start-up of the unit; during the heating-up process, the maximum equivalent stress band also gradually becomes wider as the noise increases. When the noise reaches 15%, the maximum equivalent stress band crosses the allowable stress line, indicating that there is a heating-up curve that does not meet the allowable stress condition at this time. When the noises are 5%, 10%, and 15% respectively, the proportions of the heating-up curves that meet the allowable stress condition are 100%, 99.8%, and 87.1% respectively, further demonstrating the excellent stability of the TD3-MSTC control model in the face of measurement and control deviations.

[0253] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.

[0254] Obviously, those skilled in the art can make various changes and deformations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and deformations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these changes and deformations.

Claims

1. A main steam temperature control method for ensuring safe operation of a steam turbine rotor, characterized in that: The following steps are involved: The rotor surface temperature, temperature change rate and corresponding equivalent stress data of the steam turbine under different rapid start-up conditions are collected. The rotor surface temperature and temperature change rate of the steam turbine under different start-up conditions are used as input, and the corresponding equivalent stress data is used as output to train the deep fully convolutional neural network and obtain the stress field reconstruction model. Define the system state, temperature control action and reward function, where the system state includes the main steam temperature and rotor surface temperature at different time steps, and the temperature control action includes the heating rate and the corresponding heating time corresponding to the current time step; The temperature control action of the previous time step and the system state of the current time step are input into the stress field reconstruction model to obtain the equivalent stress of the current time step, and the reward function is obtained based on the equivalent stress of the current time step; The neural network is used as an intelligent agent and interacts with the rapid startup process of the steam turbine. During the interaction, the intelligent agent receives the system status of the steam turbine and outputs the temperature control action. The steam turbine adjusts the temperature according to the temperature control action and feeds back rewards to the intelligent agent based on the reward function. The intelligent agent evaluates the quality of the temperature control action through rewards until the accumulated rewards converge. When the accumulated rewards converge, the agent outputs the optimal temperature control action, through which the temperature of the main steam of the turbine during the rapid startup process is controlled.

2. A main steam temperature control method for ensuring safe operation of a steam turbine rotor according to claim 1, characterized in that: The reward function is specifically as follows: In the formula, r is the reward, σ p is the allowable stress, σ max is the maximum equivalent stress.

3. A main steam temperature control method for ensuring safe operation of a steam turbine rotor as claimed in claim 2, characterized in that: The intelligent agent includes an actor network and a dual commentator network. The actor network is used to receive the system state and output the temperature control action. The dual commentator network is used to receive the system state and the main steam temperature control action and output the Q value. The network parameters of the actor network and the dual commentator network are updated through the Q value.

4. A main steam temperature control method for ensuring safe operation of a steam turbine rotor as claimed in claim 3, characterized in that: The actor network includes a first input layer, a first hidden layer and a first output layer; the first input layer receives the system state S at the current time step t t : h μ (0) =S t ; In the formula, h is the output value, μ is the actor network; The first hidden layer performs a nonlinear transformation on the system state at the current time step, as shown below: h μ (l) =f(W μ (l) h μ (l-1) +b μ (l) ); Where W μ (l) is the weight matrix of the first hidden layer, b is the bias vector, and f(·) is the activation function; The first output layer generates the main steam temperature control action for the current time step, as shown below: Where A is the temperature control action at the current time step, α is the heating rate, and Δt is the heating time.

5. A main steam temperature control method for ensuring safe operation of a steam turbine rotor as claimed in claim 4, characterized in that: The dual critic network includes two critic networks, each critic including a second input layer, a second hidden layer, and a second output layer; The second input layer receives the system state S at the current time step t And the main steam temperature control action, as shown below: Where, X t is the comprehensive input vector; The second hidden layer performs a nonlinear transformation on the integrated input vector as follows: In the formula, Q i is the i-th commentator network; The Q value corresponding to the second output layer input is as follows: In the formula, is the network parameter of the i-th critic network, and L is the total number of layers of the critic network.

6. A main steam temperature control method for ensuring safe operation of a steam turbine rotor as claimed in claim 5, characterized in that: The neural network is used as an intelligent agent and interacts with the rapid startup process of the steam turbine, specifically including the following steps: Collect the system state of the current time step, input the system state of the current time step into the actor network, and generate the temperature control action of the current time step; Adjust the main steam temperature through temperature control actions, collect the reward after the main steam temperature is adjusted and the system state of the next time step; Take the system state at the current time step, the temperature control action, the reward, and the system state at the next time step as a single sample; Repeat the process of obtaining a single sample to obtain multiple samples; Input the system state and temperature control action of each sample into the dual critic network to obtain the Q value of each sample; and obtain the target Q value based on the reward of each sample; The parameters of the dual critic network and the actor network are iteratively updated through the Q value of each sample and the corresponding target Q value to obtain the optimal control parameters.

7. A main steam temperature control method for ensuring safe operation of a steam turbine rotor as claimed in claim 6, characterized in that: The target Q value is obtained based on the reward for each sample, and the target Q value is specifically as follows: In the formula, y t is the tth target Q value, γ is the discount factor, Q' i is the i-th target critic network, A′ t+1 For state S t+1 Below, actions generated by the target actor network.

8. A main steam temperature control method for ensuring safe operation of a steam turbine rotor as claimed in claim 7, characterized in that: The parameters of the dual critic network and the actor network are iteratively updated by the Q value of each sample and the corresponding target Q value, specifically including the following steps: Get the first loss function of the dual critic network, which is as follows: In the formula, is the first loss function, N is the batch size, S j is the jth state sampled from the experience replay buffer, A j is the action of the jth sample, y j is the target Q value of the jth sample; The gradient descent method is used to minimize the loss function and update the commentator network parameters as follows: In the formula, is the gradient symbol; Get the second loss function of the actor network, which is as follows: Where, L μ (θ μ ) is the second loss function; The gradient descent method is used to minimize the loss function and update the commentator network parameters as follows:

9. A main steam temperature control method for ensuring safe operation of a steam turbine rotor as claimed in claim 4, characterized in that: The constraint range of the heating rate is α∈[α min ,α max ],α min is the minimum heating rate, α max is the maximum heating rate, and the constraint range of the heating time is Δt∈[Δt min ,Δt max ], Δt min is the minimum heating time, Δt max The maximum heating time.

Citation Information

Cited By

  • Gas turbine exhaust temperature control method, device and equipment and storage medium

    CN120426139A

  • A gas turbine exhaust temperature control method, device, equipment and storage medium

    CN120426139B