Boiler combustion strategy optimization method, system, equipment and medium
Through the adaptive width triangular convolutional neural network and the actor-criticist network introduced by the maximum entropy parameter, the boiler combustion strategy is optimized, and the problem of combustion optimization in the existing technology is easily trapped in local optimality and inability to deal with high-dimensional nonlinear combustion dampers, and the effect of boiler hot steam temperature stability, NOx emission reduction and thermal efficiency improvement is achieved.
Patent Information
- Application Number
- CN202510465112.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The existing boiler combustion optimization methods are prone to falling into local optimization under dynamic combustion conditions, and cannot effectively deal with continuous space optimization tasks. When dealing with high-dimensional nonlinear combustion dampers, the problems of poor thermal steam temperature stability, NOx emission concentration and boiler thermal efficiency cannot be guaranteed.
The triangular convolutional neural network TR-CNN with adaptive width was used to establish a predictive model of boiler thermal efficiency, NOx emissions and superheated steam temperature, and a boiler multi-objective combustion optimization agent based on the actor-critician network SAC was constructed, and the maximum entropy parameter was introduced to obtain the optimal strategy for all boiler combustion.
The parameters are streamlined through the TR-CNN model of adaptive width, reducing the inference time and avoiding local optimal problems; the introduction of the randomization strategy of maximum entropy parameters ensures the optimality and stability of combustion control, and achieves the effects of boiler hot steam temperature stability, reduced NOx emissions and improved thermal efficiency.
Smart Images

Figure CN119983323A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of combustion strategy optimization, and in particular to a boiler combustion strategy optimization method, system, equipment and medium. Background Art
[0002] When the load changes rapidly, the combustion stability in the furnace deteriorates, which can easily cause the flame center to deviate and the quality of steam parameters to deteriorate, thus endangering the safe operation of the steam turbine. At the same time, large-scale variable load operation will reduce the combustion efficiency in the furnace.
[0003] Constructing a predictive model for boiler combustion state is the first step to achieve combustion optimization. In this field, data-driven modeling methods have shown obvious advantages. Based on more than 5,000 data points collected from published experimental results, an artificial neural network (ANN) was trained for supercritical water heat transfer prediction, which promoted the study of supercritical fluid thermal cycles. Li et al. selected all operating condition data covering superheated steam temperature in an Australian coal-fired unit to establish an ELMAN network model, and used this model as a feedback compensator to calculate the predicted values of feedback gain and feedforward gain. Deep learning incorporates time dimension and feature dimension into input features, has the ability to infer boiler transient load, and is a key technology for establishing boiler dynamic models. A document constructed a multimodal hybrid mechanism and LSTM modeling method based on physical loss function to predict superheated steam, and verified the accuracy of the hybrid model under a multimodal switching strategy based on attention mechanism. However, the errors in the memory modules of the long short-term memory network LSTM and the recurrent neural network GRU will accumulate over time, affecting the model accuracy. The model needs to be updated frequently to meet the accuracy requirements. Therefore, an optimization algorithm is designed as the second step to achieve combustion optimization. The existing technology often uses genetic algorithms GA and particle swarm PSO for combustion optimization. The genetic algorithm GA encodes the combustion parameters as binary or real chromosomes, and realizes a global search of the parameter space through selection, crossover, and mutation operations. The particle swarm PSO represents the combustion parameter combination with the particle position, and guides the search direction through the individual optimal (pBest) and group optimal (gBest), thereby performing combustion optimization.
[0004] It can be seen that the existing boiler multi-objective optimization method is prone to fall into local optimality under dynamic combustion conditions when performing combustion optimization. It cannot directly handle the optimization tasks in continuous space and needs to discretize continuous actions. For deterministic strategy agents, the iteration effect is easily affected by the strategy. When dealing with high-dimensional nonlinear combustion dampers, it faces problems such as poor hot steam temperature stability, NOx emission concentration and boiler thermal efficiency that cannot be guaranteed. Summary of the invention
[0005] The purpose of the present invention is to provide a boiler combustion strategy optimization method, system, equipment and medium to solve the problems in the prior art in view of the above-mentioned deficiencies in the prior art.
[0006] The present invention specifically provides the following technical solutions: A boiler combustion strategy optimization method comprises the following steps: A prediction model of boiler thermal efficiency, NOx emission and superheated steam temperature is established by using a triangular convolutional neural network TR-CNN with adaptive width, and the prediction model is trained; A boiler multi-objective combustion optimization intelligent agent based on the actor-critic network SAC is constructed, and a maximum entropy parameter is introduced into the boiler multi-objective combustion optimization intelligent agent, with the goal of obtaining the optimal combustion strategy of all boilers. Based on the current state, any combustion strategy with a value higher than a threshold is obtained by introducing the boiler multi-objective combustion optimization intelligent agent with the maximum entropy parameter; the combustion strategy includes boiler combustion parameters; Based on the combustion strategy, the next state is output through the trained prediction model, and the reward corresponding to the current state is obtained, where the current state represents the current boiler thermal efficiency, NOx emissions and superheated steam temperature parameters, and the reward corresponding to the current state represents the stability of the superheated steam temperature; According to the current state, combustion strategy, next state and the reward corresponding to the current state, the boiler multi-objective combustion optimization agent with the maximum entropy parameter is iteratively optimized until the combustion strategy with the optimal reward is obtained as the optimal boiler combustion strategy.
[0007] Preferably, the construction process of the adaptive width triangular convolutional neural network TR-CNN is specifically as follows: For the width interval of the triangular convolutional neural network, randomly sample n-2 widths to train the full-width network and transfer the learned features to the sub-network; The soft labels obtained in the previous stage are used to train the sub-network, and the trained full-width network and the trained sub-network are used as the trained network. The trained network is used to construct a triangular convolutional neural network TR-CNN model with adaptive width.
[0008] Preferably, training the prediction model comprises: Obtain auxiliary variables related to boiler thermal efficiency and NOx concentration; Select variables whose importance to the prediction target is higher than a threshold value from the auxiliary variables, establish input variables, reconstruct the input variables into a 2D tensor, input the reconstructed input variables into the prediction model, and determine the hyperparameters of the prediction model; the input variables include feature dimension and time dimension; Among them, when selecting variables whose importance to the prediction target is higher than the threshold in the auxiliary variables, the importance expression of each variable in the auxiliary variables is: ; ; in, is the importance of each variable, GI is the Gini index, m represents the number of input features, K Indicates the types of variables in the data set. For Category k The probability of GI m Its Gini index.
[0009] Preferably, the goal is to obtain the optimal combustion strategy for all boilers, and based on the current state, a boiler multi-objective combustion optimization intelligent agent with a maximum entropy parameter is introduced to obtain any combustion strategy with a value higher than a threshold, including: After introducing the maximum entropy parameter into the boiler multi-objective combustion optimization agent, t Status at the moment s t and actions a t , the specific expression of the maximization strategy of the boiler multi-objective combustion optimization agent is: ; in, , indicating an action a t The entropy of α represents the regularization coefficient of entropy, Representation strategy; Introducing a discount factor γ , get the optimization target of the maximization strategy, the specific expression is: ; in, represents the optimization goal, represents the state distribution of the strategy, T represents the maximum moment, p is the initial state distribution, E is the expected value, R (·)and r (·) are all reward functions.
[0010] Preferably, the iterative optimization of the boiler multi-objective combustion optimization agent with the maximum entropy parameter introduced according to the current state, the combustion strategy, the next state and the reward corresponding to the current state includes: The state, action, next state and reward are combined into a tuple of replay buffers ( s t , a t ,s t+1 , r t ); Update the critic network Q using the current policy and the tuples collected in the replay buffer, and evaluate the actions output by the policy network through the updated critic network; Update the actor network based on the action value estimate of the updated critic network to output a better burning instruction; The value network V is updated through the loss function, and the evaluation value of the current state is output through the updated value network V; the output and loss function of the value network V are: ; ; in, is the output of the value network V, is the loss function, MSE is the mean square error function, and Represent the two critic networks Q respectively.
[0011] Preferably, the critic network is updated using the current strategy and the tuples collected by the replay buffer, specifically expressed as: ; in, represents the updated target, denotes the parameters of the two critic networks Q, represents the parameters of the two critic networks Q under the playback buffer, E is the expected value, for t Moment status, for t Always in action, for t Rewards at all times, for t+ 1 moment status, is the discount factor, α is the regularization coefficient of entropy, and the states also include unit load, total coal volume, main steam flow, main steam pressure, feed water flow, flue gas oxygen content, SCR inlet flue gas temperature and wind box pressure difference.
[0012] The present invention provides a boiler combustion strategy optimization system, comprising: A model building module, for establishing a prediction model of boiler thermal efficiency, NOx emission and superheated steam temperature by using a triangular convolutional neural network TR-CNN with adaptive width, and training the prediction model; A strategy acquisition module is used to construct a boiler multi-objective combustion optimization intelligent agent based on an actor-critic network SAC, and introduce a maximum entropy parameter to the boiler multi-objective combustion optimization intelligent agent, with the goal of obtaining the optimal combustion strategy for all boilers. Based on the current state, the boiler multi-objective combustion optimization intelligent agent with the maximum entropy parameter is introduced to obtain any combustion strategy with a value higher than a threshold; the combustion strategy includes boiler combustion parameters; The state output module is used to output the next state through the trained prediction model based on the combustion strategy and obtain the reward corresponding to the current state, where the current state represents the current boiler thermal efficiency, NOx emissions and superheated steam temperature parameters, and the reward corresponding to the current state represents the stability of the superheated steam temperature; The strategy optimization module is used to iteratively optimize the boiler multi-objective combustion optimization agent with the maximum entropy parameter according to the current state, combustion strategy, next state and the reward corresponding to the current state, until the combustion strategy with the optimal reward is obtained as the optimal boiler combustion strategy.
[0013] The present invention provides a computer device, comprising a memory and a processor, wherein a program is stored in the memory, and when the program is executed by the processor, the processor executes the steps of the above-mentioned boiler combustion strategy optimization method.
[0014] The present invention provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned boiler combustion strategy optimization method are implemented.
[0015] Compared with the prior art, the present invention has the following significant advantages: The present invention adopts a triangular convolutional neural network with adaptive width to construct a prediction model. Under the premise of ensuring prediction accuracy, the model parameters can be simplified through the adaptive width, which effectively reduces the reasoning time. The maximum entropy parameter is introduced into the boiler multi-objective combustion optimization intelligent agent to randomize the strategy, that is, the probability of each optimization action output is as dispersed as possible, reducing the problem of falling into local optimality, allowing the intelligent agent to achieve optimal control of combustion in the furnace through countless paths without missing any strategy. Based on the combustion strategy, the next state is output through the trained prediction model to obtain the reward corresponding to the current state. The current superheated steam temperature stability of the boiler is used as the reward function. According to the current state, combustion strategy, next state and the reward corresponding to the current state, the boiler multi-objective combustion optimization intelligent agent with the maximum entropy parameter introduced is iteratively optimized, which effectively prevents the control strategy from causing steam parameters to deteriorate and affecting the performance of the unit participating in deep adjustment. Under the premise of ensuring the stability of the hot steam temperature, the optimal boiler combustion strategy is used to achieve the effect of reducing the NOx emission concentration and ensuring the stability of the boiler thermal efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1The three-dimensional structure and combustion system diagram of the boiler of the present invention; Figure 2 It is the TR-CNN network diagram; Figure 3 This is a full-condition analysis diagram of the accuracy comparison of different models and sampling data; Figure 3 (a) is the relationship diagram between load and overheat temperature. Figure 3 (b) is the relationship diagram between NOx emission and thermal efficiency; Figure 4 Feature importance plot evaluated for the Gini index; Figure 5 Schematic diagram for data reconstruction; Figure 6 This is a framework diagram for optimizing boiler combustion decisions based on a random strategy SAC agent; Figure 7 for Figure 6 Each network structure diagram in Figure 7 (a) is the structure diagram of Stochastic Actor network, Vcritic network and Target V critic network. Figure 7 (b) is the structure diagram of Critic Q1-network and Q2-network; Figure 8 It is the loss curve of the two CNN models on the training set and the test set; Figure 8 (a) is the training loss, Figure 8 (b) is the test loss; Fig. 9 This is the prediction result of 3-layer CNN on the test set; Fig. 9 (a) is the thermal efficiency diagram of 3-layer CNN on the test set. Fig. 9 (b) is the NOx emission map of 3-layer CNN on the test set. Fig. 9 (c) is the superheated steam temperature of 3-layer CNN on the test set; Fig.10 is the prediction result diagram of TR-CNN on the test set; among them, Fig.10 (a) is the thermal efficiency diagram of TR-CNN on the test set. Fig.10 (b) is the NOx emission graph of TR-CNN on the test set. Fig.10 (c) is the superheated steam temperature map of TR-CNN on the test set; Fig.11 Iteration curve diagram of the mean rewards of the two agents; Fig.12 It is the optimization curve 1 of the three target variables; Fig.12 (a) is the thermal efficiency Figure 1 , Fig.12 (b) is the error diagram when obtaining thermal efficiency. Fig.12 (c) is NOx emissions Figure 1 , Fig.12 (d) is the error diagram when obtaining NOx emissions. Fig.12 (e) is the superheated steam temperature Figure 1 , Fig.12 (f) is the error diagram when obtaining the superheated steam temperature; Fig.13 It is the optimization curve 2 of the three target variables; Fig.13 (a) is the thermal efficiency Figure 2 , Fig.13 (b) is the error diagram when obtaining thermal efficiency. Fig.13 (c) is NOx emission Figure 2 , Fig.13 (d) is the error diagram when obtaining NOx emissions. Fig.13 (e) is the superheated steam temperature Figure 2 , Fig.13 (f) is the error diagram when obtaining the superheated steam temperature; Fig.14 A flow chart of a boiler combustion strategy optimization method provided by the present invention. DETAILED DESCRIPTION
[0017] The following is a clear and complete description of the technical solutions of the embodiments of the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0018] The object of this invention is a supercritical coal-fired boiler with a low nitrogen oxide concentric ring system (LNCFS). The target boiler is a subcritical parameter, with a single intermediate reheat, balanced ventilation natural circulation drum boiler, a positive pressure direct blowing pulverizing system, and a direct flow pulverized coal burner arranged in four corners.
[0019] like Figure 1 The figure shows the 3D structure and combustion system of the boiler. The boiler adopts a single furnace with tangential combustion at four corners. It is equipped with 5 coal mills, 4 of which are in operation and 1 is on standby, which provide fuel to the primary air nozzles of the AE layer. The secondary air accounts for the largest proportion of the air volume required for combustion in the furnace, and the 7 secondary air dampers are arranged near the primary air damper.
[0020] When the boiler is burning the designed coal, the minimum load for stable combustion without oil injection is 38% B-MCR, and the oil burner corresponds to OA / OB / OC nozzles. In order to effectively reduce the NOx concentration in the burnout zone, the boiler is equipped with a three-layer compact burnout air.
[0021] The optimization goal of the present invention is to improve boiler efficiency and reduce NOx emissions while ensuring the stability of superheater temperature. Based on the impact of various heat losses on boiler efficiency and the analysis of the NOx generation mechanism in the furnace, the combustion state and manipulated variables with high correlation with boiler thermal efficiency and NOx concentration are selected, and then a list of 32 auxiliary variables is constructed after merging them (see Table 1).
[0022] Although CNN performs well in boiler combustion modeling, the low latency and real-time feedback characteristics of the simulation environment become critical in the deep reinforcement learning combustion optimization framework. An efficient decision-making mechanism not only requires a sensitive agent, but also a lightweight boiler simulation environment can reduce the interaction cycle between the agent and the environment.
[0023] Table 1 32 auxiliary variables
[0024] The present embodiment is described in detail with reference to the accompanying drawings. Fig.14 As shown, a boiler combustion strategy optimization method in this embodiment includes the following steps: Step S1: A prediction model for boiler thermal efficiency, NOx emissions and superheated steam temperature is established using a triangular convolutional neural network TR-CNN with adaptive width, and the prediction model is trained.
[0025] The specific expression of the feature extraction process of the convolutional layer is: (1); in, y is the result of the convolution operation, is the weight of the convolution kernel, is the input feature vector, and , b is the deviation, and the maximum number of nodes in each layer of the convolutional layer is the width n , here the width n is the number of channels. For convolutional neural networks, the more channels there are, the higher the model accuracy will be, but a large number of hyperparameters will affect the inference speed, so n It plays a key role in the trade-off between model accuracy and efficiency. The performance of a narrower network structure is worse than that of a wider network. The residual between channels is shown below: (2); in, k 0 means the minimum width of the network. kThe sum of the channels The features extracted by the full-width network and the partial-width network have an upper bound on the residual , the network width is from k 0 to k n There are multiple network training processes, and the number of networks depends on the sparsity of the width. At the same time, the present invention sets the upper and lower limits of the width to 0.125 and 1, and randomly samples within this range. n -2 width for network training.
[0026] The construction process of the adaptive width triangular convolutional neural network TR-CNN is as follows: For the width interval of the triangular convolutional neural network, n-2 widths are randomly sampled to train the full-width network, and the learned features are transferred to the sub-network; the sub-network is trained using the soft labels obtained in the previous stage, and the trained full-width network and the trained sub-network are used as the trained network, and a triangular convolutional neural network TR-CNN model with adaptive width is constructed through the trained network.
[0027] Specifically, for the width interval [0.2,1], when the width coefficient is 1, it is a full-width network, and when the width coefficient is 0.2, it means that the number of neurons in the nodes of the sub-network is 20% of the full-width network. First, the full-width network is trained, the learned features are transferred to the sub-network, and the sub-network is trained with the soft labels of the previous stage to obtain the trained network, and the prediction model is constructed through the trained network. This training strategy naturally obtains knowledge refinement, and experimental results confirm that the prediction accuracy will not drop rapidly as the network width decreases.
[0028] During the network switching process, it is necessary to consider the impact of setting different widths on weight parameters. k =1, 2, 3 as an example, ignoring the deviation b, the specific expression of formula (1) is: (3); (4); (5); Simplifying further, y 1 can be converted into formula (6): (6); The network width is different. y The expected value of 1 is also different, as shown in formula (7): (7); Due to the existence of the BN layer, when switching between multiple widths y1, which reduces the prediction accuracy. In order to ensure that the expected values in other cases are close, or is essential. y The expected value of 2 is as follows: (8); when When and The weights on the filter can be transformed as shown in formula (9): (9); After completing the above feature extraction, the last convolutional layer is followed by a fully connected layer and an output layer, and the weights in the model are derived with the help of the loss function.
[0029] (10); When training a predictive model, this includes: Obtain auxiliary variables related to boiler thermal efficiency and NOx concentration.
[0030] Variables whose importance to the prediction target is higher than a threshold are selected from the auxiliary variables, input variables are established, and the input variables are reconstructed into a 2D tensor. The reconstructed input variables are input into the prediction model to determine the hyperparameters of the prediction model; the input variables include feature dimension and time dimension.
[0031] A feature importance threshold screening mechanism is adopted to select multiple variables from auxiliary variables as input variables. Multi-scale feature interaction between variables is realized through two-dimensional tensor reconstruction, and the hyperparameters of the prediction model are determined. This improves the model's sensitivity to boiler combustion and reduces the impact of exploration noise and hyperparameters.
[0032] The 20-day historical data of the variables in Table 1 were extracted from the DCS of the target boiler, with a sampling period of 1 minute, totaling about 28,800 sampling points. Through further analysis of the data, the operating intervals with a large number of missing data were eliminated, and the historical data of about 7 days of continuous operation that met the requirements were selected, totaling about 10,100 sampling points. The operating intervals of the variables during this period were also counted in Table 1. Abnormal data were processed using the 3σ-rule. Figure 3 The load, superheated steam temperature, NOx emissions and thermal efficiency curves in 7 days of historical data are shown. The sampling data covers a variety of operating conditions such as steady state and variable load, which is sufficient to support the comparison and analysis of various operating conditions during the experiment.
[0033] In addition, in order to avoid large errors in the predicted values of the model output layer and improve the training efficiency of the model, the original data is normalized by min-max scaling.
[0034] The input features determine the upper limit of the model's prediction performance. Selecting variables that are highly important to the prediction target from the auxiliary variables and establishing high-quality input features can not only improve the model training efficiency, but also avoid model overfitting. As a supervised dimensionality reduction method, RF is particularly suitable for evaluating the importance of high-dimensional nonlinear industrial variables such as boilers. In RF, the average decrease in the Gini index is considered to be the most important representation of Gini impurities. If the Gini index of a feature is larger, it means that the feature has a more significant impact on the target variable. The importance of the variable is represented by VIM, and the Gini index is represented by GI, as shown in equations (11) and (12). Specifically: When selecting variables whose importance to the predicted target is higher than the threshold in the auxiliary variables, the importance expression of each variable in the auxiliary variables is: (11); (12); in, VIM is the importance of each variable (importance), GI is the Gini index, m represents the number of input features, K Indicates the types of variables in the data set. For Category k The probability of GI m The auxiliary variables with importance higher than the threshold are used as input variables.
[0035] Assuming that the number of decision trees is n, the importance of the jth feature in the ith decision tree is shown in formula (13), and its importance in all decision trees is shown in formula (14). Formula (15) normalizes the features: (13); (14); (15); Among them, the sum of the importance of all features is 1, that is, .
[0036] Based on the Gini index, the importance of input variables to boiler thermal efficiency and NOx emissions was evaluated. With the sum of the importance reaching 0.9 as the limit, 13 and 18 key variables were selected (see Figure 4). Among them, Fuelair is fuel air and SOFA is overburned air. From the evaluation results, the importance of unit load ranks first for thermal efficiency. In actual operation, the higher the unit load, the higher the boiler thermal efficiency. Under deep peak load conditions, the boiler thermal efficiency will be significantly reduced at low load. For NOx emissions, SOFA-C has the highest variable importance, which is consistent with the fact that overburned air can effectively reduce NOx emissions.
[0037] In general, fuel air and coal mill capacity air have a greater impact on thermal efficiency, while secondary air and overburnt air have a greater impact on NOx emissions.
[0038] Table 2 Test variables and target variables
[0039] The two sets of key variables are merged and repeated variables are removed to serve as the input variables of the multi-objective prediction model, as shown in Table 2. The test variables represent the input boiler combustion parameters, and the feature dimension includes at least 22 test variables. In this embodiment, the target variables refer to the thermal efficiency, NOx emissions and superheated steam temperature at the first 20 moments. The time dimension is set to 20 and the moving step is set to 1. The prediction target is the boiler thermal efficiency, NOx emissions and superheated steam temperature at the 21st moment. The data reconstruction diagram is shown in Figure 5 As shown. At the same time, in order to facilitate data management, the target variable serial number (21-10100) is rearranged to be consistent with the window serial number (1-10080). The entire data set is divided into a training set and a test set at a ratio of 80% and 20%, and the hyperparameters of the prediction model are determined by the training set. Three model performance evaluation criteria are used: the determination coefficient R 2 , root mean square error RMSE and mean absolute error MAE, which are defined as follows: (16); (17); (18); in, , , represent the mean of the true value, predicted value and true value respectively. Indicates the total number of samples.
[0040] Table 3 Hyperparameter table
[0041] Step S2: Construct a boiler multi-objective combustion optimization agent based on the actor-critic network SAC, and introduce the maximum entropy parameter into the boiler multi-objective combustion optimization agent, with the goal of obtaining the optimal combustion strategy for all boilers. Based on the current state, the boiler multi-objective combustion optimization agent with the maximum entropy parameter is introduced to obtain any combustion strategy with a value higher than the threshold; the combustion strategy includes the boiler combustion parameters.
[0042] The variables that control combustion in the boiler have high-dimensional and nonlinear characteristics, and the controlled valves are mainly continuous variables. Therefore, it is necessary to consider an intelligent agent that is suitable for continuous action space. Intelligent agents that meet the conditions include DDPG, TD3, and SAC. However, DDPG and TD3 aim to maximize the expected reward. Under a boiler combustion state, only one optimal action is considered, and there is a possibility of missing other optimal strategies. SAC sets an entropy term in the optimization function. It seeks a strategy that can maximize both the expected reward and the entropy. Entropy is a measure of randomness in the strategy. The core idea of maximum entropy is not to miss any valuable strategy, which effectively solves the problem of deterministic strategies and improves the robustness of the strategy. For t Status at the moment s t and actions a t , the maximization strategy of SAC is shown in formula (19).
[0043] The goal is to obtain the optimal combustion strategy for all boilers. Based on the current state, a boiler multi-objective combustion optimization agent with maximum entropy parameters is introduced to obtain any combustion strategy with a value higher than the threshold, including: (19); in, , indicating an action a t entropy, α represents the regularization coefficient of entropy, which is used to control the randomness of the optimal strategy, In the SAC algorithm, in order to ensure the convergence of the reward function, a discount factor is introduced γ , obtain the optimization objective of the maximization strategy, so the optimization objective is defined as formula (20): (20); in, represents the optimization goal, represents the state distribution of the strategy, T represents the maximum moment, p is the initial state distribution, E is the expected value, R (·)and r (·) are all reward functions.
[0044] This goal is equivalent to maximizing the state-action pair ( s t , a t )'s discounted expected return and entropy of future states. SoftActor-Critic uses a random strategy, which has certain advantages over deterministic strategies. For example, the increased randomness of the strategy can prevent the strategy from converging to the local optimum too early, thereby encouraging the agent to conduct more extensive exploration to achieve the goal of approaching the optimal solution. If multiple actions to control boiler combustion are equally important, the agent will set the same probability weights for these actions in the strategy distribution to ensure the stability of the combustion strategy.
[0045] Step S3: Based on the combustion strategy, the next state is output through the trained prediction model, and the reward corresponding to the current state is obtained, wherein the current state represents the parameters of the current boiler thermal efficiency, NOx emissions and superheated steam temperature, and the reward corresponding to the current state represents the stability of the superheated steam temperature.
[0046] Step S4: Iteratively optimize the boiler multi-objective combustion optimization agent with the maximum entropy parameter according to the current state, combustion strategy, next state and the reward corresponding to the current state, until the combustion strategy with the optimal reward is obtained as the optimal boiler combustion strategy.
[0047] SAC is an off-policy reinforcement learning algorithm. Its structure consists of three types of functions and five networks, including policy function, value function and state value function. The random action network is used to approximate the policy function, the v critic network is used to approximate the state value function, and the Q network is used to approximate the action value function. Its algorithm steps include policy evaluation, policy improvement, state value network update and entropy weight adjustment. The specific details are as follows: The multi-objective combustion optimization agent of the boiler with the maximum entropy parameter is iteratively optimized according to the current state, combustion strategy, next state and the reward corresponding to the current state, including: The state, action, next state and reward are combined into a tuple of replay buffers ( s t , a t , s t+1 , r t ).
[0048] The critic network Q is updated using the current strategy and the tuples collected in the experience replay buffer. The action output by the policy network is evaluated through the updated critic network. The actor network is updated according to the action value estimate of the updated critic network to output a better combustion instruction, that is, the combustion instruction is output for the current state of the boiler through the policy evaluation network Q, including: Policy evaluation: Use the current policy and the mini-batch data collected in the experience replay buffer to update the critic network, where the critic network is the policy evaluation network Q. The update goal is to minimize the soft Bellman residual, specifically: (twenty one); in, represents the updated target, denotes the parameters of the two critic networks Q, represents the parameters of the two critic networks Q under the playback buffer. The dual Q network structure is used to prevent overestimation of the Q value and then select the smallest of the two Q values. E is the expected value. for t Moment status, for t Always in action, for t Rewards at all times, for t+ 1 moment status, is the discount factor, α is the regularization coefficient of entropy, D A tuple representing the playback buffer; the state represents the parameters of boiler combustion, including unit load, total coal quantity, main steam flow, main steam pressure, feed water flow, flue gas oxygen content, SCR inlet flue gas temperature, wind box pressure difference, boiler thermal efficiency, NOx emissions and superheated steam temperature.
[0049] Policy update: Update the actor network based on the action value estimate of the critic network to output a better burning instruction; the update of the policy network parameters is to minimize the KL divergence, as shown in formula (22). By weakening the partition function Contribution to the gradient, the loss function Simplified to formula (23): (twenty two); (twenty three); In order to convert random sampling into differentiable deterministic operations, a reparameterization technique is used when the policy evaluation network outputs actions. The action sampling with Gaussian distribution is shown in (24): (twenty four); in, represents the mean of the strategy distribution, represents the variance of the policy distribution, represents the added noise.
[0050] The value network V is updated through the loss function to make its estimated state value more accurate, and the evaluation value of the current state is output through the updated value network V; the output and loss function of the value network V are: (25); (26); in, is the output of the value network V, is the loss function, MSE is the mean square error function, and Represent the two critic networks Q respectively.
[0051] The current state value is evaluated through the updated value network V, and the entropy weight is adjusted based on the entropy of the current policy. Entropy weight adjustment: A key innovation in the SAC algorithm is to automatically adjust the entropy weight. α , to adapt to different tasks. This adaptive adjustment mechanism ensures that while maintaining sufficient exploration, effective strategies can also be effectively learned. As shown in formula (27), the entropy weight is adjusted based on the entropy of the current strategy. α .
[0052] (27); in, D A tuple representing a playback buffer, Indicates that the condition The following strategy.
[0053] Under the premise of safe operation, improving boiler thermal efficiency and reducing pollutant emissions are the research focuses of this invention. Figure 6 and Figure 7 The figure shows the boiler combustion decision optimization framework based on the random strategy SAC agent. Among them, TR-CNN acts as a boiler simulation environment and interacts with the SAC agent. The agent adopts the corresponding combustion strategy according to the current state of the boiler, and the boiler simulation system TR-CNN outputs the next state s t+1 , and feedback reward r t The tuple that makes up the playback buffer ( s t , a t , s t+1 , r t), for experience playback and sampling. According to the requirements of boiler combustion strategy optimization, the present invention designs the strategy network and value network (Q and V).
[0054] The strategy network is responsible for outputting combustion instructions according to the current state of the boiler. Based on the analysis of factors affecting combustion in the furnace, 11 variables including unit load, total coal volume, main steam flow, main steam pressure, feed water flow, flue gas oxygen content, SCR inlet flue gas temperature, wind box pressure difference, boiler thermal efficiency, NOx emissions and superheated steam temperature are selected as the characteristics of the boiler combustion state. In order to optimize the coal and air distribution methods, the variables in Table 2 are selected. x 11 - x 13 Three fuel air openings, x 14 and x 17 Four secondary air damper openings, two OFA damper openings, SOFA-A and SOFA-B, x 21 - x 22 The two coal mill inlet primary air volumes have a total of 12 controlled variables as combustion control variables. The strategy network is designed as a four-layer fully connected neural network. The network input layer consists of 11 neurons, corresponding to 11 variables that characterize the boiler combustion state. The output layer consists of 24 neurons, corresponding to the mean and covariance of the 12 controlled variables. Each of the two hidden layers has 512 neurons.
[0055] The Q-network is responsible for evaluating the actions output by the strategy evaluation network. Its input is the state-action pair, and its output is the score of the current boiler state - the Q value. In the present invention, the input of the Q-network includes 11 states and 12 actions, a total of 23 variables. The Q-network is designed as a four-layer fully connected neural network, with 512 neurons in each of the two hidden layers.
[0056] The input features of V-network and the policy evaluation network Q are the same, and the output is the evaluation of the current state. The evaluation value is the expected cumulative reward. The network structure of V-network is different from that of the policy evaluation network only in the output layer, and the other structures are the same.
[0057] The boiler simulation environment and SAC agent were tested and analyzed. First, the prediction performance of the boiler simulation environment based on TR-CNN was evaluated, and the optimal width coefficient W value was explored. Then, in order to compare the decision-making ability of the proposed agent, the combustion optimization results were analyzed through the interaction between TR-CNN and SAC.
[0058] In order to explore the optimal performance of the multi-objective prediction model, this section conducted comparative experiments on the test set by setting different width coefficient W values of TR-CNN (W=0.125, 0.25, 0.5, 0.75, 1.0). Since there is a trade-off between accuracy and inference speed, the inference time of 2016 test samples is also counted. Among them, the hyperparameters of TR-CNN continue to use the determined values. From the experimental results, as the width coefficient decreases, the model parameter usage will decrease, and the calculation time will be significantly reduced. However, the model prediction error gradually increases. When the width coefficient is reduced from W=0.25 to W=0.125, although the inference time is slightly reduced, the prediction error is significantly reduced. Compared with a larger width coefficient, W=0.25 significantly reduces the calculation time, and the inference time is 11.894s, which is 28.92% lower than the inference time when the width coefficient is 1. Comprehensively considering the inference time and prediction accuracy, the subsequent TR-CNN comparative experiment of the present invention selects a width coefficient of 0.25.
[0059] Table 4 RMSE of TR-CNN predicting various target variables under different width coefficients
[0060] In order to verify the prediction accuracy of the multi-target model, a 3-layer CNN was selected as the baseline model and compared with TR-CNN. The loss curves of the two CNN models on the training set and test set are shown in Figure 2. Figure 8 As shown in the figure, the loss value on the training set is lower than that on the test set. Whether in the training set or the test set, the loss value of TR-CNN is lower than that of the baseline model 3-layer CNN, indicating that the convergence effect of TR-CNN is better than that of 3-layer CNN. Correspondingly, the prediction results of 3-layerCNN and TR-CNN on the test set are shown in the figure. Fig. 9 and Fig.10 As shown in the thermal efficiency prediction results, Fig. 9 The test results of (a) are scattered on both sides of the perfect line, and there is a large error between the true value and the predicted value. Fig.10 The test results of Figure (a) are concentrated near the perfect line, and the actual value is close to the predicted value. The same prediction results also apply to NOx emissions and superheated steam temperature.
[0061] The prediction performance of the baseline model and TR-CNN is analyzed by combining the prediction indicators. Taking superheated steam temperature as an example, the RMSE of the baseline model is 6.021℃, the MAE is 4.741℃, and the R 2 The RMSE of TR-CNN is 4.922℃, MAE is 3.495℃, and R 2The RMSE and MAE decreased by 18% and 26% respectively. 2 It is improved by 3%. It can be seen that the model proposed in the present invention significantly improves the prediction performance compared with the baseline model, and the prediction accuracy meets the requirements of interaction with the intelligent agent.
[0062] To verify the performance of the SAC agent in combustion decision-making, it was compared with the traditional reinforcement learning agent-DDPG on the test set. Through the interaction of the two agents with the simulation environment, the optimization effect on boiler thermal efficiency, NOx emissions and superheated steam temperature was analyzed. The agent-related hyperparameters are shown in Table 5. The agent iteratively learns through the scheme tested in training, that is, as the training progresses, the strategy is verified every 200 rounds, and the total step length of the test experiment is 6×104.
[0063] Table 5 Hyperparameters of the two agents
[0064] With the iterative curve of the average reward of the two agents (see Fig.11 ), analyze the learning process of the agent. From the whole learning process, the reward mean of the two agents converged at the end, and the reward obtained by SAC was higher than that of DDPG, indicating that SAC has mastered the combustion decision-making level higher than DDPG. From the details, it is worth noting that in the early stage of the interaction between the agent and the simulation environment, DDPG obtained higher rewards than SAC, indicating that DDPG learned faster in the initial stage. At the end of the interaction, the reward obtained by SAC was significantly higher than that of DDPG. The highest reward of the SAC agent appeared at the 5.32×104 step, with a reward value of 475.206.
[0065] The optimization results of the two agents for thermal efficiency and NOx emissions are statistically analyzed. The optimization curves of the three objective variables are shown in Fig.12 and Fig.13 As shown. For the DDPG agent, the optimization interval before and after the thermal efficiency is optimized is [0.069, 0.303], that is, the thermal efficiency of all samples is improved. Before and after the NOx emission is optimized, the optimization interval is [-50.239, 17.201], that is, the NOx concentration is reduced by 50.239 mg / m3 at most. At the same time, there are also sample points where the NOx concentration increases, and the maximum increase is 17.201 mg / m3. According to Fig.12Figure (d) shows that among the 2016 test samples, 202 samples failed to optimize NOx emissions, accounting for 10.02%. For the SAC agent, the optimization interval before and after the thermal efficiency was optimized was [0.229, 0.462]. Similarly, the thermal efficiency of all samples was improved. Before and after the NOx emissions were optimized, the optimization interval was [-59.739, 4.061], that is, the NOx concentration was reduced by 59.739 mg / m3 at most. At the same time, there were also sample points where the NOx concentration increased, with a maximum increase of 4.061 mg / m3. According to Fig.13 According to the statistics of Figure (d), 13 samples failed to optimize NOx emissions, accounting for 0.64%. Fig.12 Figure (e) Fig.12 The (f) graph and Fig.13 Figure (e) Fig.13 From the statistical superheated steam temperature data in Figure (f), both agents can ensure the stability of superheated steam temperature in combustion optimization. From the optimization mean of performance variables, DDPG improves thermal efficiency by 0.196% and reduces NOx emissions by 10.549 mg / m 3 SAC improves thermal efficiency by 0.357% and reduces NOx emissions by 20.244 mg / m 3 From the perspective of the average thermal efficiency, SAC is 0.161% higher than DDPG, and from the perspective of the average NOx emission, SAC is 9.695 mg / m3 higher than DDPG. SAC has a better optimization effect.
[0066] In order to analyze the physical factors behind the intelligent agent's boiler combustion decision, the optimization trend of each performance variable was analyzed in combination with the load curve. Fig.12 Figure (a) and Fig.13 Figure (a) shows that at the 1000-1100 test points, the boiler is close to full load, the combustion efficiency in the furnace is also very high, and the boiler thermal efficiency is also at a high level. However, the space for optimizing thermal efficiency becomes smaller. Fig.12 Figure (b) and Fig.13 It can be clearly seen from Figure (b) that in this range, the amplitude of the improvement in thermal efficiency is generally very small, and the lowest value also appears at the 1094th test point. Near the 1500-1800th test point, the boiler is close to 30% of the rated load. The lower the load, the more unstable the combustion in the furnace, which can easily cause the flame center to deviate. On the one hand, the airflow scours the heating surface, and the temperature deviation of the flue gas on both sides of the furnace outlet becomes larger, which will cause water-cooled wall wear and even tube burst, local overheating of the convection heating surface and other consequences. The local high temperature of the water-cooled wall and other convection heating surfaces causes fluctuations in the steam side parameters, which has a serious impact on the safety of the boiler. From Fig.12 Figure (e) and Fig.13 As can be seen from Figure (e), the superheated steam temperature in this range fluctuates greatly. Fig.12 The (f) graph and Fig.13 As can be seen from Figure (f), due to the large fluctuation of superheated steam temperature, the two agents also showed a large deviation trend when stabilizing the superheated steam temperature in this interval. On the other hand, under low load, the NOx reduction environment in the burnout zone is unbalanced, and the NOx emission at the SCR inlet increases significantly and fluctuates violently in the 1500-1800 test interval (see Fig.12 Figure (c) and Fig.13 This harsh combustion atmosphere provides more room for combustion optimization. For example, the two agents achieved the maximum value of thermal efficiency improvement and NOx emission reduction near the 1800th test point (see Fig.12 Figure (b) Fig.12 Figure (d) and Fig.13 Figure (b) Fig.13 (d) Figure ).
[0067] Based on the above method, the present invention provides a boiler combustion strategy optimization system, including: a model building module, a strategy acquisition module, a state output module and a strategy optimization module.
[0068] Among them, the model construction module is used to establish a prediction model of boiler thermal efficiency, NOx emissions and superheated steam temperature by using a triangular convolutional neural network TR-CNN with adaptive width, and train the prediction model; the strategy acquisition module is used to construct a boiler multi-objective combustion optimization agent based on the actor-critic network SAC, and introduce the maximum entropy parameter to the boiler multi-objective combustion optimization agent, with the goal of obtaining the optimal combustion strategy for all boilers. Based on the current state, the boiler multi-objective combustion optimization agent with the maximum entropy parameter is used to obtain any combustion strategy with a value higher than the threshold; the combustion strategy includes boiler combustion parameters; the state output module is used to output the next state based on the combustion strategy through the trained prediction model, and obtain the reward corresponding to the current state, wherein the current state represents the current parameters of boiler thermal efficiency, NOx emissions and superheated steam temperature, and the reward corresponding to the current state represents the stability of the superheated steam temperature; the strategy optimization module is used to iteratively optimize the boiler multi-objective combustion optimization agent with the maximum entropy parameter according to the current state, combustion strategy, next state and the reward corresponding to the current state, until the combustion strategy with the optimal reward is obtained as the optimal boiler combustion strategy.
[0069] The present invention also provides a computer device, including a memory and a processor. The memory stores a program. When the program is executed by the processor, the processor executes the steps of a boiler combustion strategy optimization method.
[0070] According to the disclosed embodiments, a computing device may communicate with one or more external devices (e.g., keyboards, pointing devices, Bluetooth communications, etc.), or with any device (e.g., routers, modems, etc.) that enables a computing device to communicate with one or more other computing devices.
[0071] The present invention also provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of a boiler combustion strategy optimization method are implemented.
[0072] According to the disclosed embodiments, the storage medium may be a non-volatile computer-readable storage medium, such as but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, the storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0073] The above content is a further detailed description of the present invention in combination with a specific preferred embodiment. For technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as belonging to the protection scope of the present invention.
Claims
1. A boiler combustion strategy optimization method, characterized in that: The steps include: A prediction model of boiler thermal efficiency, NOx emission and superheated steam temperature is established by using a triangular convolutional neural network TR-CNN with adaptive width, and the prediction model is trained; A boiler multi-objective combustion optimization intelligent agent based on the actor-critic network SAC is constructed, and a maximum entropy parameter is introduced into the boiler multi-objective combustion optimization intelligent agent, with the goal of obtaining the optimal combustion strategy of all boilers. Based on the current state, any combustion strategy with a value higher than a threshold is obtained by introducing the boiler multi-objective combustion optimization intelligent agent with the maximum entropy parameter; the combustion strategy includes boiler combustion parameters; Based on the combustion strategy, the next state is output through the trained prediction model, and the reward corresponding to the current state is obtained, where the current state represents the current boiler thermal efficiency, NOx emissions and superheated steam temperature parameters, and the reward corresponding to the current state represents the stability of the superheated steam temperature; According to the current state, combustion strategy, next state and the reward corresponding to the current state, the boiler multi-objective combustion optimization agent with the maximum entropy parameter is iteratively optimized until the combustion strategy with the optimal reward is obtained as the optimal boiler combustion strategy.
2. A boiler combustion strategy optimization method according to claim 1, characterized in that: The construction process of the adaptive width triangular convolutional neural network TR-CNN is specifically as follows: For the width interval of the triangular convolutional neural network, randomly sample n-2 widths to train the full-width network and transfer the learned features to the sub-network; The soft labels obtained in the previous stage are used to train the sub-network, and the trained full-width network and the trained sub-network are used as the trained network. The trained network is used to construct a triangular convolutional neural network TR-CNN model with adaptive width.
3. A boiler combustion strategy optimization method according to claim 1, characterized in that: Training the prediction model includes: Obtain auxiliary variables related to boiler thermal efficiency and NOx concentration; Select variables whose importance to the prediction target is higher than a threshold value from the auxiliary variables, establish input variables, reconstruct the input variables into a 2D tensor, input the reconstructed input variables into the prediction model, and determine the hyperparameters of the prediction model; the input variables include feature dimension and time dimension; Among them, when selecting variables whose importance to the prediction target is higher than the threshold in the auxiliary variables, the importance expression of each variable in the auxiliary variables is: ; ; in, is the importance of each variable, GI is the Gini index, m represents the number of input features, K Indicates the types of variables in the data set. For Category k The probability of GI m Its Gini index.
4. A boiler combustion strategy optimization method according to claim 1, characterized in that: The objective of obtaining the optimal combustion strategy for all boilers is to obtain any combustion strategy with a value higher than a threshold by introducing a boiler multi-objective combustion optimization agent with a maximum entropy parameter based on the current state, including: After introducing the maximum entropy parameter into the boiler multi-objective combustion optimization agent, t Status at the moment s t and actions a t , the specific expression of the maximization strategy of the boiler multi-objective combustion optimization agent is: ; in, , indicating an action a t The entropy of α represents the regularization coefficient of entropy, Representation strategy; Introducing a discount factor γ , get the optimization target of the maximization strategy, the specific expression is: ; in, represents the optimization goal, represents the state distribution of the strategy, T represents the maximum moment, p is the initial state distribution, E is the expected value, R (·)and r (·) are all reward functions.
5. A boiler combustion strategy optimization method according to claim 4, characterized in that: The iterative optimization of the boiler multi-objective combustion optimization intelligent agent introducing the maximum entropy parameter according to the current state, the combustion strategy, the next state and the reward corresponding to the current state includes: The state, action, next state and reward are combined into a tuple of replay buffers ( s t , a t , s t+1 , r t ); Update the critic network Q using the current policy and the tuples collected in the replay buffer, and evaluate the actions output by the policy network through the updated critic network; Update the actor network based on the action value estimate of the updated critic network to output a better burning instruction; The value network V is updated through the loss function, and the evaluation value of the current state is output through the updated value network V; the output and loss function of the value network V are: ; ; in, is the output of the value network V, is the loss function, MSE is the mean square error function, and Represent the two critic networks Q respectively.
6. A boiler combustion strategy optimization method according to claim 5, characterized in that: The critic network is updated using the current strategy and the tuples collected in the replay buffer. The specific expression is: ; in, represents the updated target, denotes the parameters of the two critic networks Q, represents the parameters of the two critic networks Q under the playback buffer, E is the expected value, for t Moment status, for t Always in action, for t Rewards at all times, for t+ 1 moment status, is the discount factor, α is the regularization coefficient of entropy, and the states also include unit load, total coal volume, main steam flow, main steam pressure, feed water flow, flue gas oxygen content, SCR inlet flue gas temperature and wind box pressure difference.
7. A boiler combustion strategy optimization system, characterized in that: include: A model building module, for establishing a prediction model of boiler thermal efficiency, NOx emission and superheated steam temperature by using a triangular convolutional neural network TR-CNN with adaptive width, and training the prediction model; A strategy acquisition module is used to construct a boiler multi-objective combustion optimization intelligent agent based on an actor-critic network SAC, and introduce a maximum entropy parameter to the boiler multi-objective combustion optimization intelligent agent, with the goal of obtaining the optimal combustion strategy for all boilers. Based on the current state, the boiler multi-objective combustion optimization intelligent agent with the maximum entropy parameter is introduced to obtain any combustion strategy with a value higher than a threshold; the combustion strategy includes boiler combustion parameters; The state output module is used to output the next state through the trained prediction model based on the combustion strategy and obtain the reward corresponding to the current state, where the current state represents the current boiler thermal efficiency, NOx emissions and superheated steam temperature parameters, and the reward corresponding to the current state represents the stability of the superheated steam temperature; The strategy optimization module is used to iteratively optimize the boiler multi-objective combustion optimization agent with the maximum entropy parameter according to the current state, combustion strategy, next state and the reward corresponding to the current state, until the combustion strategy with the optimal reward is obtained as the optimal boiler combustion strategy.
8. A computer device, characterized in that: It comprises a memory and a processor, wherein a program is stored in the memory, and when the program is executed by the processor, the processor executes the steps of a boiler combustion strategy optimization method as claimed in any one of claims 1 to 6.
9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a boiler combustion strategy optimization method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Coal-fired power plant boiler ash deposition prediction method
CN115759279A
Boiler combustion control method under deep adjustment
CN117404650A
Automatic control method for emission reduction of multi-source biomass blended combustion flue gas of coal-fired boiler
CN117471906A
Deep learning combustion optimization control method and system based on working condition of combustor
CN118466427A
Multi-objective combustion optimization method based on economic predictive control
CN119802566A
Cited By
Copper recovery treatment system and method for flameless combustion of copper-containing sludge
CN120292514A
Copper recovery treatment system and method for copper-containing sludge non-flame combustion
CN120292514B
Vacuum furnace heating strategy dynamic adjustment method based on reinforcement learning driving
CN120993783A
Multi-stage dynamic heterogeneous proxy model combustion optimization method for thermal power generating unit boiler system
CN121503234A
Boiler combustion multi-target cooperative control method based on PINN and reinforcement learning
CN121557514A