A method, system, device and medium for optimizing a boiler combustion strategy

By adopting the triangular convolutional neural network TR-CNN with adaptive width and a multi-objective combustion optimization agent based on the actor-criticist network SAC in boiler combustion optimization, the problem of boiler combustion optimization in the existing technology is solved, and the boiler thermal steam temperature stability, NOx emission reduction and thermal efficiency improvement are achieved.

CN119983323BActive Publication Date: 2025-06-20CHINA UNIV OF MINING & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510465112.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-06-20
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The existing boiler combustion optimization methods are prone to falling into local optimization under dynamic combustion conditions, and cannot effectively deal with continuous space optimization tasks. When dealing with high-dimensional nonlinear combustion dampers, the problems of poor thermal steam temperature stability, NOx emission concentration and boiler thermal efficiency cannot be guaranteed.

Method used

The triangular convolutional neural network TR-CNN with adaptive width was used to establish a predictive model of boiler thermal efficiency, NOx emissions and superheated steam temperature, and a boiler multi-objective combustion optimization agent based on the actor-critician network SAC was constructed, and the maximum entropy parameter was introduced to obtain the optimal strategy for all boiler combustion.

Benefits of technology

The parameters are streamlined through the TR-CNN model of adaptive width, reducing the inference time, and through the introduction of maximum entropy parameters, the problem of falling into local optimality is reduced, the optimal control of combustion in the furnace is achieved, the boiler thermal steam temperature stability is ensured, the NOx emission concentration is reduced, and the boiler thermal efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119983323B_ABST
    Figure CN119983323B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, device and medium for optimizing a boiler combustion strategy, relating to the technical field of combustion strategy optimization, including the steps of: establishing prediction models for boiler thermal efficiency, NOx emissions and superheated steam temperature by using a triangular convolutional neural network with an adaptive width; constructing a multi-objective combustion optimization agent for the boiler and introducing a maximum entropy parameter, aiming at obtaining all the optimal boiler combustion strategies, to obtain any strategy with a value higher than the threshold; based on the combustion strategy, outputting the next state through the trained prediction model and obtaining the reward corresponding to the current state, and iteratively optimizing the multi-objective combustion optimization agent for the boiler according to the current state, combustion strategy, next state and the reward corresponding to the current state until the combustion strategy when the reward is optimal is used as the optimal boiler combustion strategy. The present invention realizes the effects of reducing the emission concentration and ensuring the stability of the boiler thermal efficiency on the premise of ensuring the stability of the hot steam temperature.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of combustion strategy optimization, and particularly relates to a method, a system, a device and a medium for optimizing a boiler combustion strategy. Background Art

[0002] When the load changes rapidly, the combustion stability in the furnace becomes poor, which easily causes the deviation of the flame center, the deterioration of the quality of steam parameters, and further endangers the safe operation of the steam turbine. At the same time, the operation with a large range of load changes results in a decrease in the combustion efficiency in the furnace.

[0003] Constructing a prediction model of the boiler combustion state is the first step to achieve combustion optimization. In this field, data-driven modeling methods have shown obvious advantages. An artificial neural network (ANN) was trained based on more than 5,000 data points collected from published experimental results for the heat transfer prediction of supercritical water, which promoted the research on the thermal cycle of supercritical fluids. Li et al. selected all the operating condition data covering the superheated steam temperature in an Australian coal-fired unit to establish an ELMAN network model, and used this model as a feedback compensator to calculate the predicted values of the feedback gain and the feedforward gain. Deep learning incorporates the time dimension and the feature dimension into the input features and has the ability to infer the transient load of the boiler, which is the key technology for establishing a boiler dynamic model. A literature constructed a multi-modal hybrid mechanism and an LSTM modeling method based on a physical loss function to predict superheated steam, and verified the accuracy of the hybrid model under a multi-mode switching strategy based on an attention mechanism. However, the long short-term memory network LSTM and the gated recurrent unit GRU have the problem that the error in their memory modules accumulates over time, which affects the model accuracy and requires frequent model updates to meet the accuracy requirements. Therefore, designing an optimization algorithm is the second step to achieve combustion optimization. The existing technologies often use the genetic algorithm GA and the particle swarm optimization PSO for combustion optimization. The genetic algorithm GA encodes the combustion parameters into binary or real-number chromosomes and realizes the global search of the parameter space through selection, crossover, and mutation operations. The particle swarm optimization PSO represents the combustion parameter combination by the particle position and guides the search direction through the individual best (pBest) and the global best (gBest), so as to perform combustion optimization.

[0004] It can be seen that when the existing boiler multi-objective optimization method performs combustion optimization, it is easy to fall into local optimum under dynamic combustion conditions, cannot directly handle the optimization tasks in the continuous space, and needs to discretize the continuous actions. For a deterministic policy agent, the iterative effect is easily affected by the policy, and when dealing with high-dimensional non-linear combustion dampers, there are problems such as poor thermal steam temperature stability, inability to guarantee the NOx emission concentration and the boiler thermal efficiency. Summary of the Invention

[0005] The object of the present invention is to provide a method, system, device and medium for optimizing the boiler combustion strategy in view of the above-mentioned deficiencies of the prior art, so as to solve the problems in the prior art.

[0006] The present invention specifically provides the following technical solutions:

[0007] A method for optimizing the boiler combustion strategy includes the following steps:

[0008] Use a triangular convolutional neural network TR-CNN with an adaptive width to establish prediction models for boiler thermal efficiency, NOx emissions and superheated steam temperature, and train the prediction models.

[0009] Construct a multi-objective combustion optimization agent for the boiler based on the actor-critic network SAC, and introduce the maximum entropy parameter into the multi-objective combustion optimization agent for the boiler. With the goal of obtaining all the optimal boiler combustion strategies, based on the current state, obtain any combustion strategy with a value higher than the threshold through the multi-objective combustion optimization agent for the boiler with the maximum entropy parameter introduced; the combustion strategy includes boiler combustion parameters.

[0010] Based on the combustion strategy, output the next state through the trained prediction model, and obtain the reward corresponding to the current state, where the current state represents the parameters of the current boiler thermal efficiency, NOx emissions and superheated steam temperature, and the reward corresponding to the current state represents the stability of the superheated steam temperature.

[0011] Iteratively optimize the multi-objective combustion optimization agent for the boiler with the maximum entropy parameter introduced according to the current state, combustion strategy, next state and the reward corresponding to the current state, until the combustion strategy when the reward is optimal is obtained as the optimal boiler combustion strategy.

[0012] Preferably, the construction process of the triangular convolutional neural network TR-CNN with an adaptive width is specifically as follows:

[0013] For the width interval of the triangular convolutional neural network, randomly sample n - 2 widths to train the full-width network, and transfer the learned features to the sub-networks.

[0014] Use the soft labels obtained in the previous stage to train the sub-networks. Use the trained full-width network and the trained sub-networks as the trained networks, and construct a triangular convolutional neural network TR-CNN model with an adaptive width through the trained networks.

[0015] Preferably, training the prediction models includes:

[0016] Obtain auxiliary variables related to boiler thermal efficiency and NOx concentration.

[0017] Select variables in the auxiliary variables whose importance to the prediction target is higher than the threshold, establish input variables, reconstruct the input variables into a 2D tensor, input the reconstructed input variables into the prediction model, and determine the hyperparameters of the prediction model; where the input variables include a feature dimension and a time dimension.

[0018] Among them, when selecting variables in the auxiliary variables whose importance to the prediction target is higher than the threshold, the importance expression of each variable in the auxiliary variables is:

[0019] ;

[0020] ;

[0021] Among them, is the importance of each variable, GI is the Gini index, m represents the number of input features, K represents the number of types contained in a certain variable in the dataset, is the category k is the probability, GI m is its Gini index.

[0022] Preferably, with the goal of obtaining the optimal combustion strategy for all boilers, based on the current state, through a boiler multi-objective combustion optimization agent that introduces the maximum entropy parameter, any combustion strategy with a value higher than the threshold is obtained, including:

[0023] After introducing the maximum entropy parameter to the boiler multi-objective combustion optimization agent, for the state t at time s t and action a t , the maximization strategy of the boiler multi-objective combustion optimization agent is specifically expressed as:

[0024] ;

[0025] Among them, , represents the entropy of action a t , α represents the regularization coefficient of entropy, represents the strategy;

[0026] Introduce the discount factor γ , and obtain the optimization objective of the maximization strategy, specifically expressed as:

[0027] ;

[0028] Among them, represents the optimization objective,​ Represents the state distribution of the policy, T Represents the maximum moment, p is the initial state distribution, E is the expected value, R (·) and r (·) are both reward functions.

[0029] Preferably, the iterative optimization of the multi-objective combustion optimization agent of the boiler introducing the maximum entropy parameter according to the current state, combustion strategy, next state, and the reward corresponding to the current state includes:

[0030] Form a tuple of the state, action, next state, and reward in the replay buffer( s t , a t , s t+1 , r t );

[0031] Use the tuples collected by the current policy and the replay buffer to update the critic network Q, and evaluate the actions output by the policy network through the updated critic network; update the actor network according to the action value estimation of the updated critic network, and output a better combustion instruction;

[0032] Update the value network V through the loss function, and output the evaluation value of the current state through the updated value network V; where the output and loss function of the value network V are respectively:

[0033] ;

[0034] ;

[0035] Among them, is the output of the value network V, is the loss function, MSE is the mean square error function, and respectively represent two critic networks Q.

[0036] Preferably, the expression for updating the critic network using the tuples collected by the current policy and the replay buffer is:

[0037] ;

[0038] Among them, represents the updated target, represents the parameters of the two critic networks Q, represents the parameters of the two critic networks Q in the replay buffer, E is the expected value, is tMoment state, is t moment action, is t moment reward, is t+ The state at time 1, is the discount factor, α is the regularization coefficient of entropy. The state also includes unit load, total coal quantity, main steam flow rate, main steam pressure, feed water flow rate, oxygen content in flue gas, SCR inlet flue gas temperature, and wind box differential pressure.

[0039] The present invention provides a boiler combustion strategy optimization system, including:

[0040] A model construction module, which is used to establish prediction models for boiler thermal efficiency, NOx emissions, and superheated steam temperature by using a triangular convolutional neural network TR-CNN with an adaptive width, and train the prediction models;

[0041] A strategy acquisition module, which is used to construct a boiler multi-objective combustion optimization agent based on the actor-critic network SAC, introduce a maximum entropy parameter into the boiler multi-objective combustion optimization agent, and aim to obtain all the optimal boiler combustion strategies. Based on the current state, through the boiler multi-objective combustion optimization agent with the maximum entropy parameter introduced, obtain any combustion strategy with a value higher than the threshold; the combustion strategy includes boiler combustion parameters;

[0042] A state output module, which is used to output the next state based on the combustion strategy through the trained prediction model, and obtain the reward corresponding to the current state, where the current state represents the parameters of the current boiler thermal efficiency, NOx emissions, and superheated steam temperature, and the reward corresponding to the current state represents the stability of the superheated steam temperature;

[0043] A strategy optimization module, which is used to iteratively optimize the boiler multi-objective combustion optimization agent with the maximum entropy parameter introduced according to the current state, combustion strategy, next state, and the reward corresponding to the current state, until the combustion strategy when the reward is optimal is obtained as the optimal boiler combustion strategy.

[0044] The present invention provides a computer device, including a memory and a processor. When a program stored in the memory is executed by the processor, the processor is caused to execute the steps of the above-mentioned boiler combustion strategy optimization method.

[0045] The present invention provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned boiler combustion strategy optimization method are implemented.

[0046] Compared with the prior art, the present invention has the following remarkable advantages:

[0047] The present invention constructs a prediction model using a triangular convolutional neural network with an adaptive width. On the premise of ensuring the prediction accuracy, the model parameters can be streamlined through the adaptive width, effectively reducing the inference time. Moreover, the maximum entropy parameter is introduced into the multi-objective combustion optimization agent of the boiler to randomize the strategy, that is, the probability of each output optimization action is as dispersed as possible, reducing the problem of falling into local optima. The agent can achieve the optimal control of in-furnace combustion through countless paths without missing any strategy. Based on the combustion strategy, the next state is output through the trained prediction model, and the reward corresponding to the current state is obtained. Taking the stability of the current superheated steam temperature of the boiler as the reward function, the multi-objective combustion optimization agent of the boiler with the maximum entropy parameter is iteratively optimized according to the current state, combustion strategy, next state, and the reward corresponding to the current state, effectively preventing the steam parameters from deteriorating due to the control strategy and affecting the performance of the unit in participating in deep regulation. On the premise of ensuring the stability of the hot steam temperature, the optimal boiler combustion strategy is used to achieve the effects of reducing the NOx emission concentration and ensuring the stability of the boiler thermal efficiency. Description of the Drawings

[0048] Figure 1 It is a 3D structure and combustion system diagram of the boiler of the present invention;

[0049] Figure 2 It is a TR-CNN network diagram;

[0050] Figure 3 It is an all-condition analysis diagram of the accuracy comparison of sampling data for different models; among which Figure 3 (a) is a diagram of the relationship between load and superheated steam temperature, Figure 3 (b) is a diagram of the relationship between NOx emissions and thermal efficiency;

[0051] Figure 4 It is a diagram of the feature importance evaluated by the Gini index;

[0052] Figure 5 It is a schematic diagram of data reconstruction;

[0053] Figure 6 It is a framework diagram of the boiler combustion decision optimization based on the stochastic policy SAC agent;

[0054] Figure 7 It is Figure 6 the network structure diagrams in Figure 7 (a) is the structure diagrams of the Stochastic Actor network, Vcritic network, and Target V critic network, Figure 7 (b) is the structure diagrams of the Critic Q1-network and Q2-network;

[0055] Figure 8Loss curves of two CNN models on the training set and the test set; among them, Figure 8 (a) of Figure 8 is the training loss,

[0056] Figure 9 Prediction result graph of 3-layer CNN on the test set; among them, Figure 9 (a) of Figure 9 is the thermal efficiency graph of 3-layer CNN on the test set, Figure 9 (b) of

[0057] Figure 10 is the NOx emission graph of 3-layer CNN on the test set, Figure 10 (c) of Figure 10 is the superheated steam temperature of 3-layer CNN on the test set; Figure 10 (a) of

[0058] Figure 11 Iterative curve graph of the average reward of two agents;

[0059] Figure 12 Optimization curve 1 graph of three target variables; among them Figure 12 (a) of Figure 1 is the thermal efficiency Figure 12 (b) of Figure 12 is the error graph when obtaining the thermal efficiency, Figure 1 (c) of Figure 12 is the NOx emission Figure 12 (d) of Figure 1 is the error graph when obtaining the NOx emission, Figure 12 (e) of

[0060] Figure 13 is the superheated steam temperature Figure 13 (f) of Figure 2 is the error graph when obtaining the superheated steam temperature; Figure 13 (b) of Figure 13 is the error graph when obtaining the thermal efficiency, Figure 2 (c) of Figure 13 is the NOx emission Figure 13 (d) of Figure 2 is the error graph when obtaining the NOx emission, Figure 13 (f) of

[0061] Figure 14 Flow chart of a method for optimizing the boiler combustion strategy provided by the present invention. Specific embodiments

[0062] The following combines the accompanying drawings in the present invention to clearly and completely describe the technical solutions of the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0063] The object of the present invention is a supercritical coal-fired boiler with a low nitrogen oxide concentric ring system (LNCFS). The target boiler is a subcritical parameter, with once-through intermediate reheat, balanced draft natural circulation drum boiler, using a positive pressure direct-fired pulverized coal system, and direct-flow pulverized coal burners arranged at the four corners.

[0064] As Figure 1 Shown is the 3D structure and combustion system diagram of the boiler. The boiler adopts a single furnace and tangential combustion at the four corners. It is equipped with 5 coal mills, 4 of which are in operation and 1 is in standby, respectively providing fuel to the primary air nozzles of layers A - E. The secondary air accounts for the largest proportion of the air volume required for combustion in the furnace, and 7 secondary air dampers are arranged near the primary air dampers.

[0065] When the boiler burns the designed coal type, the minimum stable combustion load without oil injection is 38%B - MCR, and the oil burners correspond to the OA / OB / OC nozzles. To effectively reduce the NOx concentration in the burnout zone, the boiler is equipped with three layers of compact overfire air.

[0066] The optimization goal of the present invention is to improve the boiler efficiency and reduce the NOx emissions on the premise of ensuring the stability of the superheater temperature. Based on the influence of various heat losses of the boiler on the efficiency and analyzing the NOx generation mechanism in the furnace, the combustion states and control variables with high correlations with the boiler thermal efficiency and NOx concentration are selected respectively, and after merging them, a list containing 32 auxiliary variables is constructed (see Table 1).

[0067] Although CNN performs well in boiler combustion modeling, in the deep reinforcement learning combustion optimization framework, the low latency and real-time feedback characteristics of the simulation environment become very crucial. An efficient decision-making mechanism not only requires a sensitive agent, but also a lightweight boiler simulation environment can reduce the interaction cycle between the agent and the environment.

[0068] Table 1 32 auxiliary variables

[0069]

[0070] Specifically illustrate this embodiment with reference to the accompanying drawings, asFigure 14 As shown in Figure 14 , a method for optimizing the boiler combustion strategy in this embodiment includes the following steps:

[0071] Step S1: Use a triangular convolutional neural network with adaptive width, TR-CNN, to establish prediction models for boiler thermal efficiency, NOx emissions, and superheated steam temperature, and train the prediction models.

[0072] The specific expression of the feature extraction process of the convolutional layer is:

[0073] (1);

[0074] Among them, y is the result of the convolution operation, is the weight of the convolution kernel, is the input feature vector, and , b is the bias. The maximum number of nodes in each layer of the convolutional layer is the width n Here, the width n is the number of channels. For a convolutional neural network, the more channels, the higher the model accuracy. However, a large number of hyperparameters will affect the inference speed. Therefore, n plays a key role in the trade-off between model accuracy and efficiency. A narrower network structure has worse performance than a wider network. The residual between channels is as follows:

[0075] (2);

[0076] Among them, k 0 represents the minimum width of the network. The sum of the first k channels is represented by . There is an upper bound on the residual between the features extracted by the full-width network and the partial-width network. The network width ranges from k 0 to k n In the process, there are multiple network training processes. The number of networks depends on the sparsity of the width. At the same time, in this invention, the upper and lower limits of the width are set to 0.125 and 1, and n -2 widths are randomly sampled within this interval for network training.

[0077] The construction process of the triangular convolutional neural network with adaptive width, TR-CNN, is specifically as follows:

[0078] For the width range of the triangular convolutional neural network, randomly sample n - 2 networks with full width for width training, transfer the learned features to the sub-networks; use the soft labels obtained in the previous stage to train the sub-networks, take the trained full-width network and the trained sub-networks as the trained networks, and construct an adaptive-width triangular convolutional neural network TR-CNN model through the trained networks.

[0079] Specifically, for the width range [0.2, 1], when the width coefficient is 1, it is a full-width network, and when the width coefficient is 0.2, it means that the number of neurons in the nodes of the sub-network is 20% of that of the full-width network. First, train the full-width network, transfer the learned features to the sub-network, and use the soft labels of the previous stage to train the sub-network to obtain the trained network, and construct a prediction model through the trained network. This training strategy naturally obtains knowledge refinement, and the experimental results confirm that the prediction accuracy does not decrease rapidly as the network width decreases.

[0080] During the network switching process, it is necessary to consider the problem of the influence of setting different widths on the weight parameters. Taking k i = 1, 2, 3 as an example, and ignoring the bias b, the specific expression of formula (1) is:

[0081] (3);

[0082] (4);

[0083] (5);

[0084] Further simplified, y 1 can be converted into formula (6):

[0085] (6);

[0086] The network widths are different, y and the expected value of 1 is also different, as shown in formula (7):

[0087] (7);

[0088] Due to the existence of the BN layer, when switching between multiple widths, y the statistics of 1 are different, which further reduces the prediction accuracy. In order to ensure that the expected values in other cases are close, or is essential. y The expected value of 2 is as follows:

[0089] (8);

[0090] When When the expected value and are consistent, the weights on the filter can be transformed as shown in Equation (9):

[0091] (9);

[0092] After the above feature extraction, a fully connected layer and an output layer follow the last convolutional layer, and the weights in the model are obtained by taking the derivative of the loss function.

[0093] (10);

[0094] When training the prediction model, it includes:[[]]

[0095] Obtaining auxiliary variables related to the boiler thermal efficiency and NOx concentration.

[0096] Selecting variables in the auxiliary variables whose importance to the prediction target is higher than the threshold, establishing input variables, reconstructing the input variables into a 2D tensor, inputting the reconstructed input variables into the prediction model, and determining the hyperparameters of the prediction model; where the input variables include a feature dimension and a time dimension.

[0097] Adopting a feature importance threshold screening mechanism to screen multiple variables from the auxiliary variables as input variables, realizing multi-scale feature interaction between variables through two-dimensional tensor reconstruction, determining the hyperparameters of the prediction model, enhancing the sensitivity of the model to boiler combustion, and reducing the influence of exploration noise and hyperparameters.

[0098] 20-day historical data of the variables in Table 1 were extracted from the DCS of the target boiler, with a sampling period of 1 minute, totaling approximately 28,800 sampling points. Through further analysis of the data, operating intervals with a large amount of missing data were excluded, and approximately 10,100 sampling points of continuous operating historical data that met the requirements were selected. The operating intervals of the variables during this period are also listed in Table 1. Abnormal data were processed using the 3σ-rule. As Figure 3 shown are the curves of load, superheated steam temperature, NOx emissions, and thermal efficiency in the 7-day historical data. The sampled data cover various operating conditions such as steady state and variable load, which are sufficient to support the comparison and analysis of various operating conditions during the experiment.

[0099] In addition, to avoid large errors in the predicted values of the model output layer and improve the training efficiency of the model, the original data was processed by min-max scaling normalization.

[0100] The input features determine the upper limit of the model's prediction performance. Selecting variables with high importance to the prediction target from the auxiliary variables and establishing high-quality input features can not only improve the model training efficiency but also avoid model overfitting. As a supervised dimensionality reduction method, RF is particularly suitable for evaluating the importance among high-dimensional non-linear industrial variables such as boilers. In RF, the average decrease in the Gini index is considered the most important characterization of Gini impurity. If the Gini index of a certain feature is larger, it means that the feature has a more significant impact on the target variable. The importance of variables is represented by VIM, and the Gini index is represented by GI, as shown in Eqs. (11) and (12). Specifically:

[0101] When selecting variables from the auxiliary variables whose importance to the prediction target is higher than the threshold, the importance expression of each variable in the auxiliary variables is:

[0102] (11);

[0103] (12);

[0104] Among them, VIM is the importance level (importance) of each variable, GI is the Gini index, m represents the number of input features, K represents the number of types contained in a certain variable in the dataset, is the category k probability, GI m is its Gini index. The auxiliary variables with importance levels higher than the threshold are used as input variables.

[0105] Assuming the number of decision trees is n, the importance of the j-th feature in the i-th decision tree is shown in Eq. (13), and the importance on all decision trees is shown in Eq. (14). Eq. (15) normalizes the features:

[0106] (13);

[0107] (14);

[0108] (15);

[0109] Among them, the sum of the importance of all features is 1, that is .

[0110] This invention's research respectively evaluated the importance of input variables for boiler thermal efficiency and NOx emissions based on the Gini index. Taking the sum of importance reaching 0.9 as the boundary, 13 and 18 key variables were respectively selected (see Figure 4). Among them, Fuelair is fuel air and SOFA is overfire air. From the evaluation results, for thermal efficiency, the importance of unit load ranks first. During actual operation, the higher the unit load, the higher the boiler thermal efficiency. Under deep peak shaving conditions, at low loads, the boiler thermal efficiency will decrease significantly. For NOx emissions, the variable importance of SOFA-C is the highest, which is consistent with the fact that overfire air can effectively reduce NOx emissions.

[0111] Generally speaking, fuel air and mill primary air have a greater impact on thermal efficiency, while secondary air and overfire air have a greater impact on NOx emissions.

[0112] Table 2 Test variables and target variables

[0113]

[0114] After merging the two groups of key variables and removing duplicate variables, they are used as the input variables of the multi-objective prediction model, as shown in Table 2. Among them, the test variables represent the input boiler combustion parameters, and the feature dimension includes at least 22 test variables. In this embodiment, the target variables refer to the thermal efficiency, NOx emissions, and superheated steam temperature at the first 20 moments. The time dimension is set to 20, and the moving step is set to 1. Then the prediction target is the boiler thermal efficiency, NOx emissions, and superheated steam temperature at the 21st moment. The schematic diagram of data reconstruction is as Figure 5 shown. At the same time, for the convenience of data management, the serial numbers (21 - 10100) of the target variables are rearranged to be consistent with the window serial numbers (1 - 10080). The entire dataset is divided into a training set and a test set in a ratio of 80% and 20%, and the hyperparameters of the prediction model are determined through the training set. Three model performance evaluation criteria are adopted: coefficient of determination R 2 , root mean square error RMSE, and mean absolute error MAE, and their definitions are as follows:

[0115] (16);

[0116] (17);

[0117] (18);

[0118] Among them, , , respectively represent the true value, the predicted value, and the mean value of the true value, represents the total number of samples.

[0119] Table 3 Hyperparameter table

[0120]

[0121] Step S2: Construct a boiler multi-objective combustion optimization agent based on the Actor-Critic network SAC, and introduce the maximum entropy parameter into the boiler multi-objective combustion optimization agent. Aiming to obtain all the optimal boiler combustion strategies, based on the current state, through the boiler multi-objective combustion optimization agent with the maximum entropy parameter introduced, obtain any combustion strategy with a value higher than the threshold; the combustion strategy includes boiler combustion parameters.

[0122] The variables controlling combustion in the boiler have characteristics such as high dimension and non-linearity, and the controlled valves are mainly continuous variables. Therefore, an agent suitable for the continuous action space needs to be considered. Agents that meet the conditions include DDPG, TD3, and SAC. However, DDPG and TD3 aim to maximize the expected reward. In a boiler combustion state, only one optimal action is considered, and there is a possibility of missing other optimal strategies. SAC sets an entropy term in the optimization function. By seeking a strategy that can both maximize the expected reward and maximize the entropy, entropy is a measure of randomness in the strategy. The core idea of maximum entropy is not to miss any valuable strategy, effectively solving the problem of deterministic strategies and also improving the robustness of the strategy. For t the state s t and action a t , the maximization strategy of SAC is shown in Equation (19).

[0123] Aiming to obtain all the optimal boiler combustion strategies, based on the current state, through the boiler multi-objective combustion optimization agent with the maximum entropy parameter introduced, obtain any combustion strategy with a value higher than the threshold, including:

[0124] (19);

[0125] Among them, , represents the action a t entropy, α represents the regularization coefficient of entropy, used to control the randomness of the optimal strategy, represents the strategy. In the SAC algorithm, in order to ensure the convergence of the reward function, the discount factor γ is introduced, and the optimization objective for obtaining the maximization strategy is thus defined as Equation (20):

[0126] (20);

[0127] Among them, represents the optimization objective, represents the state distribution of the strategy, T represents the maximum time, p is the initial state distribution, E is the expected value,R (·) and r (·) are both reward functions.

[0128] This goal is equivalent to maximizing the discounted expected return and entropy of the future state of the state-action pair ( s t , a t ). SoftActor-Critic uses a stochastic policy, which has certain advantages compared to a deterministic policy. For example, it increases the randomness of the policy, can prevent the policy from converging to a local optimum prematurely, and thus encourages the agent to conduct a wider exploration to achieve the purpose of approaching the optimal solution. If multiple actions for controlling boiler combustion are equally important, the agent will set the same probability weights for these actions in the policy distribution to ensure the stability of the combustion policy.

[0129] Step S3: Based on the combustion policy, the trained prediction model outputs the next state and obtains the reward corresponding to the current state, where the current state represents the parameters of the current boiler thermal efficiency, NOx emissions, and superheated steam temperature, and the reward corresponding to the current state represents the stability of the superheated steam temperature.

[0130] Step S4: Iteratively optimize the multi-objective combustion optimization agent for the boiler with the maximum entropy parameter according to the current state, combustion policy, next state, and the reward corresponding to the current state until the combustion policy when the reward is optimal is obtained as the optimal boiler combustion policy.

[0131] SAC is an off-policy reinforcement learning algorithm, and its structure consists of three types of functions and five networks, including a policy function, a value function, and a state value function. A random action network is used to approximate the policy function, a v critic network is used to approximate the state value function, and a Q network is used to approximate the action value function. Its algorithm steps include policy evaluation, policy improvement, state value network update, and entropy weight adjustment. The specific details are as follows:

[0132] Iteratively optimize the multi-objective combustion optimization agent for the boiler with the maximum entropy parameter according to the current state, combustion policy, next state, and the reward corresponding to the current state, including:

[0133] Form a tuple of the state, action, next state, and reward in the replay buffer ( s t , a t , s t+1 , r t ).

[0134] Update the critic network Q with the tuples collected using the current policy and the experience replay buffer. Evaluate the actions output by the policy network using the updated critic network. Update the actor network based on the action value estimates of the updated critic network to output a better combustion instruction, that is, output a combustion instruction for the current state of the boiler through the policy evaluation network Q, including:

[0135] Policy evaluation: Update the critic network, which is the policy evaluation network Q, using the mini-batch data collected by the current policy and the experience replay buffer. The update objective is to minimize the soft Bellman residual, specifically:

[0136] (21);

[0137] Where, represents the update objective, represents the parameters of the two critic networks Q, represents the parameters of the two critic networks Q under the replay buffer. The structure of the double Q network is adopted to prevent overestimation of the Q value, and then the smaller one of the two Q values is selected. E is the expected value, is t the state at time is t the action at time is t the reward at time is t+ the state at time t+1, is the discount factor, α is the regularization coefficient of entropy, D represents the tuples in the replay buffer; the state represents the parameters of the boiler combustion, specifically including unit load, total coal consumption, main steam flow rate, main steam pressure, feed water flow rate, flue gas oxygen content, SCR inlet flue gas temperature, wind box pressure difference, boiler thermal efficiency, NOx emissions, and superheated steam temperature.

[0138] Policy update: Update the actor network based on the action value estimates of the critic network to output a better combustion instruction; for the update of the Policy network parameters, it is to minimize the KL divergence, as shown in Equation (22). By weakening the contribution of the partition function to the gradient, the loss function is simplified to Equation (23):

[0139] (22);

[0140] (23);

[0141] To convert random sampling into a differentiable deterministic operation, when the policy evaluation network outputs an action, a reparameterization technique is adopted. The action sampling with a Gaussian distribution is shown as in (24):

[0142] (24);

[0143] where, represents the mean of the policy distribution, represents the variance of the policy distribution, represents the added noise.

[0144] The value network V is updated through the loss function to make the estimated state value more accurate, and the evaluation value of the current state is output through the updated value network V. The output and loss function of the value network V are respectively:

[0145] (25);

[0146] (26);

[0147] where, is the output of the value network V, is the loss function, MSE is the mean square error function, and respectively represent two critic networks Q.

[0148] The current state value is evaluated through the updated value network V, and the entropy weight is adjusted based on the entropy of the current policy. Entropy weight adjustment: A key innovation in the SAC algorithm is to automatically adjust the entropy weight α , to adapt to different tasks. This adaptive adjustment mechanism ensures that while maintaining sufficient exploration, an effective policy can also be effectively learned. The entropy weight is adjusted based on the entropy of the current policy as shown in equation (27) α .

[0149] (27);

[0150] where, D represents the tuple of the replay buffer, represents the policy under the condition .

[0151] On the premise of safe operation, improving the boiler thermal efficiency and reducing pollutant emissions are the research focuses of the present invention, Figure 6 and Figure 7Shown is a boiler combustion decision-making optimization framework based on a stochastic policy SAC agent. Among them, TR-CNN serves as the boiler simulation environment and interacts with the SAC agent. The agent adopts corresponding combustion strategies according to the current state of the boiler, and the boiler simulation system TR-CNN outputs the next state s t+1 , and feedback rewards r t . The tuples that make up the replay buffer ( s t , a t , s t+1 , r t ) are used for sampling by experience replay. According to the requirements of boiler combustion strategy optimization, this invention designs the policy network and value networks (Q and V).

[0152] The policy network is responsible for outputting combustion instructions according to the current state of the boiler. Based on the analysis of factors affecting in-furnace combustion, 11 variables including unit load, total coal quantity, main steam flow rate, main steam pressure, feed water flow rate, flue gas oxygen content, SCR inlet flue gas temperature, air box differential pressure, boiler thermal efficiency, NOx emissions, and superheated steam temperature are selected as the characteristics representing the boiler combustion state. To optimize the coal blending method and air distribution method, three fuel air opening degrees, x 11 - x 13 and x 14 and x 17 four secondary air damper opening degrees, two OFA damper opening degrees of SOFA-A and SOFA-B, x 21 - x 22 and the primary air flow rates at the inlets of two coal mills, a total of 12 controlled variables, are selected as the control quantities for combustion. The policy network is designed as a four-layer fully connected neural network. The input layer of the network consists of 11 neurons, corresponding to the 11 variables representing the boiler combustion state. The output layer consists of 24 neurons, corresponding to the means and covariances of the 12 controlled quantities. Each of the two hidden layers is set with 512 neurons.

[0153] The Q-network is responsible for evaluating the actions output by the policy evaluation network. Its input is the state-action pair, and the output is the score - Q value of the current boiler state. In this invention, the input of the Q-network includes 11 states and 12 actions, a total of 23 variables. The Q-network is designed as a four-layer fully connected neural network, and each of the two hidden layers is set with 512 neurons.

[0154] The V-network has the same input features as the policy evaluation network Q, and its output is the evaluation of the current state. The evaluation value is the expected cumulative reward. The network structure of the V-network is only different from that of the policy evaluation network in the output layer, and the other structures are the same.

[0155] The boiler simulation environment and the SAC agent were tested and analyzed. First, the prediction performance of the boiler simulation environment based on TR-CNN was evaluated, and the optimal width coefficient W value was explored. Then, to compare the decision-making ability of the proposed agent, the combustion optimization results were analyzed through the interaction between TR-CNN and SAC.

[0156] To explore the optimal performance of the multi-objective prediction model, this section conducted comparative experiments on the test set by setting different width coefficient W values of TR-CNN (W = 0.125, 0.25, 0.5, 0.75, 1.0). Since there is a trade-off between accuracy and inference speed, the inference time of 2,016 test samples was also counted. Among them, the hyperparameters of TR-CNN followed the determined values. From the experimental results, as the width coefficient decreased, the amount of model parameters used decreased, and the calculation time-consuming decreased significantly. However, the model prediction error gradually increased. When the width coefficient was reduced from W = 0.25 to W = 0.125, although the inference time decreased slightly, the prediction error decreased significantly. Compared with larger width coefficients, W = 0.25 significantly reduced the calculation time-consuming, and the inference time was 11.894 s, which was 28.92% lower than the inference time when the width coefficient was 1. Considering both the inference time and the prediction accuracy, the subsequent comparative experiment of TR-CNN in the present invention selected the width coefficient of 0.25.

[0157] Table 4 RMSE of TR-CNN predicting each target variable under different width coefficients

[0158]

[0159] To verify the prediction accuracy of the multi-objective model, a 3-layer CNN was selected as the baseline model, and a comparative experiment was conducted with TR-CNN. The loss curves of the two CNN models on the training set and the test set are as Figure 8 shown. From the changing trend, the loss value on the training set is lower than that on the test set. Whether on the training set or the test set, the loss value of TR-CNN is lower than that of the baseline model 3-layer CNN, indicating that the convergence effect of TR-CNN is better than that of 3-layer CNN. Correspondingly, the prediction results of 3-layer CNN and TR-CNN on the test set are as Figure 9 and Figure 10 shown. From the prediction results of the thermal efficiency, Figure 9The test results in Figure (a) are scattered on both sides of the perfect line, with a large error between the true value and the predicted value. However, Figure 10 The test results in Figure (a) are concentrated near the perfect line, and the true value is close to the predicted value. The same prediction results also apply to NOx emissions and superheated steam temperature.

[0160] Combining the prediction indicators to analyze the prediction performance of the baseline model and TR-CNN. Taking the superheated steam temperature as an example, the RMSE of the baseline model is 6.021 °C, the MAE is 4.741 °C, and the R 2 is 0.912. The RMSE of TR-CNN is 4.922 °C, the MAE is 3.495 °C, and the R 2 is 0.942. Among them, RMSE and MAE decreased by 18% and 26% respectively, and the R 2 increased by 3%. It can be seen that the model proposed in the present invention significantly improves the prediction performance compared with the baseline model, and the prediction accuracy meets the requirements of interacting with the agent.

[0161] To verify the performance of the SAC agent in combustion decision-making, it was compared with the traditional reinforcement learning agent - DDPG on the test set. By interacting the two agents with the simulation environment, the optimization effects on boiler thermal efficiency, NOx emissions and superheated steam temperature were analyzed. The hyperparameters related to the agents are shown in Table 5. The agents iteratively learn through the tested schemes during training, that is, as the training progresses, the policy is verified every 200 episodes, and the total step length of the test experiment is 6×10^4.

[0162] Table 5 Hyperparameters of the two agents

[0163]

[0164] By means of the iterative curves of the reward means of the two agents (see Figure 11 ), the learning process of the agents was analyzed. From the entire learning process, the reward means of the two agents both converged in the end, and the reward obtained by SAC was higher than that of DDPG, indicating that the combustion decision-making level mastered by SAC is higher than that of DDPG. In terms of details, it is worth noting that at the initial stage of the interaction between the agent and the simulation environment, DDPG obtained a higher reward than SAC, indicating that DDPG has a faster learning speed in the initial stage. At the end of the interaction, the reward obtained by SAC was significantly higher than that of DDPG. The highest reward of the SAC agent appeared at the 5.32×10^4 step, and the reward value was 475.206.

[0165] The optimization results of the two agents on thermal efficiency and NOx emissions were statistically analyzed. The optimization curves of the three target variables are shown in Figure 12 and Figure 13As shown. For the DDPG agent, before and after optimizing the thermal efficiency, the optimization interval is [0.069, 0.303], that is, the thermal efficiency of all samples has been improved. Before and after optimizing the NOx emissions, the obtained optimization interval is [-50.239, 17.201], that is, the NOx concentration has been reduced by up to 50.239 mg / m3, and there are also sample points where the NOx concentration increases, with a maximum increase of 17.201 mg / m3. According to Figure 12 The statistical data in the (d) graph of, among the 2016 test samples, 202 samples failed to optimize the NOx emissions, accounting for 10.02%. For the SAC agent, before and after optimizing the thermal efficiency, the optimization interval is [0.229, 0.462], and again, the thermal efficiency of all samples has been improved. Before and after optimizing the NOx emissions, the obtained optimization interval is [-59.739, 4.061], that is, the NOx concentration has been reduced by up to 59.739 mg / m3, and there are also sample points where the NOx concentration increases, with a maximum increase of 4.061 mg / m3. According to Figure 13 The statistical data in the (d) graph of, 13 samples failed to optimize the NOx emissions, accounting for 0.64%. According to Figure 12 The (e) graph of Figure 12 The (f) graph of and Figure 13 The (e) graph of Figure 13 From the superheated steam temperature data statistically analyzed in the (f) graph of, both agents can ensure the stability of the superheated steam temperature during combustion optimization. From the optimization mean values of the performance variables, DDPG improves the thermal efficiency by 0.196% and reduces the NOx emissions by 10.549 mg / m 3 . SAC improves the thermal efficiency by 0.357% and reduces the NOx emissions by 20.244 mg / m 3 . From the mean value of the thermal efficiency, SAC improves by 0.161% more than DDPG. From the mean value of the NOx emissions, SAC reduces by 9.695 mg / m3 more than DDPG. The optimization effect of SAC is better.

[0166] To analyze the physical factors behind the boiler combustion decision-making of the agent, the optimization trends of various performance variables are analyzed in combination with the load curve. From Figure 12 The (a) graph of and Figure 13 The (a) graph of, it can be seen that among the 1000-1100th test points, the boiler is in a near-full load state, and the in-furnace combustion efficiency is also very high. At this time, the boiler thermal efficiency is also at a relatively high level. However, the space for optimizing the thermal efficiency becomes relatively small. From Figure 12 The (b) graph of and Figure 13As can be clearly seen from Figure (b), in this interval, the amplitude of the thermal efficiency improvement is generally very small, and the lowest value also appears at the 1094th test point. Near the 1500-1800th test points, the boiler is at a load close to 30% of the rated load. The lower the load, the more unstable the combustion in the furnace, which is likely to cause the deviation of the flame center. On the one hand, the gas flow scours the heating surface, and the flue gas temperature deviation on both sides of the furnace outlet becomes larger, which will lead to consequences such as the wear and even bursting of the water wall tubes and the local overheating of the convective heating surface. The local high temperature of the water wall and other convective heating surfaces causes fluctuations in the steam side parameters, seriously affecting the safety of the boiler. From Figure 12 Figure (e) of Figure 13 and Figure (e) of Figure 12 it can be seen that there are large fluctuations in the superheated steam temperature in this interval. At the same time, from Figure 13 Figure (f) of Figure 12 and Figure (f) of Figure 13 it can be seen that due to the large fluctuations in the superheated steam temperature, there is also a trend of large deviation when the two agents stabilize the superheated steam temperature in this interval. On the other hand, at low loads, the reduction environment of NOx in the burnout zone is out of balance, and the NOx emissions at the SCR inlet increase significantly and fluctuate violently in the 1500-1800th test interval (see Figure 12 Figure (b) of Figure 12 and Figure (d) of Figure 13 and Figure (b) of Figure 13 and Figure (d) of

[0167] Based on the above method, the present invention provides a boiler combustion strategy optimization system, including: a model construction module, a strategy acquisition module, a state output module, and a strategy optimization module.

[0168] Among them, the model construction module is used to establish a prediction model for boiler thermal efficiency, NOx emissions, and superheated steam temperature by using a triangular convolutional neural network TR-CNN with an adaptive width, and train the prediction model; the policy acquisition module is used to construct a boiler multi-objective combustion optimization agent based on the actor-critic network SAC, and introduce a maximum entropy parameter into the boiler multi-objective combustion optimization agent, aiming to obtain all the optimal boiler combustion strategies. Based on the current state, through the boiler multi-objective combustion optimization agent with the maximum entropy parameter introduced, any combustion strategy with a value higher than the threshold is obtained; the combustion strategy includes boiler combustion parameters; the state output module is used to output the next state based on the combustion strategy through the trained prediction model, and obtain the reward corresponding to the current state, where the current state represents the parameters of the current boiler thermal efficiency, NOx emissions, and superheated steam temperature, and the reward corresponding to the current state represents the stability of the superheated steam temperature; the policy optimization module is used to iteratively optimize the boiler multi-objective combustion optimization agent with the maximum entropy parameter introduced according to the current state, combustion strategy, next state, and the reward corresponding to the current state until the combustion strategy when the reward is optimal is obtained as the optimal boiler combustion strategy.

[0169] The present invention also provides a computer device, including a memory and a processor. When a program stored in the memory is executed by the processor, the processor executes the steps of a method for optimizing a boiler combustion strategy.

[0170] According to the disclosed embodiments, the computer device can communicate with one or more external devices (such as a keyboard, a pointing device, Bluetooth communication, etc.), or communicate with any device (such as a router, a demodulator, etc.) that enables the computing device to communicate with one or more other computing devices.

[0171] The present invention also provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a method for optimizing a boiler combustion strategy are implemented.

[0172] According to the disclosed embodiments, the storage medium can be a non-volatile computer-readable storage medium, for example, it can include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, apparatus, or device.

[0173] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, which should all be regarded as belonging to the protection scope of the present invention.

Claims

1. A boiler combustion strategy optimization method, characterized in that: The steps include: A prediction model of boiler thermal efficiency, NOx emission and superheated steam temperature is established by using a triangular convolutional neural network TR-CNN with adaptive width, and the prediction model is trained; A boiler multi-objective combustion optimization intelligent agent based on the actor-critic network SAC is constructed, and a maximum entropy parameter is introduced into the boiler multi-objective combustion optimization intelligent agent, with the goal of obtaining the optimal combustion strategy of all boilers. Based on the current state, any combustion strategy with a value higher than a threshold is obtained by introducing the boiler multi-objective combustion optimization intelligent agent with the maximum entropy parameter; the combustion strategy includes boiler combustion parameters; Based on the combustion strategy, the next state is output through the trained prediction model, and the reward corresponding to the current state is obtained, where the current state represents the current boiler thermal efficiency, NOx emissions and superheated steam temperature parameters, and the reward corresponding to the current state represents the stability of the superheated steam temperature; According to the current state, combustion strategy, next state and the reward corresponding to the current state, the boiler multi-objective combustion optimization agent with the maximum entropy parameter is iteratively optimized until the combustion strategy with the optimal reward is obtained as the optimal boiler combustion strategy.

2. A boiler combustion strategy optimization method according to claim 1, characterized in that: The construction process of the adaptive width triangular convolutional neural network TR-CNN is specifically as follows: For the width interval of the triangular convolutional neural network, randomly sample n-2 widths to train the full-width network and transfer the learned features to the sub-network; The soft labels obtained in the previous stage are used to train the sub-network, and the trained full-width network and the trained sub-network are used as the trained network. The trained network is used to construct a triangular convolutional neural network TR-CNN model with adaptive width.

3. A boiler combustion strategy optimization method according to claim 1, characterized in that: Training the prediction model includes: Obtain auxiliary variables related to boiler thermal efficiency and NOx concentration; Select variables whose importance to the prediction target is higher than a threshold value from the auxiliary variables, establish input variables, reconstruct the input variables into a 2D tensor, input the reconstructed input variables into the prediction model, and determine the hyperparameters of the prediction model; the input variables include feature dimension and time dimension; Among them, when selecting variables whose importance to the prediction target is higher than the threshold in the auxiliary variables, the importance expression of each variable in the auxiliary variables is: ; ; in, is the importance of each variable, GI is the Gini index, m represents the number of input features, K Indicates the types of variables in the data set. For Category k The probability of GI m Its Gini index.

4. A boiler combustion strategy optimization method according to claim 1, characterized in that: The objective of obtaining the optimal combustion strategy for all boilers is to obtain any combustion strategy with a value higher than a threshold by introducing a boiler multi-objective combustion optimization agent with a maximum entropy parameter based on the current state, including: After introducing the maximum entropy parameter into the boiler multi-objective combustion optimization agent, t Status at the moment s t and actions a t , the specific expression of the maximization strategy of the boiler multi-objective combustion optimization agent is: ; in, , indicating an action a t The entropy of α represents the regularization coefficient of entropy, Representation strategy; Introducing a discount factor γ , get the optimization target of the maximization strategy, the specific expression is: ; in, represents the optimization goal, represents the state distribution of the strategy, T represents the maximum moment, p is the initial state distribution, E is the expected value, R (·)and r (·) are all reward functions.

5. A boiler combustion strategy optimization method according to claim 4, characterized in that: The iterative optimization of the boiler multi-objective combustion optimization agent with the maximum entropy parameter introduced according to the current state, the combustion strategy, the next state and the reward corresponding to the current state includes: The state, action, next state and reward are combined into a tuple of replay buffers ( s t , a t , s t+1 , r t ); Update the critic network Q using the current policy and the tuples collected in the replay buffer, and evaluate the actions output by the policy network through the updated critic network; Update the actor network based on the action value estimate of the updated critic network to output a better burning instruction; The value network V is updated through the loss function, and the evaluation value of the current state is output through the updated value network V; the output and loss function of the value network V are: ; ; in, is the output of the value network V, is the loss function, MSE is the mean square error function, and Represent the two critic networks Q respectively.

6. A boiler combustion strategy optimization method according to claim 5, characterized in that: The critic network is updated using the current strategy and the tuples collected in the replay buffer. The specific expression is: ; in, represents the updated target, denotes the parameters of the two critic networks Q, represents the parameters of the two critic networks Q under the playback buffer, E is the expected value, for t Moment status, for t Always in action, for t Rewards at all times, for t+ 1 moment status, is the discount factor, α is the regularization coefficient of entropy, and the states also include unit load, total coal volume, main steam flow, main steam pressure, feed water flow, flue gas oxygen content, SCR inlet flue gas temperature and wind box pressure difference.

7. A boiler combustion strategy optimization system, characterized in that: include: A model building module, for establishing a prediction model of boiler thermal efficiency, NOx emission and superheated steam temperature by using a triangular convolutional neural network TR-CNN with adaptive width, and training the prediction model; A strategy acquisition module is used to construct a boiler multi-objective combustion optimization intelligent agent based on an actor-critic network SAC, and introduce a maximum entropy parameter to the boiler multi-objective combustion optimization intelligent agent, with the goal of obtaining the optimal combustion strategy for all boilers. Based on the current state, the boiler multi-objective combustion optimization intelligent agent with the maximum entropy parameter is introduced to obtain any combustion strategy with a value higher than a threshold; the combustion strategy includes boiler combustion parameters; The state output module is used to output the next state through the trained prediction model based on the combustion strategy and obtain the reward corresponding to the current state, where the current state represents the current boiler thermal efficiency, NOx emissions and superheated steam temperature parameters, and the reward corresponding to the current state represents the stability of the superheated steam temperature; The strategy optimization module is used to iteratively optimize the boiler multi-objective combustion optimization agent with the maximum entropy parameter according to the current state, combustion strategy, next state and the reward corresponding to the current state, until the combustion strategy with the optimal reward is obtained as the optimal boiler combustion strategy.

8. A computer device, characterized in that: It comprises a memory and a processor, wherein a program is stored in the memory, and when the program is executed by the processor, the processor executes the steps of a boiler combustion strategy optimization method as claimed in any one of claims 1 to 6.

9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a boiler combustion strategy optimization method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Coal-fired power plant boiler ash deposition prediction method

    CN115759279A

  • Boiler combustion control method under deep adjustment

    CN117404650A