Boiler combustion optimization method and device, computer equipment and storage medium

By using the combination method of channel selection convolution layer and TD3 agent in the boiler, boiler combustion control is optimized, the combustion control problem is solved under deep peak condition, and efficient and environmentally friendly boiler operation is achieved.

CN119957943APending Publication Date: 2025-05-09CHINA UNIV OF MINING & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510093311.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art is difficult to effectively optimize combustion control under boiler depth peak condition, resulting in problems such as low boiler thermal efficiency, increased NOx emissions and water-cooled wall overtemperature.

Method used

The convolutional neural network (CNN) is replaced with the channel selection convolution layer, and a dual-delay depth deterministic strategy gradient (TD3) agent is constructed, coupled with the boiler state prediction model, and the combustion control strategy is optimized through reinforcement learning.

Benefits of technology

Real-time and precise control of the boiler combustion process is achieved, the thermal efficiency and operating efficiency of the boiler are improved, NOx emissions are reduced, and the boiler operates under safe and environmentally friendly conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119957943A_ABST
    Figure CN119957943A_ABST
Patent Text Reader

Abstract

The invention provides a boiler combustion optimization method and device, computer equipment and a storage medium, and belongs to the technical field of boiler control. The method comprises the steps that an input variable is determined; constructing a boiler state prediction model based on a channel selection convolutional neural network CS-CNN; constructing a double-delay depth deterministic strategy gradient TD3 intelligent agent, and interacting the TD3 intelligent agent with the boiler state prediction model to obtain a combustion control optimization model; and inputting the current state of the boiler into the combustion control optimization model to obtain a combustion control optimization strategy of the boiler, and optimizing the input variable according to the combustion control optimization strategy. Thus, the CS-CNN is coupled with the TD3, so that the TD3 can perform optimization control on the combustion state under the variable load working condition of the boiler, and the TD3 intelligent agent has the capability of coping with the transient load of the boiler, thereby being beneficial to improving the decision-making efficiency of the boiler.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of boiler control, and in particular relates to a boiler combustion optimization method, device, computer equipment and storage medium. Background Art

[0002] In the process of power supply, the power grid needs coal-fired power generation to undertake deep peak-shaving and frequency-regulating tasks. Under deep peak-shaving, the unit is forced to operate at ultra-low load, the combustion atmosphere in the furnace deteriorates, resulting in worse heat transfer, which can easily cause overheating and tube bursting of the heating surface; a series of chain reactions such as parameter quality damage lead to low boiler thermal efficiency, and the flue gas temperature decreases and the flow field is uneven under ultra-low load, resulting in an imbalance in the catalytic reduction environment of pollutants. In addition, frequent overheating of the water-cooled wall will affect the deep regulation capability of the unit. In this context, it is urgent to establish a combustion control method suitable for deep peak-shaving and load transient conditions.

[0003] The research on boiler combustion process modeling includes two branches: mechanism and data-driven. The mechanism method is based on the kinetic model and energy conservation law involved in combustion and related processes, and then establishes a mathematical model that reflects the process mechanism. Compared with the mechanism method, the data-driven method uses machine learning algorithms to achieve fast modeling speed and high accuracy.

[0004] In the prior art, for the optimization of boiler combustion process, algorithm models including particle swarm optimization algorithm PSO and its variants and genetic algorithm GA and its variants have been established. However, each optimization process of the above optimization algorithms is independent of each other. For objects with strong time series such as boilers, there is a correlation between the previous and subsequent working conditions. Therefore, traditional optimization methods such as PSO and GA will cause lag in boiler control and are difficult to apply to the transient load conditions under the current deep peak load regulation conditions. Summary of the invention

[0005] In order to solve the problem of boiler combustion optimization under load transient conditions, the present invention provides a boiler combustion optimization method, device, computer equipment and storage medium.

[0006] In order to achieve the above object, the present invention provides the following technical solutions:

[0007] First, a boiler combustion optimization method is provided, the method comprising:

[0008] Obtain relevant variables that have an impact on the combustion state of the boiler, perform feature importance screening on the relevant variables, and obtain a preset number of strongly relevant variables as input variables; the relevant variables include boiler operation parameters and fuel parameters;

[0009] The conventional convolutional layer in the convolutional neural network (CNN) model is replaced with a channel selection convolutional layer to obtain a channel selection convolutional neural network (CS-CNN); the CS-CNN is trained to obtain a boiler state prediction model;

[0010] Construct a double-delayed deep deterministic policy gradient TD3 agent, which determines the optimization strategy for boiler combustion based on the current state of the boiler, adjusts the values ​​of input variables according to the optimization strategy, and outputs the next state of the boiler through a prediction model;

[0011] Determine the network parameters of the Actor network and the Critic network in the TD3 agent according to the current state, the optimization strategy and the next state, determine the loss value based on the network parameters, train the Actor network and the Critic network with the goal of minimizing the loss value, and obtain a combustion control optimization model;

[0012] The state parameters of the current state of the boiler are input into the combustion control optimization model, and the combustion control optimization strategy of the boiler is determined by the combustion control optimization model, and then the input variables are adjusted according to the combustion control optimization strategy; and the combustion state of the boiler is controlled according to the adjusted input variables.

[0013] Optionally, the related variables are screened by feature importance to obtain a preset number of strongly related variables as input variables, specifically including:

[0014] The importance of the features of the relevant variables is evaluated by the random forest method RF, and the formula is:

[0015]

[0016] in, are the prediction errors of the data of the relevant variables before and after the interference on the decision tree b, and B is the number of decision trees. The sum of the importance of all input variables is 1, that is, i=1,2,…,m represents the number of related variables;

[0017] According to the evaluation results, the related variables are sorted by feature importance, and a preset number of related variables at the top of the sorting are determined as strongly related variables;

[0018] The strongly correlated variable is determined as an input variable.

[0019] Optionally, the channel selection convolution layer includes an expected channel damage matrix, and the expected channel damage matrix is ​​used to identify the importance of input channels in the channel selection convolution layer, screen a preset number C of high-importance channels, and release hyperparameters of remaining low-importance channels.

[0020] Optionally, the channel selection convolution layer is also used to perform channel unblocking, channel reallocation and spatial shifting; the channel unblocking prevents related variables with low importance from being used in subsequent model calculations and releases the parameters of these channels; the channel reallocation replaces the blocked channels with preset C high-importance channels; the spatial shifting offsets the feature loss caused by too many similar feature channels.

[0021] Optionally, the training of CS-CNN to obtain a boiler state prediction model includes:

[0022] Obtaining a boiler combustion training sample; the training sample includes a sample input variable and its corresponding real state;

[0023] Input the sample input variable into CS-CNN to obtain a predicted state;

[0024] The CS-CNN is trained with the goal of minimizing the deviation between the true state and the predicted state to obtain the boiler state prediction model.

[0025] Optionally, determining network parameters of the Actor network in TD3 according to the current state and the optimization strategy, and determining the loss value based on the network parameters includes:

[0026] The Q values ​​of the two Critic networks are determined according to the current state and the optimization strategy, and the loss value of the Actor network is determined based on the smaller value of the two Q values. The Q values ​​are Q1(s, a) and Q2(s, a), respectively. The specific formula is:

[0027] loss actor =-min(Q 1,2 (s,a));

[0028] Among them, s is the current state and a is the optimization strategy.

[0029] Optionally, determining network parameters of a Critic network in TD3 according to the current state and the optimization strategy, and determining the loss value based on the network parameters includes:

[0030] The loss functions of the two online updated Critic networks use the temporal difference error algorithm to calculate the mean square error between Q1(s,a) and r+γ·min(Q1′,Q2′), and between Q2(s,a) and r+γ·min(Q1′,Q2′) as the loss values. The specific formulas are:

[0031] loss critic1 =MSE[Q1(s,a),r+γ·min(Q1′,Q2′)];

[0032] losscritic2 =MSE[Q2(s,a),r+γ·min(Q1′,Q2′)]

[0033] Among them, Q1′ and Q2′ are the Q values ​​of the two delayed update target Critic networks, r is the immediate reward obtained after adjusting the input variables according to the optimization strategy in the current state, and γ is the discount factor that balances the immediate reward and future rewards.

[0034] Secondly, a boiler combustion optimization device is provided, the device comprising:

[0035] An acquisition module is used to acquire relevant variables that affect the state of the boiler, perform random forest RF feature importance evaluation on the relevant variables, and determine input variables;

[0036] A construction module is provided for constructing a convolutional neural network (CNN) model, replacing the conventional convolutional layer in the CNN model with a channel selection convolutional layer to obtain a channel selection convolutional neural network (CS-CNN); the CS-CNN includes an input layer, a channel selection convolutional layer, a fully connected layer, and an output layer; the CS-CNN is trained to obtain a boiler state prediction model; a dual-delay deep deterministic policy gradient (TD3) agent is constructed, including the design and training process of an Actor network and a Critic network, and the TD3 agent interacts with the boiler state prediction model, including: the agent adjusts the input variables according to the current state of the boiler to obtain an optimization strategy for boiler combustion, adjusts the state variables according to the optimization strategy, and outputs the next state of the boiler through a prediction model; the loss values ​​of the Actor network and the Critic network in TD3 are determined according to the current state and the optimization strategy, and TD3 is trained with the goal of minimizing the loss value to obtain a combustion control optimization model;

[0037] The optimization module is used to input the current state of the boiler into the combustion control optimization model, determine the combustion control optimization strategy of the boiler through the combustion control optimization model, and then adjust the input variables according to the combustion control optimization strategy.

[0038] In addition, a computer-readable storage medium is provided, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned boiler combustion optimization method is implemented.

[0039] Finally, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned boiler combustion optimization method when executing the program.

[0040] The boiler combustion optimization method provided by the present invention has the following beneficial effects:

[0041] First, relevant variables are obtained and screened, which can not only reduce the computational complexity of subsequent models, but also improve the prediction accuracy and optimization effect of the model; secondly, a boiler state prediction model is constructed, and the conventional convolution layer is replaced by a channel selection convolution layer. This can evaluate the importance of the input channel, screen and retain the information of the key channel, and release the resources of the low-importance channel at the same time, which not only improves the parameter utilization efficiency of the model, but also effectively prevents the overfitting problem caused by input redundancy; then a TD3 agent is constructed, and coupled and interacted with the boiler state prediction model, and the TD3 agent is trained based on the interaction results. In this way, the network parameters of the TD3 agent are continuously iterated and optimized through model coupling, so that the TD3 agent can more accurately evaluate the long-term value of different state-action pairs, and output a better control strategy to better adapt to various complex working conditions in the boiler combustion process. Finally, the current state of the boiler is input into the trained TD3 agent, and the control strategy based on the current state output by the TD3 agent can be obtained. The values ​​of the input variables are adjusted according to these strategies, which can realize real-time and precise control of the boiler combustion process. This control method can perfectly cope with the transient load of the boiler under deep peak-shaving conditions, and can instantly adjust the input variables according to the current state of the boiler. It not only improves the operating efficiency of the boiler, but also ensures that the boiler operates under safe and environmentally friendly conditions. It can significantly improve the thermal efficiency of the boiler and reduce NOx emissions, thereby achieving the goals of energy conservation, emission reduction and sustainable development. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the embodiment of the present invention and its design scheme, the following briefly introduces the drawings required for this embodiment. The drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0043] Figure 1 The present invention is a schematic flow chart of a boiler combustion optimization method provided according to an exemplary embodiment of the present invention.

[0044] Figure 2 The variable importance evaluation results of NOx emission and thermal efficiency based on RF are provided according to an exemplary embodiment of the present invention.

[0045] Figure 3 A schematic diagram of reconstruction of input data of a CS-CNN prediction model provided by the present invention according to an exemplary embodiment.

[0046] Figure 4 A schematic diagram of a neural network for approximating a strategy function according to an exemplary embodiment of the present invention.

[0047] Figure 5A schematic diagram of a neural network for approximating a value function provided by the present invention according to an exemplary embodiment.

[0048] Figure 6 A schematic diagram of updating a combustion optimization instruction according to an exemplary embodiment of the present invention is provided.

[0049] Figure 7 It is a block diagram of a boiler combustion optimization device provided according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0050] In order to enable those skilled in the art to better understand the technical solution of the present invention and implement it, the present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the scope of protection of the present invention.

[0051] The present invention constructs thermal efficiency, NO x Multi-objective integrated prediction model for emissions and wall temperature; focus on coupling the multi-objective combustion optimization agent with the boiler state prediction model to build a multi-objective combustion process decision optimization framework based on policy iteration. An approximation network of the policy function and the value function is designed for the boiler state and manipulated variables, giving the agent the ability to flexibly respond to unfamiliar boiler conditions. In particular, the reward function is carefully designed to motivate the agent to make more reasonable combustion decision instructions, and interaction rules for optimization strategies and state updates are designed for the agent and the boiler state prediction model. In order to avoid the problem that the Q value of the value network is overestimated and the policy function destroys the combustion decision, the Twin Delayed Depth Deterministic Policy Gradient (TD3) is selected as the agent. In order to improve the optimization effect of the TD3 agent on the combustion process, noise is superimposed on the manipulated variables output by the policy function to enhance the exploration ability of the optimization strategy.

[0052] The technical solutions provided by various embodiments of the present invention are described in detail below in conjunction with the accompanying drawings.

[0053] First, the present invention provides a boiler combustion optimization method, specifically as follows Figure 1 As shown, the following steps are included:

[0054] S101. Obtain input variables.

[0055] Specifically, the relevant variables that affect the state of the boiler can be obtained first, and the relevant variables are screened by feature importance to obtain a preset number of strongly relevant variables as input variables. Among them, the feature importance of the relevant variables can be evaluated by the random forest method RF, and the relevant variables are ranked by feature importance according to the evaluation results, and the preset number of relevant variables in the top of the ranking are determined as strongly relevant variables, and the strongly relevant variables are determined as input variables.

[0056] Taking the W flame boiler as an example, the combustion state of the boiler includes NO x One or more of the various combustion conditions such as emissions and thermal efficiency; for example, combined with the characteristics of the W-flame boiler combustion system, the influence of boiler NO x The variables related to emissions and thermal efficiency were integrated to form a comprehensive table of related variables based on the combustion process mechanism, expert experience, and the advice of operation engineers. About one month of historical data was extracted from the boiler's DCS, with a sampling period of 1 minute. Through further analysis of the data, the operating intervals with a large number of missing data were eliminated, and the data that met the requirements and was running continuously were selected.

[0057] Impact on boiler thermal efficiency and NO x The emission-related variables were evaluated for their importance using random forest RF features, such as Figure 2 shown.

[0058] For example, the relevant variables include unit load, total fuel quantity, capacity air flow of each coal mill, burnout air damper opening, F layer secondary air damper opening, etc., a total of 52 variables. The specific relevant variables can be shown in Table 1 below:

[0059] Table 1 Related variables

[0060]

[0061]

[0062] RF uses several independent decision trees for parallel prediction. The random sampling with replacement (BootstrapSampling) makes the out-of-bag data naturally have the function of evaluating the deterioration of the model prediction error. The degree of deterioration indirectly indicates the importance of the relevant variable. For each decision tree, the in-bag data is used to train the tree, and the out-of-bag data OOB is used to evaluate the tree. The deterioration of prediction accuracy will be achieved by applying noise interference to important variables. Before and after the interference, the importance measurement of the input related variables is given by the following formula:

[0063]

[0064] in, are the prediction errors of the data of the relevant variables before and after the interference on the decision tree b, and B is the number of decision trees. The sum of the importance of all relevant variables is 1, that is, i=1,2,…,m represents the number of related variables.

[0065] The NO obtained by RF x The importance variables of emissions and thermal efficiency are merged, duplicate variables are removed, and new variables are obtained after reordering as input variables.

[0066] For example, from the NOx emission calculation results, the total weight of the top 16 variables in importance reached 0.902, so 16 variables were selected as input variables for the NOx emission model. From the thermal efficiency calculation results, the total weight of the top 10 variables in importance reached 0.933, so 10 variables were selected as input variables for the thermal efficiency model. The above two groups of input variables are merged, duplicate variables are removed, and re-arranged to obtain the 23 input variable tables in Table 2, which are used as input variables for the boiler NOx emission, thermal efficiency and maximum wall temperature multi-objective prediction model. The input variables can be shown in Table 2 as follows:

[0067] Table 2 Input variables

[0068]

[0069]

[0070] S102: Construct a boiler state prediction model.

[0071] Specifically, a convolutional neural network (CNN) model is constructed, and the conventional convolutional layer in the CNN model is replaced with a channel selection convolutional layer to obtain a channel selection convolutional neural network (CS-CNN); the CS-CNN includes an input layer, a channel selection convolutional layer, a fully connected layer, and an output layer; the CS-CNN is trained to obtain a boiler state prediction model. The training process of the CS-CNN can be as follows: obtain a training sample of boiler combustion; the training sample includes a sample input variable and its corresponding true state; the sample input variable is input into the CS-CNN to obtain a predicted state; the CS-CNN is trained with the goal of minimizing the deviation between the true state and the predicted state to obtain a boiler state prediction model.

[0072] In one embodiment, the channel selection convolution layer includes an expected channel damage matrix, which is used to identify the importance of input channels in the channel selection convolution layer, screen a preset number of high-importance channels, and release the hyperparameters of the remaining low-importance channels, thereby maximizing the utilization of limited parameters of the model and preventing overfitting problems caused by input redundancy.

[0073] In addition, the channel selection convolution layer is also used for channel de-alloc, channel reallocation and spatial shift; channel de-alloc prevents related variables with low importance from being used in subsequent model calculations and releases the parameters of these channels; channel reallocation replaces blocked channels with a preset number of high-importance channels; and spatial shift offsets feature loss caused by too many similar feature channels.

[0074] For example, in the channel importance identification task, let Conv(X,W) represent the conventional convolution operation, X represents the input variable of the convolution layer, and W represents the weight matrix of the convolutional layer, and Among them, I and O represent the number of input and output channels, respectively, and H and W represent the height and width of the input channel, respectively. Since the boiler historical data are arranged in chronological order, setting the convolution kernel size needs to consider the feature extraction at the same time, so the convolution kernel size is set to K l ×1. And i=1,2,...,I,h=1,2,...,H,w=1,2,...,W. Define ECDM(X;W) to calculate the expected mean in the time dimension, as follows:

[0075]

[0076] When W i =0, it means that the i-th channel has been set as a low-importance channel and needs to be blocked.

[0077] In defining the channel selection and spatial shift tasks, the channel selection convolution layer is defined as SelectConv(X,W), which implements the selection function of the input channel X based on the conventional convolution layer, so the channel selection convolution layer can be expressed as:

[0078] SelectConv(X;W)=Conv(SelectChannel(X);W);

[0079] Among them, SelectChannel needs to perform channel blocking and re-indexing for De-alloc and Re-alloc respectively, as follows Figure 3 As shown. So we introduce the gate variable g i ∈{0,1} and index π i ∈{1,2,…,I}. The input channel X can be expressed as if Does not contribute to the output, the corresponding g i = 0. Apply spatial offset to the reallocated channels SelectChannel is described as follows:

[0080]

[0081] i={1,2,…,I},

[0082] π i ={1,2,…,I},

[0083] g i ∈{0,1},

[0084]

[0085] Among them, shift(X,b) represents the spatial shift operation of the input channel X. For each element (x, y) on the channel, shift(X,b) x,y Defined as:

[0086]

[0087] The variable b is continuous and is learned jointly with other hyperparameters via stochastic gradient descent during model training.

[0088] Training scheme for channel selection convolutional layer: In order to use ECDM to identify the channels to be de-allocated / reallocated, the parameter S = (X; g, π, b) is cleverly designed in the channel selection convolutional layer, and then S is iteratively updated during the training process of de-alloc and re-alloc. For de-allocating channels, by setting g i = 0. Given the expected damage level γ>0, the unlocking goal can be designed as the following optimization problem:

[0089]

[0090] g i =0or1;i=1,2,…,I;

[0091]

[0092] Assuming j = 1, 2, ..., O, after calculating nECDM, the minimum channel || nECDM (X; W) i || ∞ Iterate to determine the channel to be released, but it is necessary to ensure that the vector nECDM(W;X) i l ∞ norm is less than γ. For channel reallocation Re-alloc, according to the l of nECDM 2 The norm selects the Top-C maximum channel, i.e. ||nECDM(W;X) i ||2.

[0093] S103, constructing a TD3 intelligent agent, and interacting the TD3 intelligent agent with a boiler state prediction model.

[0094] A double-delayed deep deterministic policy gradient TD3 is constructed, and the TD3 agent interacts with the boiler state prediction model. The interaction process includes: the TD3 agent determines the optimization strategy of boiler combustion according to the current state of the boiler (that is, the action of TD3 on the input variables of the boiler), adjusts the values ​​of the input variables according to the optimization strategy, and outputs the next state of the boiler through the prediction model.

[0095] In this step, for multi-objective optimization problems under nonlinear complex constraints, a reinforcement learning TD3 agent suitable for continuous action space is introduced, focusing on the design of the function approximation network in the agent. The core of the boiler combustion process optimization method based on reinforcement learning is the agent. The agent interacts with the boiler state prediction model to explore and learn the combustion optimization strategy, improve the boiler thermal efficiency and reduce NO while ensuring that the boiler wall temperature does not exceed the warning value. x The controlled variables include the unit coal feed rate, secondary air and OFA damper opening, and the means of combustion optimization include adjusting the unit coal feed rate, secondary air and OFA damper opening.

[0096] The present invention constructs a combustion process optimization algorithm architecture based on TD3. In this architecture, the TD3 agent interacts with the boiler state prediction model, and the interaction records will be collected by the buffer. Each record is a four-tuple (state s, action a, reward r, next state s′). The experience in the playback buffer may come from different strategies, which is more conducive to the Q network learning different experiences. If the buffer reaches the set capacity, the old experience is abandoned. The buffer provides a batch of policy experience for experience playback and provides a data basis for the Q network. In order to solve the problem of overestimation of Q value, the TD3 agent adds two Q networks on the basis of deep deterministic policy gradient DDPG, namely: the online updated Q2 network and the corresponding delayed update target Q2′ network. Therefore, the agent includes the policy function, value function and reward function, and the number of neural networks used for function approximation reaches six.

[0097] Based on the above description, the input of the strategy function in TD3 constructed by the present invention is the observable state quantity in the boiler state prediction model, and the output is the control quantity that acts on the boiler state prediction model to prompt the boiler to transfer its state to the next moment. Therefore, based on the analysis of the boiler combustion operation law, the unit load, total coal volume, main steam flow, main steam pressure, flue gas oxygen content, exhaust gas temperature, boiler NO xEmissions, thermal efficiency and maximum wall temperature are used as state variables to observe the current combustion status of the boiler. The reason why variables such as unit load, total coal volume and main steam flow are added as state variables is that they have the same boiler NO x There are multiple combustion conditions that can affect emissions and thermal efficiency. x Emissions and thermal efficiency cannot distinguish the comprehensive state of the boiler. In order to optimize the coal and air distribution methods, 14 controlled variables, including four pulverizer capacity air damper openings, two OFA damper openings, and eight secondary air damper openings, are selected as the control variables for optimizing the boiler combustion process.

[0098] In order to meet the policy function approximation, a policy network is designed. The policy network is a four-layer deep neural network - Multilayer Perceptron (MLP), such as Figure 4 The network input layer shown is composed of 9 neurons, corresponding to 9 variables representing the combustion state of the boiler. The network output layer is composed of 14 neurons, corresponding to 14 combustion optimization controlled variables. In addition, the network is set up with two hidden layers, each of which is set up with 256 neurons.

[0099] like Figure 5 The figure shows a neural network that approximates the value function. The value network is a four-layer deep neural network MLP. The network input layer consists of 23 neurons, the output layer consists of 1 neuron, and the network has two hidden layers, each with 256 neurons.

[0100] The reward function belongs to the objective function in combustion optimization. It is an indicator for evaluating the quality of the state change before and after the control amount acts on the boiler state prediction model. It is a function that maps from the boiler environment state and the actions taken by the intelligent agent to the actual reward value, that is, the immediate effect of the decision-making behavior. The reward function provides an immediate, local feedback to guide the intelligent agent to make appropriate decisions at each step. However, the evaluation feedback of the real boiler environment has a certain delay, and focuses more on short-term evaluation and lacks long-term cumulative benefit evaluation. Therefore, it is necessary to introduce a value function. Reward function design, the goal of boiler combustion optimization is to improve the thermal efficiency of the boiler and reduce NO x Discharge, to prevent water wall overheating, so the design optimization objective function is as follows:

[0101] minΨ=-max(C η ·Δη+C NOx ·ΔNO x +C wall ·(T wall -T warn ))

[0102] sA min ≤A≤Amax ,

[0103] C η ≥0,

[0104] C NOx ≤0,

[0105] C wall ≤0

[0106] Among them, variables η, NO x and T wall They are respectively the boiler thermal efficiency, NO x Discharge and water wall maximum temperature, variables Δη and ΔNO x are the changes in boiler thermal efficiency and NO before and after optimization. x Emission change, T warn is the water wall overtemperature alarm limit. The optimization objective function is composed of reward boiler thermal efficiency term, penalty NO x The emission item and the over-temperature alarm item are composed of C η , C NOx , C wall are their weight coefficients respectively. In order to encourage the intelligent agent to make decisions to improve thermal efficiency and reduce pollutants within the safe wall temperature range, the boiler thermal efficiency weight coefficient C η ≥0, NO x Emission weight coefficient C NOx ≤0, over-temperature alarm weight coefficient C wall ≤0. In order to ensure the rationality of decision variables, set the upper and lower limits of all controlled variables A, that is, A max , A min .

[0107] S104. Train the Actor network and the Critic network in the TD3 agent according to the current state and the interaction result to obtain a combustion control optimization model.

[0108] Specifically, the network parameters of the Actor network and the Critic network in the TD3 agent can be determined according to the current state, the optimization strategy and the next state, and the loss value can be determined based on the network parameters. The Actor network and the Critic network are trained with the goal of minimizing the loss value to obtain a combustion control optimization model.

[0109] In one embodiment, specifically Figure 4 The policy function is approximated by the neural network and Figure 5 The neural network that approximates the value function shown in the figure outputs the optimal control quantity at the next moment according to the state of the boiler combustion system through the policy network that approximates the policy function, and maps the boiler combustion state-action pair to a real value through the value function, providing important guiding information for the decision-making of the intelligent agent.

[0110] In the above steps, the strategy network is first designed. The strategy network used to approximate the strategy function outputs the optimal control quantity at the next moment according to the state of the boiler combustion system. Among the six networks of the intelligent agent, the strategy network is the Actor network, the input is the state, and the output is the action. For the Actor network, the current state of the boiler is input and the optimization strategy is output, that is, the action space of the boiler controlled variable. It should be noted that the variable representing the current state of the boiler needs to be unique and can be represented by multiple variables. The purpose of the Actor network is to output the value that makes Q based on the current state s. 1,2 (s,a) The largest action a.

[0111] For the Actor network, the Q values ​​of the two Critic networks can be determined based on the current state and optimization strategy, and the loss value of the Actor network can be determined based on the smaller value of the two Q values. The Q values ​​are Q1(s, a) and Q2(s, a), respectively. The specific formula is:

[0112] loss actor =-min(Q 1,2 (s,a));

[0113] Among them, s is the current state and a is the optimization strategy.

[0114] In terms of network parameter optimization, the loss function of the Actor network is shown in the above formula.

[0115] The value network evaluates the expected cumulative reward that the agent will receive in the long term, that is, the expected value of future rewards that the agent can obtain in a specific state. The value function helps the agent measure the long-term value of different state-action pairs, thereby guiding the agent to consider future impacts when making decisions. By combining with the reward function, the agent can better balance short-term rewards and long-term goals. The value network can be a Critic network, which forms part of the Actor-Critic algorithm with the Actor network.

[0116] In the optimization problem of boiler combustion process, more emphasis is placed on emphasizing the influence of the output controlled quantity on the performance of the boiler, so the state-action function is selected as the expression form of the value function in the subsequent description. Similar to the method of approximating the strategy function, the present invention uses a deep neural network to approximate the value function, and its input includes state (unit load, total coal volume, main steam flow, main steam pressure, flue gas oxygen content, exhaust gas temperature, NOx emission, thermal efficiency and maximum wall temperature, a total of 9 state variables) and action (x3-x6 four coal mill capacity air damper openings, x14 and x15 two OFA damper openings, x16-x23 eight secondary air damper openings, a total of 14 controlled variables) a total of 23 variables, and the output is the evaluation of the change in boiler performance after the action acts on the boiler state prediction model.

[0117] In reinforcement learning tasks, rewards usually appear on a long time scale, requiring the agent to make correct decisions in the short term to achieve long-term goals. The reward function quantifies the goal in the optimization problem and evaluates the performance of the control quantity after it acts on the combustion system. It is equivalent to the objective function in the general optimization problem. Since the optimization goal of the present invention is to improve the thermal efficiency of the boiler and reduce NOx emissions under the constraint of safe wall temperature, when designing the reward function, it is necessary to reward the improvement of thermal efficiency, punish the increase of NOx emissions, and punish the wall temperature exceeding 475°C, as shown in the following formula:

[0118] reward=C η ·Δη+C NOx ·ΔNO x +C wall ·(T wall -475℃)

[0119] sA min ≤A≤A max ,

[0120] C η ≥0,

[0121] C NOx ≤0,

[0122] C wall ≤0;

[0123] In the formula, various parameters are finally determined through continuous debugging. The items that need to be clarified include: and optimization weight allocation and optimization instruction constraints:

[0124] Optimization weight allocation: Boiler combustion optimization first needs to prioritize economic indicators under the premise of ensuring safe operation, so thermal efficiency is the main optimization target. Since the thermal efficiency change Δη value is small, C ηThe reward weight coefficient is set to 200; and the NOx emission change ΔNO x The value is large, C NOx The penalty weight coefficient is set to -30; the penalty weight coefficient of wall temperature C wall is set to -75, and an appropriate reward is given under safe operating conditions where the wall temperature is below 475°C. At this time, the penalty weight coefficient of the wall temperature is C wall If set to -1, the reward value fluctuates with the distance from the alarm limit.

[0125] Optimization instruction constraints: The optimized instructions include 14 controlled variables, including four coal mill capacity air damper openings x3-x6, two OFA damper openings x14 and x15, and eight secondary air damper openings x16-x23. Set the upper and lower limits of all controlled variables A, i.e. A max =100%, A min =0%.

[0126] For the Critic network, the two online updated Critic network loss functions use the temporal difference error algorithm to calculate the mean square error between Q1(s,a) and r+γ·min(Q1′,Q2′), and between Q2(s,a) and r+γ·min(Q1′,Q2′). The specific formulas are:

[0127] loss critic1 =MSE[Q1(s,a),r+γ·min(Q1′,Q2′)];

[0128] loss critic2 =MSE[Q2(s,a),r+γ·min(Q1′,Q2′)]

[0129] Among them, Q1′ and Q2′ are the Q values ​​of the delayed updated target Critic network, r is the immediate reward obtained after adjusting the input variables according to the optimization strategy in the current state, and γ is the discount factor that balances the immediate reward and future rewards.

[0130] In the above steps, Q is updated online 1,2 The network is the Critic network, which takes state and action as input and outputs Q value. The purpose of the Critic network is to calculate the action value Q(s,a) based on the state-action pair (s,a). The loss function of the two Critic networks uses the time difference error to calculate Q 1,2 The mean square error between (s,a) and r+γ·min(Q1′,Q2′) is as shown above.

[0131] For the delayed updated target policy network, i.e., the target actor network, its input is the state s' at the next moment, and its output is the estimated action a'. In order to make the two online updated critic networks converge stably, the target actor network softly updates the parameters of the actor network to the target actor network through the delayed update method. The two delayed updated target Q1' and Q2' networks, i.e., the target critic networks, are also intended to make the two online updated critic networks converge stably, and their parameters are softly updated from the Q1 and Q2 networks. The reason why the TD3 algorithm uses two target Q1' and Q2' networks is that in practical applications, the critic network always overestimates the Q value. It draws on the idea of ​​DDQN, uses two networks to estimate the target Q value, and then selects the smaller Q value of the target Q1' and Q2' to avoid overestimating the target Q value as much as possible. In addition, TD3 also adds noise to the action estimated by the target actor network to generate a', which is used as the input of the two target Q1' and Q2' networks, so as to encourage the agent to explore, thereby making the next step of the Q value more accurate.

[0132] It should be noted that the Actor network is crucial and directly determines the quality of the agent's output decision (only the Actor network among the six networks in the agent interacts with the boiler state prediction model). To train a satisfactory Actor network, an accurate Critic network is needed to evaluate it. Therefore, the goal of the other five networks of TD3 is to create a Critic network that is as accurate as possible.

[0133] In addition, since the input format of the prediction model in the combustion process decision optimization system is a two-dimensional tensor, including the time dimension and the feature dimension, and the agent and the prediction model only update a single condition each time they interact, the update method of the optimization action in the two-dimensional tensor is designed. t When , according to the policy network output, optimize action a t For example, the current state is the boiler state variable of the 20th time dimension in a single tensor, and the optimization action will be updated to the action variable on the 20th time dimension in the next tensor, as shown in Figure 6 shown.

[0134] S105. Control the combustion state of the boiler through the combustion control optimization model.

[0135] Specifically, the state parameters of the current state of the boiler are input into the combustion control optimization model, the combustion control optimization strategy of the boiler is determined by the combustion control optimization model, and then the input variables are adjusted according to the combustion control optimization strategy; and the combustion state of the boiler is controlled according to the adjusted input variables.

[0136] Based on the boiler state prediction model and combustion control optimization model designed in the above steps, the boiler combustion control strategy is optimized to obtain the optimized controlled variables, and then the boiler combustion is controlled according to the controlled variables.

[0137] The present invention has carried out data analysis and optimization simulation tests on a W-flame coal-fired boiler of a 600MW supercritical thermal power unit in a power plant. The boiler of this unit is a direct current boiler with supercritical parameters and variable pressure operation. The boiler is equipped with 6 double-inlet and double-outlet coal mills, 24 double-cyclone direct current pulverized coal burners that are specially used to burn low-volatile coal, and each coal mill has 4 burners, wherein the pulverized coal burners are arranged on the front and rear arches of the furnace, with 12 burners on each front and rear arches. The secondary air of each burner is controlled separately, and the air distribution unit consists of an upper wind box and a lower wind box. The A, B, and D dampers are set to open during combustion adjustment and are not adjusted during daily operation. The C damper mainly controls the air volume required for the combustion of the ignition and stabilization oil gun. The F damper has the largest air volume and controls the main secondary air volume required for combustion. Therefore, the opening of the F-layer secondary air damper of the 24 burners has a significant effect on the combustion process and should be included in the auxiliary variables required for modeling. In addition, in order to achieve low NOx emissions from the boiler, 26 OFA air regulators are arranged on the water-cooled walls on the front and rear walls of the boiler arch. The optimal position of the OFA air regulator is determined during the combustion adjustment experiment during the boiler trial operation. As long as the coal type does not change significantly, no adjustment is required. The OFA total air volume is adjusted by the OFA wind box inlet damper actuator, and two A and B are set on the front and rear walls of the actuator. About one month of historical data was extracted from the DCS of the target boiler, with a sampling period of 1 minute, totaling about 43,200 sampling points. The sampling data covers various operating conditions such as steady state and variable load, which is sufficient to support the comparison and analysis of various operating conditions during the method verification. Through further analysis of the data, the operating intervals with more vacant data were eliminated, and historical data that met the requirements and operated continuously for about 14 days were selected, totaling about 20,160 sampling moments. The operating intervals of the variables during this period are also counted in Table 1.

[0138] In this embodiment, the CS-CNN model was trained and parameterized using the training set. The learning rate during model training was set to 0.001 and the batch size was set to 64. The RMSE of NOx emissions predicted by CS-CNN was 9.099±4.072mg / m3, the RMSE of thermal efficiency was 0.143±0.028%, and the RMSE of the maximum wall temperature was 2.586±0.622℃. The prediction results on the test set show that the CS-CNN multi-objective dynamic prediction model can effectively and accurately predict NOx emissions, thermal efficiency and wall temperature of coal-fired boilers, and the prediction accuracy can meet the needs of interaction with deep reinforcement learning agents.

[0139] In this embodiment, in order to ensure that the agent has enough learning samples to be familiar with as many working conditions as possible, the first 19540 samples are used as the training set of the optimizer, and the last 600 samples are used as the test set. The test set samples cover the boiler's rapid load change and stable load conditions. The agent-related hyperparameters are set as follows: the hidden layers of the policy network and the Q network in the two agents are both 2 layers, each layer contains 256 neurons, the policy network learning rate is 10-4, and the Q network learning rate is 10-3. The number of Mini-batch samples during buffer training is 128, the number of episodes of interaction between the agent and the prediction model is 200, and the trajectory length of each round is 4. The agent adopts the training and testing scheme for iterative learning, that is, as the training progresses, a test is performed every 200 rounds, the test set contains 600 continuous working conditions, and the Epoch of the test experiment is 300. Through statistical analysis, the range of boiler thermal efficiency before and after optimization is 0.108-0.767%, and the range of NOx emissions is -34.847--1.574 mg / m3, that is, the thermal efficiency of all test conditions has been improved and NOx emissions have been reduced. From the optimization mean, the mean thermal efficiency has increased by 0.411%, and the mean NOx emissions have decreased by 17.701 mg / m3. It can be seen that the decisions made based on the TD3 agent can ensure that the wall temperature does not exceed the temperature.

[0140] In actual operation, we will not simply pursue economic efficiency or nitrogen oxide emissions, but usually comprehensively optimize economic efficiency and pollutant generation. The strategy iteration time of the TD3 agent was counted. Among the 600 test samples, the maximum single calculation time of the agent was 0.004s, which achieved instantaneous decision-making and theoretically met the time requirements of real-time optimization. This shows that reinforcement learning has the advantage of less time consumption in strategy iteration.

[0141] By adopting the above method, firstly, relevant variables are obtained and screened, which can not only reduce the computational complexity of subsequent models, but also improve the prediction accuracy and optimization effect of the model; secondly, a boiler state prediction model is constructed, and the conventional convolution layer is replaced by a channel selection convolution layer, which can screen and retain the information of key channels by evaluating the importance of input channels, and release the resources of low-importance channels at the same time, which not only improves the parameter utilization efficiency of the model, but also effectively prevents the overfitting problem caused by input redundancy; then, a TD3 agent is constructed, and coupled and interacted with the boiler state prediction model, and the TD3 agent is trained based on the interaction results. In this way, the network parameters of the TD3 agent are continuously iterated and optimized through model coupling, so that the TD3 agent can more accurately evaluate the long-term value of different state-action pairs, and output a better control strategy to better adapt to various complex working conditions in the boiler combustion process. Finally, the current state of the boiler is input into the trained TD3 agent, and the control strategy based on the current state output by the TD3 agent can be obtained. The values ​​of the input variables are adjusted according to these strategies, so that real-time and precise control of the boiler combustion process can be achieved. This control method can perfectly cope with the transient load of the boiler under deep peak-shaving conditions, and can instantly adjust the input variables according to the current state of the boiler. It not only improves the operating efficiency of the boiler, but also ensures that the boiler operates under safe and environmentally friendly conditions. It can significantly improve the thermal efficiency of the boiler and reduce NOx emissions, thereby achieving the goals of energy conservation, emission reduction and sustainable development.

[0142] Secondly, the present invention also provides a boiler combustion optimization device, such as Figure 7 As shown, including:

[0143] The acquisition module 701 is used to acquire relevant variables that have an impact on the state of the boiler, perform random forest RF feature importance evaluation on the relevant variables, and determine input variables.

[0144] Construction module 702 is used to construct a convolutional neural network (CNN) model, replace the conventional convolutional layer in the CNN model with a channel selection convolutional layer, and obtain a channel selection convolutional neural network (CS-CNN); the CS-CNN includes an input layer, a channel selection convolutional layer, a fully connected layer, and an output layer; train the CS-CNN to obtain a boiler state prediction model; construct a dual-delay deep deterministic policy gradient TD3 agent, including the design and training process of the Actor network and the Critic network, and interact the TD3 agent with the boiler state prediction model, including: the agent adjusts the input variable according to the current state of the boiler to obtain an optimization strategy for boiler combustion, adjusts the state variable according to the optimization strategy, and outputs the next state of the boiler through the prediction model; determines the loss value of the Actor network and the Critic network in TD3 according to the current state and the optimization strategy, trains TD3 with the goal of minimizing the loss value, and obtains a combustion control optimization model.

[0145] The optimization module 703 is used to input the current state of the boiler into the combustion control optimization model, determine the combustion control optimization strategy of the boiler through the combustion control optimization model, and then adjust the input variables according to the combustion control optimization strategy.

[0146] By using the above device, firstly, relevant variables are obtained and screened, which can not only reduce the computational complexity of subsequent models, but also improve the prediction accuracy and optimization effect of the model; secondly, a boiler state prediction model is constructed, and the conventional convolution layer is replaced by a channel selection convolution layer, so that by evaluating the importance of the input channel, the information of the key channel can be screened and retained, and the resources of the low-importance channel can be released at the same time, which not only improves the parameter utilization efficiency of the model, but also effectively prevents the overfitting problem caused by input redundancy; then a TD3 agent is constructed, and coupled and interacted with the boiler state prediction model, and the TD3 agent is trained based on the interaction results, so that the network parameters of the TD3 agent are continuously iterated and optimized through model coupling, so that the TD3 agent can more accurately evaluate the long-term value of different state-action pairs, and output a better control strategy, so that it can better adapt to various complex working conditions in the boiler combustion process, and finally the current state of the boiler is input into the trained TD3 agent, and the control strategy based on the current state output by the TD3 agent can be obtained. According to these strategies, the values ​​of the input variables are adjusted to achieve real-time and precise control of the boiler combustion process. This control method can perfectly cope with the transient load of the boiler under deep peak-shaving conditions, and can instantly adjust the input variables according to the current state of the boiler. It not only improves the operating efficiency of the boiler, but also ensures that the boiler operates under safe and environmentally friendly conditions. It can significantly improve the thermal efficiency of the boiler and reduce NOx emissions, thereby achieving the goals of energy conservation, emission reduction and sustainable development.

[0147] The present invention also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 Steps of a boiler combustion optimization method are provided.

[0148] The present invention also provides a computer device. At the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 Steps of a boiler combustion optimization method are provided.

[0149] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0150] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as a combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0151] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0152] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0153] It should be noted that the above specific implementation method can enable those skilled in the art to understand the invention more comprehensively, but does not limit the invention in any way. Therefore, although the present invention has been described in detail in this specification, those skilled in the art should understand that the invention can still be modified or replaced by equivalents; and all technical solutions and improvements that do not deviate from the spirit and scope of the invention are included in the protection scope of the patent for the invention. Any figure mark in the claims should not be regarded as limiting the claims involved.

Claims

1. A boiler combustion optimization method, characterized in that: The method comprises: Obtain relevant variables that have an impact on the combustion state of the boiler, perform feature importance screening on the relevant variables, and obtain a preset number of strongly relevant variables as input variables; the relevant variables include boiler operation parameters and fuel parameters; The conventional convolutional layer in the convolutional neural network (CNN) model is replaced with a channel selection convolutional layer to obtain a channel selection convolutional neural network (CS-CNN); the CS-CNN is trained to obtain a boiler state prediction model; Construct a double-delayed deep deterministic policy gradient TD3 agent, which determines the optimization strategy for boiler combustion based on the current state of the boiler, adjusts the values ​​of input variables according to the optimization strategy, and outputs the next state of the boiler through a prediction model; Determine the network parameters of the Actor network and the Critic network in the TD3 agent according to the current state, the optimization strategy and the next state, determine the loss value based on the network parameters, train the Actor network and the Critic network with the goal of minimizing the loss value, and obtain a combustion control optimization model; The state parameters of the current state of the boiler are input into the combustion control optimization model, and the combustion control optimization strategy of the boiler is determined by the combustion control optimization model, and then the input variables are adjusted according to the combustion control optimization strategy; and the combustion state of the boiler is controlled according to the adjusted input variables.

2. A boiler combustion optimization method according to claim 1, characterized in that: The relevant variables are screened by feature importance to obtain a preset number of strongly relevant variables as input variables, specifically including: The importance of the features of the relevant variables is evaluated by the random forest method RF, and the formula is: in, are the prediction errors of the data of the relevant variables before and after the interference on the decision tree b, and B is the number of decision trees. The sum of the importance of all input variables is 1, that is, Indicates the number of related variables; According to the evaluation results, the related variables are sorted by feature importance, and a preset number of related variables at the top of the sorting are determined as strongly related variables; The strongly correlated variable is determined as an input variable.

3. A boiler combustion optimization method according to claim 1, characterized in that: The channel selection convolution layer includes an expected channel damage matrix, through which the importance of the input channels in the channel selection convolution layer is identified, a preset number C of high-importance channels are screened, and the hyperparameters of the remaining low-importance channels are released.

4. A boiler combustion optimization method according to claim 3, characterized in that: The channel selection convolution layer is also used to perform channel unblocking, channel reallocation and spatial shifting; through the channel unblocking, related variables with low importance are prevented from being used in subsequent model calculations and the parameters of these channels are released; through the channel reallocation, the blocked channels are replaced with preset C high-importance channels; through the spatial shifting, the feature loss caused by too many similar feature channels is offset.

5. A boiler combustion optimization method according to claim 1, characterized in that: The CS-CNN is trained to obtain a boiler state prediction model including: Obtaining a boiler combustion training sample; the training sample includes a sample input variable and its corresponding real state; Input the sample input variable into CS-CNN to obtain a predicted state; The CS-CNN is trained with the goal of minimizing the deviation between the true state and the predicted state to obtain the boiler state prediction model.

6. A boiler combustion optimization method according to claim 1, characterized in that: Determining network parameters of the Actor network in TD3 according to the current state and the optimization strategy, and determining the loss value based on the network parameters includes: The Q values ​​of the two Critic networks are determined according to the current state and the optimization strategy, and the loss value of the Actor network is determined based on the smaller value of the two Q values. The Q values ​​are Q1(s, a) and Q2(s, a), respectively. The specific formula is: loss actor =-min(Q 1,2 (s,a)); Among them, s is the current state and a is the optimization strategy.

7. A boiler combustion optimization method according to claim 1, characterized in that: Determining the network parameters of the Critic network in TD3 according to the current state and the optimization strategy, and determining the loss value based on the network parameters includes: The loss functions of the two online updated Critic networks use the temporal difference error algorithm to calculate the mean square error between Q1(s,a) and r+γ·min(Q1′,Q2′), and between Q2(s,a) and r+γ·min(Q1′,Q2′) as the loss values. The specific formulas are: loss critic1 =MSE[Q1(s,a),r+γ·min(Q1′,Q2′)]; loss critic2 =MSE[Q2(s,a),r+γ·min(Q1′,Q2′)] Among them, Q1′ and Q2′ are the Q values ​​of the two delayed update target Critic networks, r is the immediate reward obtained after adjusting the input variables according to the optimization strategy in the current state, and γ is the discount factor that balances the immediate reward and future rewards.

8. A boiler combustion optimization device, characterized in that: The device comprises: An acquisition module is used to acquire relevant variables that affect the state of the boiler, perform random forest RF feature importance evaluation on the relevant variables, and determine input variables; A construction module is provided for constructing a convolutional neural network (CNN) model, replacing the conventional convolutional layer in the CNN model with a channel selection convolutional layer to obtain a channel selection convolutional neural network (CS-CNN); the CS-CNN includes an input layer, a channel selection convolutional layer, a fully connected layer, and an output layer; the CS-CNN is trained to obtain a boiler state prediction model; a dual-delay deep deterministic policy gradient (TD3) agent is constructed, including the design and training process of an Actor network and a Critic network, and the TD3 agent interacts with the boiler state prediction model, including: the agent adjusts the input variables according to the current state of the boiler to obtain an optimization strategy for boiler combustion, adjusts the state variables according to the optimization strategy, and outputs the next state of the boiler through a prediction model; the loss values ​​of the Actor network and the Critic network in TD3 are determined according to the current state and the optimization strategy, and TD3 is trained with the goal of minimizing the loss value to obtain a combustion control optimization model; The optimization module is used to input the current state of the boiler into the combustion control optimization model, determine the combustion control optimization strategy of the boiler through the combustion control optimization model, and then adjust the input variables according to the combustion control optimization strategy.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the program.