Garbage incinerator emission automatic control method based on reinforcement learning

By applying an automatic control method based on reinforcement learning in waste incinerators and adjusting control parameters using DDPG algorithm, the problems of harmful gas pollution and high energy consumption during waste incineration are solved, and automated control and energy consumption optimization are achieved.

CN119914883AActive Publication Date: 2025-05-02CHINA CONSTRUCTION THIRD ENGINEERING BUREAU WUCHUANG YUNWEI TECHNOLOGY (WUHAN) CO LTD +2

Patent Information

Application Number
CN202411923894.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-02
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

The harmful gases generated during waste incineration cause serious pollution to the environment. Traditional control methods cannot achieve automated and intelligent control, and flue gas purification and heat recovery equipment consumes a large amount of energy.

Method used

The automatic control method of waste incinerator emissions based on reinforcement learning is adopted, and the control model is constructed through the deep deterministic strategy gradient (DDPG) algorithm, and the control parameters during the incineration process are automatically adjusted to achieve flue gas emissions comply with environmental protection standards and reduce energy consumption.

Benefits of technology

Automatic and intelligent control of waste incineration flue gas emissions has been achieved, ensuring that emissions meet environmental protection standards, and at the same time reducing energy consumption for flue gas treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_4
    Figure SMS_4
  • Figure SMS_5
    Figure SMS_5
  • Figure SMS_7
    Figure SMS_7
Patent Text Reader

Abstract

The invention discloses a garbage incinerator emission automatic control method based on reinforcement learning. The method comprises the following steps that all control parameters of a garbage incinerator are set and represented by a control parameter vector x = [v1, v2, d1, d2, vm, h, vl, m]; setting a toxic emission concentration vector W = [w1, w2,..., wn] and an energy consumption vector e = [e1, e2,..., ej] for representation, finding a group of control parameter vectors x, and enabling the toxic emission concentration vector W and the energy consumption vector e to be minimum: constructing an Ev model, an Ac model and a Q model by using a depth deterministic strategy gradient; historical monitoring data in an incinerator DCS (Distributed Control System) are used for training an Ev model, then the trained Ev model is used for calculating corresponding data after actions to train an Ac model and a Q model, the trained Ac model can predict an optimal control parameter vector x according to the current state S of the incinerator, and automatic control of waste incineration flue gas emission is realized. The control parameters in the incineration process are automatically adjusted by using the reinforcement learning algorithm and combining with the historical monitoring data of the incinerator, so that the energy consumption is reduced while the flue gas emission reaches the environmental protection standard.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a waste incineration flue gas emission control technology, and specifically to a waste incinerator emission automatic control method based on reinforcement learning, which is used to ensure that the emitted flue gas complies with environmental protection regulations while reducing the energy consumption of flue gas treatment. Background Art

[0002] Waste incineration can not only effectively reduce the volume of waste, but also convert heat energy into electricity or other energy resources, with good environmental protection and economic benefits. However, the impact of flue gas and harmful gases generated during waste incineration on the environment cannot be ignored. During the waste incineration process, if the flue gas containing harmful gases (such as nitrogen oxides, sulfur dioxide, carbon monoxide, volatile organic compounds, etc.) is not effectively controlled, it will seriously pollute the air, endanger human health, and affect the ecological environment. In order to ensure that the emission of harmful substances during the incineration process meets the standards of environmental protection regulations, the emission control system of the waste incinerator needs to be precisely controlled. According to the "Air Pollutant Emission Standard" (GB 13271-2014) and other relevant environmental protection regulations, the concentration of pollutants in waste incineration flue gas needs to be strictly controlled, especially the emission of harmful gases such as sulfur dioxide, nitrogen oxides, and carbon monoxide. These standards require continuous and real-time monitoring of gas emissions during waste incineration, and dynamically adjust the combustion conditions in the furnace based on the data to ensure that the emission concentration does not exceed the standard.

[0003] Existing waste incineration flue gas emission control technologies usually rely on traditional control methods, such as PID control, fuzzy control, etc. Although these methods can stabilize the combustion process in the furnace to a certain extent, due to the involvement of multiple variables in the waste incineration process, such as the calorific value, humidity, and composition of the waste, traditional control methods are difficult to cope with the complex and changeable incineration process, which can easily lead to problems such as delayed system response, low combustion efficiency, and excessive emission concentration. In addition, traditional manual adjustment and static optimization methods are difficult to cope with complex operating conditions. In the actual incineration process, factors such as the composition and characteristics of the garbage, climatic conditions, and furnace temperature will change. The impact of these factors on the incineration process and its emissions is dynamic and complex. For the control of these uncertain factors, the effect of traditional methods is often not ideal. Therefore, how to achieve automated and intelligent control in the incineration process is the key to improving the efficiency of the waste incineration emission control system, reducing energy consumption, and ensuring emission compliance.

[0004] In addition, the treatment and purification of flue gas during waste incineration consumes a lot of energy, including air flow regulation, temperature control, reaction time control, etc. Although the heat energy generated by incineration can be used for self-power supply during waste incineration, the energy consumption of flue gas purification equipment and heat recovery equipment is still relatively large. Therefore, how to reduce the energy consumption of flue gas treatment and optimize the flue gas purification process has become a difficult problem that needs to be solved in waste incineration technology. Summary of the invention

[0005] The purpose of the present invention is to solve the problem that harmful gases generated during the incineration of garbage cause serious pollution to the environment, traditional control methods cannot achieve automated and intelligent control of the garbage incineration process, and flue gas purification and heat recovery equipment consume a lot of energy during the treatment process. A garbage incinerator emission automatic control method based on reinforcement learning is provided. The method provided by the present invention realizes automatic control of garbage incineration through reinforcement learning. Reinforcement learning (RL) is a paradigm of machine learning that learns how to make decisions through interaction with the environment; its main difference from supervised learning and unsupervised learning is that the learning process in reinforcement learning is completed through the interaction between an agent and an environment; the agent performs actions in the environment and adjusts strategies according to the rewards fed back by the environment to ultimately achieve a predetermined goal.

[0006] In a waste incinerator, waste goes through three stages: drying, burning, and combustion:

[0007] 1) The garbage is first transported to the drying grate through the feeder. Under the action of primary air and furnace temperature, the moisture in the garbage begins to evaporate;

[0008] 2) After the drying process is completed, the garbage is sent to the combustion grate, and the primary air volume is used to make the garbage reach the ignition point and start to burn, and volatile matter is precipitated. The high-temperature flue gas generated forms a highly turbulent flow under the action of oxygen provided by the secondary air entering from the top of the furnace, which fully decomposes the harmful gases in the flue gas;

[0009] 3) Finally, the garbage that has passed the combustion stage is sent to the incineration grate to ensure that the remaining unburned garbage is completely burned out, and the residue is sent to the residue pool by the slag conveyor.

[0010] In order to optimize the above incineration process, the following operating parameters must be controlled: primary air volume v1, secondary air volume v2, primary air direction d1, secondary air direction d2, row movement speed v m , grate movement mode h, including forward, backward and pause; garbage feeding rate v l ; Adsorbent dosage m;

[0011] The control goal is to make the garbage discharge meet the standards while minimizing the energy consumption of the garbage incineration process. The garbage emission index is mainly characterized by the concentration of toxic substances such as CO, NOx, and SOx, which is recorded as a toxic emission concentration vector W = [w1, w2, ..., w n ], where w n represents the concentration of the nth toxic emission, n is a natural number greater than 0, and the control energy consumption is also represented by an energy consumption vector e = [e1, e2, ..., e j ] indicates that e j represents the electric power or fuel flow under the jth control parameter vector; at this time, the automatic control problem of waste incineration flue gas emissions is converted into a minimization problem, that is, to find a set of control parameter vectors x that minimize the concentration vector W of toxic emissions and the energy consumption vector e:

[0012] x is the action;

[0013] The above minimization problem is a highly nonlinear non-convex optimization problem. The method of the present invention designs a waste incineration flue gas control algorithm based on a deep reinforcement learning architecture, and specifically uses a deep deterministic policy gradient (DDPG) to construct a control model.

[0014] The technical solution adopted by the present invention is:

[0015] A method for automatic emission control of a garbage incinerator based on reinforcement learning, characterized by comprising the following steps:

[0016] Set the control parameters of the garbage incinerator: primary air volume v1, secondary air volume v2, primary air direction d1, secondary air direction d2, row movement speed v m , grate movement mode h, including forward, backward and pause; garbage feeding rate v l ; Adsorbent dosage m; The above waste incinerator control parameters are represented by a control parameter vector x = [v1, v2, d1, d2, v m ,h,v l , m] to indicate;

[0017] The goal of control is to make the garbage discharge meet the standards and minimize the energy consumption of the garbage incineration process. The garbage discharge index is characterized by the concentration of toxic emissions, which is recorded as a toxic emission concentration vector W = [w1, w2, ..., w n ], where w n represents the concentration of the nth toxic emission, n is a natural number greater than 0, and the energy consumption is controlled by an energy consumption vector e = [e1, e2, ..., e j ] indicates that ej represents the electric power or fuel flow under the jth control parameter vector; at this time, the automatic control problem of waste incineration flue gas emissions is converted into a minimization problem, that is, to find a set of control parameter vectors x that minimize the concentration vector W of toxic emissions and the energy consumption vector e:

[0018] x is the action;

[0019] Construct control model: Use deep deterministic policy gradient to construct the simulated environment network Ev model, action network Ac model and criticism network Q model; use the historical monitoring data in the incinerator DCS to train the Ev model, and then use the trained Ev model to calculate the action and then use the corresponding data to train the Ac model and Q model. The trained Ac model can predict the optimal control parameter vector x according to the current state S of the incinerator, thereby realizing automatic control of flue gas emissions from garbage incineration.

[0020] First, define the environmental state S = [W, e], where W is the concentration vector of toxic emissions, e is the energy consumption vector, and action x is the control parameter vector x; construct three neural network models: the simulation environment network Ev model, the action network Ac model, and the criticism network Q model; the input and output of the three are as follows:

[0021] 1) Input current environment state S of the simulation environment network Ev model t and the current action x t ; Output the environmental state S after the action t+1 ; t is the current time;

[0022] 2) Input of the action network Ac model: current environment state S t ; Output current action x t ;

[0023] 3) Critique the input environment state S and action x of the network Q model; output the value of the action Q(S, x);

[0024] The above three network models all adopt the structure of multi-layer perceptron (MLP). The training of each model is carried out twice. The first training is to train the simulated environment network Ev model, and the second training is to simultaneously train the action network Ac model and the criticism network Q model.

[0025] The training steps of the simulated environment network Ev model are as follows:

[0026] 1) Construct a simulated environment dataset, which contains pairs of current environment states S t , current action x t , and the state S after the action t+1The data set is constructed using the historical monitoring data in the incinerator DCS. The time series of the historical monitoring data in the DCS is differentiated to obtain the above sample pairs ([S t , x t ]; S t+1 ); the historical monitoring data in the incinerator DCS are primary air volume v1, secondary air volume v2; primary air direction d1, secondary air direction d2, row movement speed v m , grate movement mode h, garbage feeding rate v l and adsorbent dosage m; concentrations of toxic emissions w1, w2, ..., w n , where w n Represents the concentration of the nth toxic emission, where n is a natural number greater than 0.

[0027] 2) Divide the sample pairs in the data set into training set and validation set in a ratio of 8:2;

[0028] 3) Randomly select a sample pair from the training set ([S t , x t ]; S t+1 );

[0029] 4) Use the simulated environment network Ev model to calculate the state prediction value after the action

[0030]

[0031] 5) Calculate the loss function L e :

[0032]

[0033] 6) L e Back propagation, optimizing the simulation environment network Ev model;

[0034] 7) Randomly select a sample pair from the verification set to perform error verification on the optimized simulation environment network Ev model, and save the parameters corresponding to the sample pair of the input Ev model in this round;

[0035] 8) Repeat steps 3) to 7) for 10,000 rounds, and finally take the parameters corresponding to a set of sample pairs of the input Ev model with the lowest error on the validation set as the training results, and obtain the trained Ev model.

[0036] After the Ev model training is completed, the parameters of the Ev model are frozen, and then the action network Ac model and the criticism network Q model are trained.

[0037] The training steps of the action network Ac model and the criticism network Q model are as follows:

[0038] 1) Initialize playback buffer D;

[0039] 2) Randomly sample a current environment state S from the data set t =[W t , e t ];

[0040] 3)Ac model takes the current environment state S t As input, calculate the current action x t :

[0041] x t =Ac(S t );

[0042] 4) Use the trained Ev model to calculate the environmental state S after the action t+1 :

[0043] [W t , e t ]=S t+1 =Ev(S t , x t );

[0044] 5) Calculate the new action x t+1 :

[0045] x t+1 =Ac(S t+1 );

[0046] 6) Calculate the reward function r t :

[0047] r t =-|W t |-|e t |;

[0048] 7) (S t , x t , r t , S t+1 ) is stored in the playback buffer D, which is used to store the experience data (S t , x t , r t , S t+1 , xt +1 );

[0049] 8) Randomly extract an experience data (S) from the playback buffer D t , x t , r t , S t+1 , x t+1 );

[0050] 9) Use the Q model to calculate the loss function L q :

[0051] y=r t +0.9*Q(S t+1 , x t+1 ), y is an intermediate variable, r t is the reward function;

[0052] L q =yQ(S t , x t );

[0053] 10) For L q Back propagation, optimizing the Q model;

[0054] 11) Using the Q model to calculate the policy gradient Where θ represents the parameters of the Ac model;

[0055]

[0056] 12) Use gradient ascent to update the parameters of the Ac model, and the learning rate α is 0.001;

[0057]

[0058] 13) Repeat steps 2) to 12) until the average change of θ is within 10 -6 If the number of repetitions exceeds 100,000, the training is terminated and the trained Ac model and Q model are obtained.

[0059] After the training of the Ac model and the Q model is completed, the trained Ac model can predict the optimal control parameter vector x according to the current state S of the incinerator, thereby realizing automatic control of flue gas emissions from garbage incineration.

[0060] The present invention uses a reinforcement learning algorithm, combined with historical monitoring data and flame data characteristics, to automatically adjust the control parameters in the incineration process, so as to achieve flue gas emissions that meet environmental protection standards while reducing energy consumption. DETAILED DESCRIPTION

[0061] The present invention will be further described in detail below in conjunction with specific embodiments to facilitate a clear understanding of the present invention, but they do not limit the present invention.

[0062] In a waste incinerator, waste goes through three stages: drying, burning, and combustion:

[0063] 1) The garbage is first transported to the drying grate through the feeder. Under the action of primary air and furnace temperature, the moisture in the garbage begins to evaporate;

[0064] 2) After the drying process is completed, the garbage is sent to the combustion grate, and the primary air volume is used to make the garbage reach the ignition point and start to burn, and volatile matter is precipitated. The high-temperature flue gas generated forms a highly turbulent flow under the action of oxygen provided by the secondary air entering from the top of the furnace, which fully decomposes the harmful gases in the flue gas;

[0065] 3) Finally, the garbage that has passed the combustion stage is sent to the incineration grate to ensure that the remaining unburned garbage is completely burned out, and the residue is sent to the residue pool by the slag conveyor.

[0066] In order to optimize the above incineration process, the following operating parameters must be controlled: primary air volume v1, secondary air volume v2, primary air direction d1, secondary air direction d2, row movement speed v m , grate movement mode h, including forward, backward and pause; garbage feeding rate v l ; Adsorbent dosage m;

[0067] The goal of control is to make the garbage discharge meet the standards while minimizing the energy consumption of the garbage incineration process. The garbage emission index is mainly characterized by the concentration of toxic substances such as CO, NOx, and SOx, which is recorded as a toxic emission concentration vector W = [w1, w2, ..., w n ], where w n represents the concentration of the nth toxic emission, n=9, w1 is the concentration of HCl emission, w2 is the concentration of HF emission, w3 is the concentration of SO2 emission, w4 is the concentration of NO x The emission concentration is w5, CO emission concentration, w6, Hg emission concentration, w7, Cd+Ti emission concentration, w8, Pb+Cr and other heavy metal emission concentrations, and w9, dioxin emission concentration. The energy consumption is also controlled by an energy consumption vector e=[e1, e2, ..., e j ] indicates that e j represents the electric power or fuel flow under the jth control parameter vector; at this time, the automatic control problem of waste incineration flue gas emissions is converted into a minimization problem, that is, to find a set of control parameter vectors x that minimize the concentration vector W of toxic emissions and the energy consumption vector e:

[0068] x is the action;

[0069] The above minimization problem is a highly nonlinear non-convex optimization problem. The method of the present invention designs a waste incineration flue gas control algorithm based on a deep reinforcement learning architecture, and specifically uses a deep deterministic policy gradient (DDPG) to construct a control model.

[0070] The technical solution adopted by the present invention is:

[0071] A method for automatic emission control of a garbage incinerator based on reinforcement learning, characterized by comprising the following steps:

[0072] Set the control parameters of the garbage incinerator: primary air volume v1, secondary air volume v2, primary air direction d1, secondary air direction d2, row movement speed v m , grate movement mode h, including forward, backward and pause; garbage feeding rate v l ; Adsorbent dosage m; The above waste incinerator control parameters are represented by a control parameter vector x = [v1, v2, d1, d2, v m ,h,v l , m] to indicate;

[0073] The goal of control is to make the garbage discharge meet the standards and minimize the energy consumption of the garbage incineration process. The garbage discharge index is characterized by the concentration of toxic emissions, which is recorded as a toxic emission concentration vector W = [w1, w2, ..., w n ], where w n represents the concentration of the nth toxic emission, n is 9, w1 is the concentration of HCl emission, w2 is the concentration of HF emission, w3 is the concentration of SO2 emission, and w4 is the concentration of NO x The emission concentration is: w5 is CO emission concentration, w6 is Hg emission concentration, w7 is Cd+Ti emission concentration, w8 is Pb+Cr and other heavy metal emission concentrations, w9 is dioxin emission concentration, and energy consumption is controlled by an energy consumption vector e=[e1, e2, ..., e j ] indicates that e j represents the electric power or fuel flow under the jth control parameter vector; at this time, the automatic control problem of waste incineration flue gas emissions is converted into a minimization problem, that is, to find a set of control parameter vectors x that minimize the concentration vector W of toxic emissions and the energy consumption vector e:

[0074] x is the action;

[0075] Construct control model: Use deep deterministic policy gradient to construct the simulated environment network Ev model, action network Ac model and criticism network Q model; use the historical monitoring data in the incinerator DCS to train the Ev model, and then use the trained Ev model to calculate the action and then use the corresponding data to train the Ac model and Q model. The trained Ac model can predict the optimal control parameter vector x according to the current state S of the incinerator, thereby realizing automatic control of flue gas emissions from garbage incineration.

[0076] First, define the environmental state S = [W, e], where W is the concentration vector of toxic emissions, e is the energy consumption vector, and action x is the control parameter vector x; construct three neural network models: the simulation environment network Ev model, the action network Ac model, and the criticism network Q model; the input and output of the three are as follows:

[0077] 1) Input current environment state S of the simulation environment network Ev model t and the current action x t ; Output the environmental state S after the action t+1 ; t is the current time;

[0078] 2) Input of the action network Ac model: current environment state S t ; Output current action x t ;

[0079] 3) Critique the input environment state S and action x of the network Q model; output the value of the action Q(S, x);

[0080] The above three network models all adopt the structure of multi-layer perceptron (MLP). The training of each model is carried out twice. The first training is to train the simulated environment network Ev model, and the second training is to simultaneously train the action network Ac model and the criticism network Q model.

[0081] The training steps of the simulated environment network Ev model are as follows:

[0082] 1) Construct a simulated environment dataset, which contains pairs of current environment states S t , current action x t , and the state S after the action t+1 The data set is constructed using the historical monitoring data in the incinerator DCS. The time series of the historical monitoring data in the DCS is differentiated to obtain the above sample pairs ([S t , x t ]; S t+1 ); the historical monitoring data in the incinerator DCS are primary air volume v1, secondary air volume v2; primary air direction d1, secondary air direction d2, row movement speed v m , grate movement mode h, garbage feeding rate v l and adsorbent dosage m; HCl emission concentration w1, HF emission concentration w2, SO2 emission concentration w3, NO x Emission concentration w4, CO emission concentration w5, Hg emission concentration w6, Cd+Ti emission concentration w7, Pb+Cr and other heavy metal emission concentration w8, dioxin emission concentration w9;

[0083] 2) Divide the sample pairs in the data set into training set and validation set according to the ratio of 8:2;

[0084] 3) Randomly select a sample pair from the training set ([S t , x t ]; S t+1 );

[0085] 4) Use the simulated environment network Ev model to calculate the state prediction value after the action

[0086] 5) Calculate the loss function L e :

[0087]

[0088] 6) L e Back propagation, optimizing the simulation environment network Ev model;

[0089] 7) Randomly select a sample pair from the verification set to perform error verification on the optimized simulation environment network Ev model, and save the parameters corresponding to the sample pair of the input Ev model in this round;

[0090] 8) Repeat steps 3) to 7) for 10,000 rounds, and finally take the parameters corresponding to a set of sample pairs of the input Ev model with the lowest error on the validation set as the training results, and obtain the trained Ev model.

[0091] After the Ev model training is completed, the parameters of the Ev model are frozen, and then the action network Ac model and the criticism network Q model are trained.

[0092] The training steps of the action network Ac model and the criticism network Q model are as follows:

[0093] 1) Initialize playback buffer D;

[0094] 2) Randomly sample a current environment state S from the data set t =[W t , e t ];

[0095] 3)Ac model takes the current environment state S t As input, calculate the current action x t :

[0096] x t =Ac(S t );

[0097] 4) Use the trained Ev model to calculate the environmental state S after the action t+1 :

[0098] [W t , e t]=S t+1 =Ev(S t , x t );

[0099] 5) Calculate the new action x t+1 :

[0100] x t+1 =Ac(S t+1 );

[0101] 6) Calculate the reward function r t :

[0102] r t =-|W t |-|e t |;

[0103] 7) (S t , x t , r t , S t+1 ) is stored in the playback buffer D, which is used to store the experience data (S t , x t , r t , S t+1 , x t+1 );

[0104] 8) Randomly extract an experience data (S) from the playback buffer D t , x t , r t , S t+1 , x t+1 );

[0105] 9) Use the Q model to calculate the loss function L q :

[0106] y=r t +0.9*Q(S t+1 , x t+1 ), y is an intermediate variable, r t is the reward function;

[0107] L q =yQ(S t , x t );

[0108] 10) For L q Back propagation, optimizing the Q model;

[0109] 11) Using the Q model to calculate the policy gradient Where θ represents the parameters of the Ac model;

[0110]

[0111] 12) Use gradient ascent to update the parameters of the Ac model, and the learning rate α is 0.001;

[0112]

[0113] 13) Repeat steps 2) to 12) until the average change of θ is within 10 -6 If the number of repetitions exceeds 100,000, the training is terminated and the trained Ac model and Q model are obtained.

[0114] After the training of the Ac model and the Q model is completed, the trained Ac model can predict the optimal control parameter vector x according to the current state S of the incinerator, thereby realizing automatic control of flue gas emissions from garbage incineration.

Claims

1. A method for automatic emission control of a waste incinerator based on reinforcement learning, characterized in that The following steps are involved: Set the control parameters of the garbage incinerator: primary air volume v1, secondary air volume v2, primary air direction d1, secondary air direction d2, row movement speed v m , grate movement mode h, including forward, backward and pause; garbage feeding rate v l ; Adsorbent dosage m; The above waste incinerator control parameters are represented by a control parameter vector x = [v1, v2, d1, d2, v m ,h,v l , m] to indicate; Control target: The garbage discharge index is characterized by the concentration of toxic emissions, which is recorded as a toxic emission concentration vector W = [w1, w2, ..., w n ], where w n represents the concentration of the nth toxic emission, n is a natural number greater than 0, and the energy consumption is controlled by an energy consumption vector e = [e1, e2, ..., e j ] indicates that e j represents the electric power or fuel flow under the jth control parameter vector; at this time, the automatic control problem of waste incineration flue gas emissions is converted into a minimization problem, that is, to find a set of control parameter vectors x that minimize the concentration vector W of toxic emissions and the energy consumption vector e: x is the action; Construct control model: Use deep deterministic policy gradient to construct the simulated environment network Ev model, action network Ac model and criticism network Q model; use the historical monitoring data in the incinerator DCS to train the Ev model, and then use the trained Ev model to calculate the action and then use the corresponding data to train the Ac model and Q model. The trained Ac model can predict the optimal control parameter vector x according to the current state S of the incinerator, thereby realizing automatic control of flue gas emissions from garbage incineration.

2. According to claim 1, a method for automatic emission control of a garbage incinerator based on reinforcement learning is characterized in that The specific steps to build the control model are as follows: First, define the environmental state S = [W, e], where W is the concentration vector of toxic emissions, e is the energy consumption vector, and action x is the control parameter vector x; construct three neural network models: the simulation environment network Ev model, the action network Ac model, and the criticism network Q model; the input and output of the three are as follows: 1) Input current environment state S of the simulation environment network Ev model t and the current action x t ; Output the environmental state S after the action t+1 ; t is the current time; 2) Input of the action network Ac model: current environment state S t ; Output current action x t ; 3) Critique the input environment state S and action x of the network Q model; output the value of the action Q(S, x); The above three network models all adopt the structure of multi-layer perceptron (MLP). The training of each model is carried out twice. The first training is to train the simulated environment network Ev model, and the second training is to simultaneously train the action network Ac model and the criticism network Q model.

3. According to claim 2, a method for automatic emission control of a garbage incinerator based on reinforcement learning is characterized in that The training steps of the simulated environment network Ev model are as follows: 1) Construct a simulated environment dataset, which contains pairs of current environment states S t , current action x t , and the state S after the action t+1 The data set is constructed using the historical monitoring data in the incinerator DCS. The time series of the historical monitoring data in the DCS is differentiated to obtain the above sample pairs ([S t , x t ]; S t+1 ); the historical monitoring data in the incinerator DCS are primary air volume v1, secondary air volume v2; primary air direction d1, secondary air direction d2, row movement speed v m , grate movement mode h, garbage feeding rate v l and adsorbent dosage m; concentrations of toxic emissions w1, w2, ..., w n , w n Represents the concentration of the nth toxic emission, where n is a natural number greater than 0; 2) Divide the sample pairs in the data set into training set and validation set in a ratio of 8:2; 3) Randomly select a sample pair from the training set ([S t , x t ]; S t+1 ); 4) Use the simulated environment network Ev model to calculate the state prediction value after the action 5) Calculate the loss function L e : 6) L e Back propagation, optimizing the simulation environment network Ev model; 7) Randomly select a sample pair from the verification set to perform error verification on the optimized simulation environment network Ev model, and save the parameters corresponding to the sample pair of the input Ev model in this round; 8) Repeat steps 3) to 7) for 10,000 rounds, and finally take the parameters corresponding to a set of sample pairs of the input Ev model with the lowest error on the validation set as the training results, and obtain the trained Ev model.

4. The method for automatic emission control of a garbage incinerator based on reinforcement learning according to claim 3 is characterized in that The training steps of the action network Ac model and the criticism network Q model are as follows: 1) Initialize playback buffer D; 2) Randomly sample a current environment state S from the data set t =[W t , e t ]; 3)Ac model takes the current environment state S t As input, calculate the current action x t : x t =Ac(S t ); 4) Use the trained Ev model to calculate the environmental state S after the action t+1 : [W t ,e t ]=S t+1 =Ev(S t ,x t ); 5) Calculate the new action x t+1 : x t+1 =Ac(S t+1 ); 6) Calculate the reward function r t : r t N-|W t |-|e t |6 7) (S t , x t , r t , S t+1 ) is stored in the playback buffer D, which is used to store the experience data (S t , x t , r t , S t+1 , x t+1 ); 8) Randomly extract an experience data (S) from the playback buffer D t , x t , r t , S t+1 , x t+1 ); 9) Use the Q model to calculate the loss function L q : y=r t +0.9*Q(S t+1 , x t+1 ), y is an intermediate variable, r t is the reward function; L q =y-Q(S t ,x t ); 10) For L q Back propagation, optimizing the Q model; 11) Using the Q model to calculate the policy gradient Where θ represents the parameters of the Ac model; 12) Use gradient ascent to update the parameters of the Ac model, and the learning rate α is 0.001; 13) Repeat steps 2) to 12) until the average change of θ is within 10 -6 If the number of repetitions exceeds 100,000, the training is terminated and the trained Ac model and Q model are obtained; the trained Ac model can predict the optimal control parameter vector x according to the current state S of the incinerator to realize the automatic control of the flue gas emission from the garbage incineration.

Citation Information

Patent Citations

  • Artificial intelligence control method for waste incineration process

    CN113405106A

  • Accurate coal blending method and system based on deep reinforcement learning

    CN118333229A

  • Method, apparatus and electronic device for constructing reinforcement learning model and medium

    US20210216686A1

Cited By

  • Garbage incineration power generation performance prediction and optimization control method based on big data

    CN121274204A