Collaborative control method for incineration pollution based on multi-objective reinforcement learning
By constructing a multi-objective reinforcement learning system, combined with fuzzy logic and multimodal sensors, the control variables in the waste incineration process are optimized, solving the problems of low pollutant emission and fly ash metal recovery efficiency in traditional waste incineration, and achieving synergistic optimization of pollutant emission reduction, resource recovery cost reduction and resource utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING INST OF ENVIRONMENTAL SCI MINIST OF ECOLOGY & ENVIRONMENT OF THE PEOPLES REPUBLIC OF CHINA
- Filing Date
- 2025-06-17
- Publication Date
- 2026-05-29
AI Technical Summary
In traditional waste incineration, pollutant emission control is difficult, valuable metal recovery efficiency in fly ash is low and resource utilization costs are high, and there is a lack of accurate data collection and intelligent decision support systems, making it difficult to achieve real-time dynamic adjustment to complex and ever-changing incineration conditions.
A dynamic feature matrix of incineration operation and a multi-source data acquisition system are constructed. A multi-objective reinforcement learning joint optimization model is adopted, and dynamic weight allocation and collaborative optimization control are realized by combining fuzzy logic. Data is collected in real time through multi-modal sensors. Objective functions of pollutant emissions, fly ash metal recovery rate and resource recovery cost are defined. The control variables are optimized by using a deep deterministic policy gradient algorithm and an Actor-Critic network structure.
It effectively inhibits the formation of dioxins and particulate matter, improves the metal recovery rate of fly ash, reduces resource utilization costs, improves the overall efficiency of waste incineration, and promotes the effective utilization of resources and environmental protection.
Smart Images

Figure CN120650717B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of waste incineration technology, and in particular to a collaborative control method for incineration pollution based on multi-objective reinforcement learning. Background Technology
[0002] With the acceleration of urbanization and the improvement of residents' living standards, the amount of urban solid waste is increasing day by day. Waste incineration, as an effective waste treatment method, is widely used because it can significantly reduce waste volume, lower environmental impact, and recover energy. However, traditional waste incineration processes face many challenges, including but not limited to difficulties in controlling pollutant emissions, low efficiency in recovering valuable metals from fly ash, and high resource recovery costs.
[0003] During waste incineration, if combustion conditions are not ideal, large amounts of harmful substances may be generated, such as persistent organic pollutants like dioxins and polycyclic aromatic hydrocarbons (PAHs), as well as air pollutants like particulate matter (PM) and nitrogen oxides (NOx). These pollutants not only cause serious environmental pollution but may also pose a threat to human health. Therefore, effectively controlling pollutant emissions has become one of the urgent problems to be solved by waste incineration plants.
[0004] Fly ash from waste incineration contains certain amounts of heavy metals, such as lead (Pb), mercury (Hg), and cadmium (Cd), which have high economic value but are also potential sources of environmental pollution. Currently, existing fly ash treatment methods mainly include solidification / stabilization followed by landfill or direct use as building materials, but neither of these methods fully utilizes the valuable metal resources in the fly ash. Furthermore, the lack of efficient recycling technologies and reasonable resource allocation strategies leads to resource waste and increased treatment costs.
[0005] To address these challenges, some studies have attempted to optimize incineration conditions by improving combustion process parameters and applying advanced purification technologies. However, most existing solutions tend to focus on optimizing a single objective, such as reducing the emission of a specific pollutant or increasing the recovery rate of a certain type of metal, while neglecting the improvement of the overall system performance. Furthermore, the lack of precise data acquisition methods and intelligent decision support systems makes it difficult to achieve real-time dynamic adjustments to complex and changing incineration conditions, thus limiting the optimization effectiveness.
[0006] In light of the above background, this invention proposes an optimized control system and method for incineration operation based on multi-source data acquisition and reinforcement learning. This system constructs a dynamic feature matrix of incineration operation conditions and a multi-source data acquisition system, establishes a multi-objective reinforcement learning joint optimization model, and combines fuzzy logic to achieve dynamic weight allocation and collaborative optimization control. The aim is to comprehensively improve the overall efficiency of waste incineration treatment, achieving the goals of minimizing pollutant emissions, maximizing fly ash metal recovery rate, and minimizing resource recovery costs. This innovation not only contributes to environmental protection but also promotes the effective utilization of resources, providing strong technical support for the sustainable development of the waste incineration industry. Summary of the Invention
[0007] To address the above problems, this invention proposes a collaborative control method for incineration pollution based on multi-objective reinforcement learning. The specific steps are as follows:
[0008] Step 1: Construct a dynamic feature matrix of incineration conditions and a multi-source data acquisition system. Deploy multi-modal sensor arrays in the furnace, flue, and fly ash treatment unit to collect parameters such as combustion temperature, flue gas residence time, denitrification agent injection amount, fly ash metal content, and dioxin / polycyclic aromatic hydrocarbon concentration in real time, and construct a high-dimensional dynamic feature matrix.
[0009] Step 2: Establish a multi-objective reinforcement learning joint optimization model. Based on the dynamic feature matrix, define three types of objective functions: minimizing pollutant emissions, maximizing fly ash metal recovery rate, and minimizing resource utilization cost. Use a deep deterministic policy gradient algorithm to construct a multi-objective reinforcement learning framework. Use combustion air ratio, secondary wind speed, and denitrification agent injection timing as control variables in the action space. A dual-network structure is used to realize policy exploration and value assessment.
[0010] Step 3: Implement dynamic weight allocation and multi-objective collaborative optimization control, construct a priority evaluation module based on fuzzy logic, and dynamically adjust the weight coefficients of the objective function according to real-time operating conditions.
[0011] As a further improvement to this invention, the construction of the dynamic feature matrix of incineration conditions and the multi-source data acquisition system in step 1 is represented as follows:
[0012] Step 1.1 Sensor selection and installation location;
[0013] The combustion temperature sensor uses a thermocouple array, deployed at 6 points at the top, middle, and bottom of the furnace, covering the main combustion zone, secondary combustion zone, and cooling zone; the pressure sensor uses a differential pressure sensor, installed at the furnace outlet and secondary air inlet, to monitor negative pressure fluctuations; the gas concentration sensor is installed at 3 points in the furnace outlet flue to monitor combustion products in real time.
[0014] A flue gas residence time sensor is installed in the middle section of the flue and calculates the residence time through optical path attenuation; a dioxin / polycyclic aromatic hydrocarbon online monitor is deployed at the front of the flue gas purification system to detect trace pollutants; an X-ray fluorescence analyzer is used to detect metal content and is integrated into the fly ash conveying pipeline to monitor the heavy metal leaching rate in real time; a particulate matter concentration sensor is installed at the dust collector outlet to capture particulate matter escaping.
[0015] Step 1.2 Data Acquisition;
[0016] Multi-source heterogeneous data is integrated into a standardized feature matrix to support the input requirements of subsequent multi-objective optimization models;
[0017] A sliding window with a window size of 10 seconds and a step size of 1 second is used to align asynchronous data with timestamps, eliminate sensor sampling delay differences, and perform One-Hot encoding on waste type classification variables to form a sensor feature matrix.
[0018] As a further improvement to the present invention, the multi-objective reinforcement learning joint optimization model established in step 2 is represented as follows:
[0019] Step 2.1 Define the objective function and associate it with the dynamic feature matrix;
[0020] Based on the high-dimensional dynamic feature matrix constructed in step 1, three types of objective functions are defined: minimizing pollutant emissions, maximizing fly ash metal recovery rate, and minimizing resource utilization cost, which provide the input basis for the dynamic weight adjustment module in step 3.
[0021] Step 2.1.1 Minimize pollutant emissions;
[0022] Using the real-time data from the flue area sensor in step 1, the calculation formula can be expressed as:
[0023]
[0024] in, and These are the weighting factors used to balance dioxins and particulate matter. This represents the rate of dioxin formation. This refers to the particulate matter concentration. The goal is to minimize pollutant emissions;
[0025] Physical constraints: furnace temperature Smoke residence time Monitored by the sensor in step 1;
[0026] Step 2.1.2 Maximize fly ash metal recovery rate;
[0027] The calculation formula can be expressed as:
[0028]
[0029] in, To maximize the recovery rate of fly ash metals These are the weighting coefficients. The heavy metal leaching rate is based on the X-ray fluorescence analyzer collected by the fly ash treatment unit in step 1. The maximum allowable heavy metal leaching rate; this expression is achieved by reducing... Improve recovery rate;
[0030] Step 2.1.3 Minimize resource recovery costs;
[0031] The calculation formula can be expressed as:
[0032]
[0033] in, To minimize resource recovery costs, These are the weighting coefficients. Energy consumption per unit of waste disposal;
[0034] Physical constraints: Denitrification agent injection volume ;
[0035] Step 2.2 Construction of Multi-Objective Joint Optimization Model
[0036] The three types of objective functions are integrated into a joint optimization objective. And a dynamic weight allocation mechanism is designed to connect the fuzzy logic module in step 3;
[0037] Step 2.2.1 Combining the objective functions:
[0038]
[0039] Weight initialization: , , ;
[0040] The reward function is defined as:
[0041]
[0042] The reinforcement learning framework is designed with an Actor-Critic dual-network structure:
[0043] Actor network: outputs continuous control actions, combustion air ratio. Secondary wind speed Denitrification agent injection timing ;
[0044] Critic Network: Evaluating Action Pairs The influence of gradient feedback is provided;
[0045] Experience replay and soft update mechanism:
[0046] Store historical experience To experience pool D;
[0047] Through soft update coefficients Stabilize the training process;
[0048] Step 2.3 Input mapping and state space design of dynamic feature matrix
[0049] Map the dynamic feature matrix from step 1 to the state space of reinforcement learning. Ensure that the model input is consistent with the physical conditions;
[0050] Step 2.3.1 Define the state space as follows: ,in These are the parameters for the furnace, including temperature, pressure, and CO concentration. The parameters are the dioxin concentration and residence time in the flue gas duct. These are parameters related to the metal content and particulate matter concentration of fly ash.
[0051] Step 2.3.2 Define the action space as follows The action must satisfy the hard constraints monitored by the sensor in step 1;
[0052] Step 2.4 Multi-objective collaborative optimization and dynamic weight connection
[0053] The weights are dynamically adjusted via the fuzzy logic module in step 3. , , To achieve a balance between multiple objectives;
[0054] Step 2.4.1 Dynamic Weight Input Interface:
[0055] Adjust the weights in step 3. , , Real-time injection of joint objective function:
[0056]
[0057] Step 2.4.2 Training process optimization:
[0058] An OU noise enhancement strategy was adopted to gradually attenuate the noise intensity:
[0059]
[0060] in, For Actor network policy functions, For policy network parameters, The output action for time step t, This is OU noise; the noise intensity is gradually reduced as the number of training iterations increases.
[0061] The loss function update formula is described as follows:
[0062] Critic loss function:
[0063]
[0064] in, For the loss of the Critic network, Let D be the mathematical expectation of the experience replay pool D, where D is the set storing tuples of historical experiences. For the target Q value, Here is the Q-value function for the Critic network, with the following parameters: ,Evaluate Action Value
[0065] Actor loss function:
[0066]
[0067] in, For the loss of the Actor network, The Critic network evaluates the actions of the current policy μ.
[0068] The three types of objective functions are integrated into a reward signal. :
[0069]
[0070] Step 2.5 provides a data interface and feedback mechanism for the dynamic weight adjustment module in Step 3.
[0071] Step 2.5.1 Real-time parameter feedback:
[0072] The parameters in the dynamic feature matrix are passed to the fuzzy logic module in step 3; when an abnormality in the actual working condition is detected, step 3 automatically adjusts the weights and updates the module. ;
[0073] Step 2.5.2 Model Iteration Update:
[0074] The reinforcement learning framework retrains the Actor-Critic network based on the adjusted weights to achieve dynamic optimization.
[0075] As a further improvement to the present invention, the dynamic weight allocation and multi-objective collaborative optimization control in step 3 is represented as follows:
[0076] Step 3.1 Design of Dynamic Weight Adjustment Module
[0077] Based on the dynamic feature matrix from step 1 and the multi-objective joint optimization model from step 2, a fuzzy logic-driven weight adaptive mechanism is constructed to dynamically adjust the weight coefficients of the three types of objective functions. , , This enables real-time optimization of combustion conditions;
[0078] Step 3.1.1 Fuzzy Logic Controller Design
[0079] Using triangular membership functions, input variables are mapped to linguistic variables "high", "medium", and "low", with thresholds set based on historical data.
[0080]
[0081] Fuzzy rule base:
[0082] High, increase , reduce and ; High; increase , reduce and ; High, increase , reduce and ; and In the middle, balancing priorities, that is ;
[0083] Step 3.1.2 Fuzzy Reasoning and Defuzzification
[0084] Trigger strength calculation:
[0085]
[0086] in, The rule trigger strength is calculated as the minimum of all input membership degrees, where k is the rule index. Input variables membership degree Input variables , , ;
[0087] Defuzzification:
[0088] The final weights are calculated using a weighted average method:
[0089]
[0090] in, This is a dynamic weight output, where N is the number of triggered rules; Assign weights to the k-th rule;
[0091] Step 3.2 Dynamic weights are fed back to reinforcement learning
[0092] Dynamic weights , , Real-time injection of the joint objective function in step 2 and reward signals This drives the policy update of the DDPG algorithm.
[0093] This invention presents a collaborative control method for incineration pollution based on multi-objective reinforcement learning, which has beneficial effects. The technical advantages of this invention are as follows:
[0094] 1. This invention utilizes a multi-objective reinforcement learning framework constructed through a deep deterministic policy gradient algorithm. This framework can precisely adjust control variables such as the combustion air ratio, secondary air velocity, and denitrification agent injection timing based on real-time operating conditions, resulting in a more stable and efficient combustion process. This effectively suppresses the formation of pollutants such as dioxins and particulate matter. Compared with traditional waste incineration technologies, this significantly reduces pollutant emissions, greatly mitigating environmental pollution and protecting the ecological environment and human health.
[0095] 2. This invention significantly improves the efficiency of fly ash metal recovery. It integrates an X-ray fluorescence analyzer into the fly ash conveying pipeline to monitor the heavy metal leaching rate in real time and incorporates this data into the calculation of the objective function for maximizing fly ash metal recovery. Through optimization and adjustment of relevant control variables using a multi-objective reinforcement learning framework, the fly ash treatment process can be dynamically adjusted based on real-time changes in the metal content of the fly ash, thereby improving the fly ash metal recovery rate. This not only achieves effective resource recycling and reduces resource waste but also creates certain economic benefits, aligning with the concept of sustainable development.
[0096] 3. This invention fully considers factors that significantly impact resource recovery costs, such as unit waste treatment energy consumption and denitrification agent injection volume, and reflects these factors in the objective function of minimizing resource recovery costs. Through the optimization and control of these factors using a multi-objective reinforcement learning framework, it achieves a reasonable reduction in energy consumption and precise denitrification agent injection, avoiding excessive use of denitrification agents and excessive energy consumption. This effectively reduces the resource recovery cost of waste incineration and improves the economic efficiency of waste incineration treatment. Attached Figure Description
[0097] Figure 1 This is a flowchart of the present invention;
[0098] Figure 2This is a structural diagram of the reinforcement learning model of the present invention. Detailed Implementation
[0099] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:
[0100] This invention focuses on waste incineration treatment, constructing a feature matrix by deploying multimodal sensors in the furnace and other locations to collect multiple parameters. Based on this, a multi-objective reinforcement learning joint optimization model is established, defining three types of objective functions. Simultaneously, a fuzzy logic module is constructed to dynamically adjust the weights. This achieves pollutant emission reduction, improved fly ash metal recovery, cost reduction, and multi-objective collaborative optimization. The invention flowchart is shown below. Figure 1 As shown, the steps of the present invention will be described in detail below.
[0101] Step 1: Constructing a dynamic feature matrix of incineration conditions and a multi-source data acquisition system
[0102] Multimodal sensor arrays are deployed in the furnace, flue, and fly ash treatment unit to collect parameters such as combustion temperature, flue gas residence time, denitrification agent injection amount, fly ash metal content, and dioxin / polycyclic aromatic hydrocarbon concentration in real time, and to construct a high-dimensional dynamic feature matrix.
[0103] Step 1.1 Sensor Selection and Installation Location
[0104] The combustion temperature sensors are thermocouple arrays deployed at six points in the furnace (top, middle, and bottom), covering the main combustion zone, secondary combustion zone, and cooling zone. Differential pressure sensors are installed at the furnace outlet and secondary air inlet to monitor negative pressure fluctuations. Gas concentration sensors are installed at three points in the furnace outlet flue to monitor combustion products in real time.
[0105] A flue gas residence time sensor is installed in the middle section of the flue, calculating residence time through optical path attenuation. An online dioxin / polycyclic aromatic hydrocarbon (PAH) monitor is deployed upstream of the flue gas purification system to detect trace pollutants. An X-ray fluorescence analyzer is used for metal content detection, integrated into the fly ash conveying pipeline to monitor heavy metal leaching rates in real time. A particulate matter concentration sensor is installed at the dust collector outlet to capture escaping particulate matter.
[0106] Step 1.2 Data Acquisition
[0107] Multi-source heterogeneous data is integrated into a standardized feature matrix to support the input requirements of subsequent multi-objective optimization models.
[0108] A sliding window with a window size of 10 seconds and a step size of 1 second is used to align asynchronous data with timestamps, eliminate sensor sampling delay differences, and perform One-Hot encoding on waste type classification variables to form a sensor feature matrix.
[0109] Step 2: Establish a multi-objective reinforcement learning joint optimization model
[0110] Based on the dynamic feature matrix, three types of objective functions are defined: minimizing pollutant emissions, maximizing fly ash metal recovery rate, and minimizing resource recovery cost. A multi-objective reinforcement learning framework is constructed using a deep deterministic policy gradient algorithm. Combustion air ratio, secondary air velocity, and denitrification agent injection timing are used as control variables in the action space. A dual-network structure is employed to realize policy exploration and value assessment. The reinforcement learning model structure diagram is shown below. Figure 2 As shown.
[0111] Step 2.1 Defining the objective function and relating it to the dynamic feature matrix
[0112] Based on the high-dimensional dynamic feature matrix constructed in step 1, three types of objective functions are defined: minimizing pollutant emissions, maximizing fly ash metal recovery rate, and minimizing resource utilization cost. These functions provide the input basis for the dynamic weight adjustment module in step 3.
[0113] Step 2.1.1 Minimize pollutant emissions
[0114] Using the real-time data from the flue area sensor in step 1, the calculation formula can be expressed as:
[0115]
[0116] in, and These are the weighting factors used to balance dioxins and particulate matter. This represents the rate of dioxin formation. This refers to the particulate matter concentration. The goal is to minimize pollutant emissions.
[0117] Physical constraints: furnace temperature Smoke residence time The sensor in step 1 monitors the data.
[0118] Step 2.1.2 Maximizing fly ash metal recovery rate
[0119] The calculation formula can be expressed as:
[0120]
[0121] in, To maximize the recovery rate of fly ash metals These are the weighting coefficients. The heavy metal leaching rate is based on the X-ray fluorescence analyzer collected by the fly ash treatment unit in step 1. This represents the maximum permissible heavy metal leaching rate. This expression is achieved by reducing... Improve the recovery rate.
[0122] Step 2.1.3 Minimize resource recovery costs
[0123] The calculation formula can be expressed as:
[0124]
[0125] in, To minimize resource recovery costs, These are the weighting coefficients. Energy consumption per unit of waste disposal.
[0126] Physical constraints: Denitrification agent injection volume .
[0127] Step 2.2 Construction of Multi-Objective Joint Optimization Model
[0128] The three types of objective functions are integrated into a joint optimization objective. A dynamic weight allocation mechanism was designed to connect the fuzzy logic module in step 3.
[0129] Step 2.2.1 Combining the objective functions:
[0130]
[0131] Weight initialization: , , .
[0132] The reward function is defined as:
[0133]
[0134] The reinforcement learning framework is designed with an Actor-Critic dual-network structure:
[0135] Actor network: outputs continuous control actions, combustion air ratio. Secondary wind speed Denitrification agent injection timing .
[0136] Critic Network: Evaluating Action Pairs The influence of gradient feedback is provided.
[0137] Experience replay and soft update mechanism:
[0138] Store historical experience To experience pool D.
[0139] Through soft update coefficients Stable training process.
[0140] Step 2.3 Input mapping and state space design of dynamic feature matrix
[0141] Map the dynamic feature matrix from step 1 to the state space of reinforcement learning. This ensures that the model input is consistent with the physical conditions.
[0142] Step 2.3.1 Define the state space as follows: ,in These are the parameters for the furnace, including temperature, pressure, and CO concentration. The parameters are the dioxin concentration and residence time in the flue gas duct. These are parameters related to the metal content and particulate matter concentration of fly ash.
[0143] Step 2.3.2 Define the action space as follows The action must meet the hard constraints monitored by the sensor in step 1.
[0144] Step 2.4 Multi-objective collaborative optimization and dynamic weight connection
[0145] The weights are dynamically adjusted via the fuzzy logic module in step 3. , , This allows for a balance between multiple objectives.
[0146] Step 2.4.1 Dynamic Weight Input Interface:
[0147] Adjust the weights in step 3. , , Real-time injection of joint objective function:
[0148]
[0149] Step 2.4.2 Training process optimization:
[0150] An OU noise enhancement strategy was adopted to gradually attenuate the noise intensity:
[0151]
[0152] in, For Actor network policy functions, For policy network parameters, The output action for time step t, This is OU noise. The noise intensity is gradually reduced as the number of training iterations increases.
[0153] The loss function update formula is described as follows:
[0154] Critic loss function:
[0155]
[0156] in, For the loss of the Critic network, Let D be the mathematical expectation of the experience replay pool D, where D is the set storing tuples of historical experiences. For the target Q value, Here is the Q-value function for the Critic network, with the following parameters: ,Evaluate Action Value
[0157] Actor loss function:
[0158]
[0159] in, For the loss of the Actor network, This is the Critic network's evaluation of the current policy μ's actions.
[0160] The three types of objective functions are integrated into a reward signal. :
[0161]
[0162] Step 2.5 provides a data interface and feedback mechanism for the dynamic weight adjustment module in Step 3.
[0163] Step 2.5.1 Real-time parameter feedback:
[0164] The parameters in the dynamic feature matrix are passed to the fuzzy logic module in step 3. When an anomaly in the actual operating condition is detected, step 3 automatically adjusts the weights and updates the module. .
[0165] Step 2.5.2 Model Iteration Update:
[0166] The reinforcement learning framework retrains the Actor-Critic network based on the adjusted weights to achieve dynamic optimization.
[0167] Step 3: Implement dynamic weight allocation and multi-objective collaborative optimization control
[0168] A priority evaluation module based on fuzzy logic is constructed to dynamically adjust the weight coefficients of the objective function according to real-time operating conditions.
[0169] Step 3.1 Design of Dynamic Weight Adjustment Module
[0170] Based on the dynamic feature matrix from step 1 and the multi-objective joint optimization model from step 2, a fuzzy logic-driven weight adaptive mechanism is constructed to dynamically adjust the weight coefficients of the three types of objective functions. , , This enables real-time optimization of combustion conditions.
[0171] Step 3.1.1 Fuzzy Logic Controller Design
[0172] Using triangular membership functions, input variables are mapped to linguistic variables "high", "medium", and "low", with thresholds set based on historical data.
[0173]
[0174] Fuzzy rule base:
[0175] High, increase , reduce and ; High. Increase , reduce and ; High, increase , reduce and ; and In the middle, balancing priorities, that is .
[0176] Step 3.1.2 Fuzzy Reasoning and Defuzzification
[0177] Trigger strength calculation:
[0178]
[0179] in, The rule trigger strength is calculated as the minimum of all input membership degrees, where k is the rule index. Input variables membership degree Input variables , , .
[0180] Defuzzification:
[0181] The final weights are calculated using a weighted average method:
[0182]
[0183] in, This is the dynamic weight output, where N is the number of trigger rules. Assign weights to the k-th rule.
[0184] Step 3.2 Dynamic weights are fed back to reinforcement learning
[0185] Dynamic weights , , Real-time injection of the joint objective function in step 2 and reward signals This drives the policy update of the DDPG algorithm.
[0186] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other way. Any modifications or equivalent changes made based on the technical essence of the present invention shall still fall within the scope of protection claimed by the present invention.
Claims
1. A collaborative control method for incineration pollution based on multi-objective reinforcement learning, comprising the following specific steps, characterized in that: Step 1: Construct a dynamic feature matrix of incineration conditions and a multi-source data acquisition system. Deploy multi-modal sensor arrays in the furnace, flue, and fly ash treatment unit to collect parameters such as combustion temperature, flue gas residence time, denitrification agent injection amount, fly ash metal content, and dioxin / polycyclic aromatic hydrocarbon concentration in real time, and construct a high-dimensional dynamic feature matrix. Step 2: Establish a multi-objective reinforcement learning joint optimization model. Based on the dynamic feature matrix, define three types of objective functions: minimizing pollutant emissions, maximizing fly ash metal recovery rate, and minimizing resource recovery cost. A multi-objective reinforcement learning framework is constructed using a deep deterministic policy gradient algorithm. The combustion air ratio, secondary wind speed, and denitrification agent injection timing are used as control variables in the action space. Policy exploration and value assessment are achieved through a dual-network structure. The multi-objective reinforcement learning joint optimization model established in step 2 is represented as follows: Step 2.1 Define the objective function and associate it with the dynamic feature matrix; Based on the high-dimensional dynamic feature matrix constructed in step 1, three types of objective functions are defined: minimizing pollutant emissions, maximizing fly ash metal recovery rate, and minimizing resource utilization cost, which provide the input basis for the dynamic weight adjustment module in step 3. Step 2.1.1 Minimize pollutant emissions; Using the real-time data from the flue area sensor in step 1, the calculation formula can be expressed as: ; in, and These are the weighting factors used to balance dioxins and particulate matter. This represents the rate of dioxin formation. This refers to the particulate matter concentration. The goal is to minimize pollutant emissions; Physical constraints: furnace temperature Smoke residence time Monitored by the sensor in step 1; Step 2.1.2 Maximize fly ash metal recovery rate; The calculation formula can be expressed as: ; in, The goal is to maximize the recovery rate of fly ash metals. These are the weighting coefficients. The heavy metal leaching rate is based on the X-ray fluorescence analyzer collected by the fly ash treatment unit in step 1. The maximum allowable heavy metal leaching rate; this expression is achieved by reducing... Improve recovery rate; Step 2.1.3 Minimize resource recovery costs; The calculation formula can be expressed as: ; in, To minimize resource recovery costs, These are the weighting coefficients. Energy consumption per unit of waste disposal; Physical constraints: Denitrification agent injection volume ; Step 2.2 Construction of a multi-objective joint optimization model; The three types of objective functions are integrated into a joint optimization objective. And a dynamic weight allocation mechanism is designed to connect the fuzzy logic module in step 3; Step 2.2.1 Combining the objective functions: ; Weight initialization: , , ; The reward function is defined as: ; The reinforcement learning framework is designed with an Actor-Critic dual-network structure: Actor network: outputs continuous control actions, combustion air ratio. Secondary wind speed Denitrification agent injection timing ; Critic Network: Evaluating Action Pairs The influence of gradient feedback is provided; Experience replay and soft update mechanism: Store historical experience To experience pool D; Through soft update coefficients Stabilize the training process; Step 2.3 Input mapping and state space design of dynamic feature matrix; Map the dynamic feature matrix from step 1 to the state space of reinforcement learning. Ensure that the model input is consistent with the physical conditions; Step 2.3.1 Define the state space as follows: ,in These are the parameters for the furnace, including temperature, pressure, and CO concentration. The parameters are the dioxin concentration and residence time in the flue gas duct. These are parameters related to the metal content and particulate matter concentration of fly ash. Step 2.3.2 Define the action space as follows The action must satisfy the hard constraints monitored by the sensor in step 1; Step 2.4 Multi-objective collaborative optimization and dynamic weight connection; The weights are dynamically adjusted via the fuzzy logic module in step 3. , , To achieve a balance between multiple objectives; Step 2.4.1 Dynamic Weight Input Interface: Adjust the weights in step 3. , , Real-time injection of joint objective function: ; Step 2.4.2 Training process optimization: An OU noise enhancement strategy was adopted to gradually attenuate the noise intensity: ; in, For Actor network policy functions, For policy network parameters, The output action for time step t, This is OU noise; the noise intensity is gradually reduced as the number of training iterations increases. The loss function update formula is described as follows: Critic loss function: ; in, For the loss of the Critic network, Let D be the mathematical expectation of the experience replay pool D, where D is the set storing tuples of historical experiences. For the target Q value, Here is the Q-value function for the Critic network, with the following parameters: ,Evaluate Action Value Actor loss function: ; in, For the loss of the Actor network, The Critic network evaluates the actions of the current policy μ. The three types of objective functions are integrated into a reward signal. : ; Step 2.5 provides a data interface and feedback mechanism for the dynamic weight adjustment module in Step 3; Step 2.5.1 Real-time parameter feedback: The parameters in the dynamic feature matrix are passed to the fuzzy logic module in step 3; when an abnormality in the actual working condition is detected, step 3 automatically adjusts the weights and updates the module. ; Step 2.5.2 Model Iteration Update: The reinforcement learning framework retrains the Actor-Critic network based on the adjusted weights to achieve dynamic optimization. Step 3: Implement dynamic weight allocation and multi-objective collaborative optimization control, construct a priority evaluation module based on fuzzy logic, and dynamically adjust the weight coefficients of the objective function according to real-time operating conditions; The dynamic weight allocation and multi-objective collaborative optimization control implemented in step 3 are represented as follows: Step 3.1 Design of Dynamic Weight Adjustment Module Based on the dynamic feature matrix from step 1 and the multi-objective joint optimization model from step 2, a fuzzy logic-driven weight adaptive mechanism is constructed to dynamically adjust the weight coefficients of the three types of objective functions. , , This enables real-time optimization of combustion conditions; Step 3.1.1 Fuzzy Logic Controller Design Using triangular membership functions, input variables are mapped to linguistic variables "high", "medium", and "low", with thresholds set based on historical data. ; Fuzzy rule base: High, increase , reduce and ; High; increase , reduce and ; High, increase , reduce and ; and In the middle, balancing priorities, that is ; Step 3.1.2 Fuzzy Reasoning and Defuzzification Trigger strength calculation: ; in, The rule trigger strength is calculated as the minimum of all input membership degrees, where k is the rule index. Input variables membership degree Input variables , , ; Defuzzification: The final weights are calculated using a weighted average method: ; in, This is a dynamic weight output, where N is the number of triggered rules; Assign weights to the k-th rule; Step 3.2: Dynamic weights are fed back to reinforcement learning; Dynamic weights , , Real-time injection of the joint objective function in step 2 and reward signals This drives the policy update of the DDPG algorithm.
2. The method for coordinated control of incineration pollution based on multi-objective reinforcement learning according to claim 1, characterized in that: The dynamic feature matrix of incineration conditions and the multi-source data acquisition system constructed in step 1 are represented as follows: Step 1.1 Sensor selection and installation location; The combustion temperature sensor uses a thermocouple array, deployed at 6 points at the top, middle, and bottom of the furnace, covering the main combustion zone, secondary combustion zone, and cooling zone; the pressure sensor uses a differential pressure sensor, installed at the furnace outlet and secondary air inlet, to monitor negative pressure fluctuations; the gas concentration sensor is installed at 3 points in the furnace outlet flue to monitor combustion products in real time. A flue gas residence time sensor is installed in the middle section of the flue and calculates the residence time through optical path attenuation; a dioxin / polycyclic aromatic hydrocarbon online monitor is deployed at the front of the flue gas purification system to detect trace pollutants; an X-ray fluorescence analyzer is used to detect metal content and is integrated into the fly ash conveying pipeline to monitor the heavy metal leaching rate in real time; a particulate matter concentration sensor is installed at the dust collector outlet to capture particulate matter escaping. Step 1.2 Data Acquisition; Multi-source heterogeneous data is integrated into a standardized feature matrix to support the input requirements of subsequent multi-objective optimization models; A sliding window with a window size of 10 seconds and a step size of 1 second is used to align asynchronous data with timestamps, eliminate sensor sampling delay differences, and perform One-Hot encoding on waste type classification variables to form a sensor feature matrix.