Intelligent Control Method for Multi-Objective Optimization of Combustion-Product Carbon Dioxide System

By building a dual-drive model and deep reinforcement learning agent, combining multi-objective reward function and layered progressive optimization control, the problem of efficient combustion and collaborative control of pollutants and carbon emissions in the garbage combustion system is solved, and the multi-objective optimization and stable operation of the system is achieved.

CN119983287BActive Publication Date: 2025-07-04CHN ENERGY JIANGSU POWER CO LTD +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510480882.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-04
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

When the existing garbage combustion control system deals with complex domestic waste, it is difficult to achieve efficient coordinated control of combustion and pollutants and carbon emissions. The traditional mechanism model is insufficiently adaptable, and the data-driven model is difficult to fully ensure the stable operation of the system and multi-target optimization.

Method used

Build a dual-drive model that coordinates combustion mechanism model and combustion data model, design a mechanism-constrained deep reinforcement learning agent, generate control strategies through multi-objective reward function, perform hierarchical progressive optimization control, and combine physical feasibility checks and dynamic weight adjustments to achieve maximum combustion efficiency, minimize carbon dioxide emissions and pollutant control standards.

Benefits of technology

It improves the control accuracy and stability of the combustion system, achieves the improvement of combustion efficiency and the reduction of carbon dioxide emissions and pollutant emissions, and promotes the efficient and green operation of the combustion carbon dioxide production system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119983287B_ABST
    Figure CN119983287B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of intelligent control of equipment, and discloses a multi-objective optimization intelligent control method for a carbon dioxide combustion production system. The method includes: constructing a combustion mechanism model; training a combustion data model; fusing the combustion mechanism model and the combustion data model to obtain a dual-drive model and determining optimization objectives; using the dual-drive model to simulate the waste combustion process, and optimizing the control strategy according to the optimization objectives to adjust the combustion equipment. The present invention effectively integrates the advantages of the two models, realizes the complementary advantages between the models, thereby improving the physical interpretability and adaptability of the models. During the optimization process of the control strategy, not only the combustion efficiency is considered, but also the carbon emissions and pollutant emissions are taken into account, realizing the improvement of the combustion efficiency and the reduction of the carbon emissions and pollutant emissions, thereby realizing the multi-objective optimization of the carbon dioxide combustion production system, and further improving the environmental protection performance and economic performance of the waste carbon dioxide combustion production system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent control of equipment, and more specifically, to a multi-objective optimization intelligent control method for a combustion-produced carbon dioxide system. Background Art

[0002] The combustion of domestic waste, as an effective treatment method, is widely adopted because it can significantly reduce the volume of waste, eliminate harmful substances and recover energy; however, a large amount of pollutants will be generated during the waste combustion process, such as carbon dioxide (CO2), nitrogen oxides (NO x ), dioxins, etc. The emissions of these pollutants not only cause pressure on the environment but also pose a potential threat to human health; in addition, the carbon emissions during the combustion process also have a direct impact on global climate change; therefore, how to effectively control pollutant emissions and achieve coordinated emission reduction of carbon emissions while ensuring efficient waste combustion has become a key problem to be solved urgently.

[0003] Traditional waste combustion control systems mainly rely on physical and chemical mechanism models to predict and adjust key parameters during the combustion process, such as the real-time prediction system and method for the calorific value of the incoming fuel of a circulating fluidized bed domestic waste incineration boiler disclosed in Chinese Patent with the authorized announcement number CN105864797B; however, due to the complex composition and large calorific value fluctuations of domestic waste, these mechanism models often face challenges such as insufficient adaptability and low prediction accuracy in practical applications.

[0004] With the development of information technology and big data technology, data-driven modeling methods have been gradually widely used. By learning a large amount of historical data, they can better handle the nonlinear behavior of the system and adapt to changing operating conditions, such as the intelligent control system for a waste incineration power plant and a waste combustion-produced carbon dioxide system disclosed in Chinese Patent Application with the publication number CN118089034A; however, due to the complex and variable composition of waste, it is difficult to accurately grasp the physical and chemical characteristics of the combustion process, and it is difficult to achieve the optimal control of the system; only relying on the thermal efficiency index for dynamic constraint relaxation may lead to the failure of local optimization and cannot comprehensively ensure the stable operation and multi-objective optimization of the system under different working conditions; therefore, there is an urgent need for an intelligent control method that combines mechanism models and data-driven models.

[0005] In view of this, the present invention proposes a multi-objective optimization intelligent control method for a combustion-produced carbon dioxide system to solve the above problems. Summary of the Invention

[0006] To overcome the above-mentioned defects of the prior art and achieve the above object, the present invention provides the following technical solution: A multi-objective optimization intelligent control method for a combustion-produced carbon dioxide system, comprising:

[0007] S1: Construct a dual - drive model synergistically composed of a combustion mechanism model and a combustion data model;

[0008] S2: Use the dual - drive model to simulate the waste combustion process, and design a mechanism - constrained deep reinforcement learning agent according to the dual - drive model;

[0009] S3: Optimize the control strategy according to the optimization objective, perform hierarchical progressive optimization control, and adjust the combustion equipment.

[0010] Further, the steps of constructing the combustion mechanism model include:

[0011] Step S101: Analyze the waste characteristics and determine the characteristic parameters;

[0012] The characteristic parameters include physical characteristic parameters and chemical characteristic parameters; the physical characteristic parameters include moisture content and thermal conductivity, and the chemical characteristic parameters include calorific value, ash content, combustible content, fixed carbon, and volatile matter;

[0013] Step S102: Analyze the reaction mechanism of waste combustion and construct a combustion mechanism model;

[0014] The combustion mechanism model consists of a pyrolysis model, a combustion model, a flue gas generation model, and a pollutant generation model;

[0015] Step S103: Use numerical calculation methods to simulate the waste combustion process, and optimize the combustion mechanism model according to the simulation results;

[0016] Convert the equations of each model in the combustion mechanism model into discrete mathematical forms; set the initial conditions; use numerical calculation methods to solve the discrete equations, simulate the evolution process of each stage in the waste combustion process; perform post - processing on the simulation results; optimize the model parameters of the combustion mechanism model according to the differences between the post - processed simulation results and the characteristic parameters.

[0017] Further, the construction process of the pyrolysis model is: determine the reaction mechanism of the pyrolysis model, and the reaction mechanism of the pyrolysis model is to describe the decomposition reaction of waste at different temperatures; select the pyrolysis model; establish mass and energy balance equations to simulate the reaction kinetics during pyrolysis; calibrate the pyrolysis model through the corresponding characteristic parameters;

[0018] The construction process of the combustion model is: determine the reaction mechanism of the combustion model, and the reaction mechanism of the combustion model is to describe the main combustion reactions; select the combustion model; establish mass, energy, and momentum balance equations for the combustion process to simulate the conversion of reactants and products; calibrate the combustion model through the corresponding characteristic parameters;

[0019] The construction process of the flue gas generation model is as follows: Determine the reaction mechanism of the flue gas generation model, and the reaction mechanism of the flue gas generation model is to describe the generation and conversion processes of each component in the flue gas; Select the flue gas generation model; Establish the mass balance and chemical reaction equations of flue gas generation; Calibrate the flue gas generation model through corresponding characteristic parameters;

[0020] The construction process of the pollutant generation model is as follows: Determine the reaction mechanism of the pollutant generation model, and the reaction mechanism of the pollutant generation model is to describe the chemical reaction path of pollutant generation; Select the pollutant generation model; Establish the chemical reaction equation of pollutant generation; Calibrate the pollutant generation model through corresponding characteristic parameters.

[0021] Furthermore, the construction method of the mechanism-constrained deep reinforcement learning agent includes:

[0022] Define the state space as the vector of combustion parameters collected in real time, including environmental parameters and characteristic parameters of the garbage;

[0023] Define the action space as the set of adjustable parameters of the combustion equipment, including the adjustment amount of oxygen supply, the increment of grate speed, and the continuous control amount of the opening of the secondary air damper;

[0024] Construct an Actor-Critic network architecture containing a physical feasibility verification gate, where the Actor network is used to generate continuous control actions, and the Critic network is used to design a multi-objective reward function and evaluate multi-objective rewards;

[0025] The design method of the multi-objective reward function includes:

[0026] Construct an efficiency improvement term, a carbon emission penalty term, and a pollutant over-standard gradient penalty term by combining the combustion efficiency, carbon emissions, and pollutant concentration through dynamic weight coefficients, and obtain the multi-objective reward function according to the efficiency improvement term, the carbon emission penalty term, and the pollutant over-standard gradient penalty term. Among them, the carbon emission penalty intensity increases non-linearly with the temperature deviating from the optimal interval.

[0027] Furthermore, the implementation method of the physical feasibility verification gate includes: Add a constraint verification module to the output layer of the Critic network, and the constraint verification module performs non-linear fusion on the predicted Q value and the physical feasibility coefficient through a neural network layer, and outputs the action value evaluation result corrected by physical constraints.

[0028] Furthermore, the adjustment method of the dynamic weight coefficient includes:

[0029] Use the operation data within a preset time period as the input of the stage recognition model to obtain the current stage, which includes the startup stage, the steady-state stage, and the transient stage; in the startup stage, increase the dynamic weight coefficient corresponding to the combustion efficiency to a preset weight threshold, and randomly allocate the dynamic weight coefficients corresponding to the carbon emission penalty term and the pollutant over-standard gradient penalty term according to the constraint condition that the sum of all dynamic weight coefficients is 1; in the steady-state stage, the dynamic weight coefficients corresponding to the carbon emission penalty term and the pollutant over-standard gradient penalty term are dynamically allocated according to the pollutant emission quota; in the transient stage, introduce an emergency adjustment factor to adjust the dynamic weight coefficient corresponding to the pollutant over-standard gradient penalty term.

[0030] Furthermore, the calibration method of the carbon emission penalty term includes:

[0031] Conduct a stepped heating experiment in different temperature ranges of the combustion furnace, collect the material expansion coefficients at each temperature point, calibrate the temperature influence factor through the material expansion coefficients at each temperature point, establish a mapping relationship between the temperature deviation degree and the carbon emission penalty gain coefficient, and calibrate the carbon emission penalty term through the mapping relationship between the temperature deviation degree and the carbon emission penalty gain coefficient.

[0032] Furthermore, the optimization objectives of the optimization control strategy include maximizing the combustion efficiency, minimizing the carbon dioxide emissions, and meeting the pollutant control standards.

[0033] Furthermore, the method for implementing hierarchical progressive optimization control includes:

[0034] Conduct pre-training under the constraint of the mechanism model in a virtual environment to generate an initial control strategy library;

[0035] Implement data-driven policy fine-tuning through a dual experience replay pool;

[0036] Dynamically relax the action space constraint according to the operation stability index;

[0037] The dual experience replay pool includes a physical compliance pool and an actual optimization pool; the physical compliance pool is used to store experience samples that meet the preset constraint conditions; the actual optimization pool is used to store samples that are met during the operation process; among them, the sampling priority of the dual experience replay pool is dynamically determined by the weighted fusion result of the physical compliance score and the reward value.

[0038] Furthermore, the method for dynamically relaxing the action space constraint includes:

[0039] When the thermal efficiency volatility in N consecutive control cycles is not higher than the preset volatility threshold, update the constraint parameters;

[0040] The method for updating the constraint parameters includes:

[0041] The adaptive relaxation adjustment of the action space constraint parameters is carried out according to the operation stability index, and when the performance index reaches the preset index threshold, the allowable control quantity change range is enlarged proportionally.

[0042] The technical effects and advantages of the multi-objective optimization intelligent control method for the carbon dioxide combustion system of the present invention are as follows:

[0043] By constructing a dual-drive model to accurately simulate the combustion process, it lays a foundation for intelligent control; based on this, a mechanism-constrained deep reinforcement learning agent is designed, which can generate reasonable control strategies according to the multi-objective reward function, realizing the maximization of combustion efficiency, the minimization of carbon dioxide emissions, and the compliance of pollutant control. The physical feasibility verification gate ensures that the control actions conform to the physical laws, and the dynamic weight coefficient is adjusted according to different stages, making the multi-objective optimization more targeted, and the carbon emission penalty term is accurately calibrated to strictly restrict carbon emissions. In the hierarchical progressive optimization control, the virtual environment is pre-trained to generate an initial policy library, the dual experience replay pool fine-tunes the policy, and the dynamic relaxation of the action space constraint enables the system to flexibly explore the optimization space. The comprehensive effect of these technical means improves the accuracy, intelligence, and stability of the system control, realizes multi-objective collaborative optimization, effectively solves the problems of incomplete combustion, unstable emissions, and low energy conversion rate in traditional combustion systems, and promotes the efficient and green operation of the carbon dioxide combustion system. Description of the Drawings

[0044] Figure 1 It is the flowchart of the multi-objective optimization intelligent control method for the carbon dioxide combustion system in Embodiment 1 of the present invention;

[0045] Figure 2 It is the schematic flowchart of the design method of the deep reinforcement learning agent in Embodiment 1 of the present invention;

[0046] Figure 3 It is the schematic flowchart of the dynamic weight coefficient adjustment method in Embodiment 1 of the present invention;

[0047] Figure 4 It is the schematic flowchart of the hierarchical progressive optimization control in Embodiment 1 of the present invention. Detailed Embodiments

[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0049] Embodiment 1

[0050] Please refer to Figure 1As shown in the figure, the multi-objective optimization intelligent control method for the combustion-produced carbon dioxide system in this embodiment includes:

[0051] S1: Construct a dual-drive model coordinated by a combustion mechanism model and a combustion data model.

[0052] The combustion mechanism model is a mathematical model that describes the combustion process based on the physical and chemical laws of the waste combustion process and uses basic theories such as thermodynamics, fluid mechanics, and chemical reaction kinetics.

[0053] The steps for constructing the combustion mechanism model include:

[0054] Step S101: Analyze the waste characteristics, determine the characteristic parameters, and provide the necessary input data for the combustion mechanism model.

[0055] The characteristic parameters include physical characteristic parameters and chemical characteristic parameters; the physical characteristic parameters at least include moisture content, thermal conductivity, etc., and the chemical characteristic parameters at least include calorific value, ash content, combustible content, fixed carbon, volatile matter, etc.; the characteristic parameters of different wastes are obtained by those skilled in the art through waste combustion experiments;

[0056] The moisture content is the proportion of moisture in the waste; the thermal conductivity is the heat conduction ability of the waste; the calorific value is the heat released when the waste is completely burned; the ash content is the proportion of inorganic residues after the waste is completely burned; the combustible content is the proportion of the components in the waste that can burn, the fixed carbon is the remaining carbon content after removing moisture and volatile organic compounds from the waste; the volatile matter is the gas components released during the waste combustion process.

[0057] Step S102: Analyze the reaction mechanism of waste combustion and construct the combustion mechanism model.

[0058] The combustion mechanism model is at least composed of a pyrolysis model, a combustion model, a flue gas generation model, and a pollutant generation model.

[0059] The construction process of the pyrolysis model is as follows: Determine the reaction mechanism of the pyrolysis model. The reaction mechanism of the pyrolysis model is to describe the decomposition reaction of the waste at different temperatures, such as the release of volatile matter and the solid carbonization process, etc.; Select the pyrolysis model. The pyrolysis model is, for example, an empirical model (such as TGA data), a chemical kinetics model (such as the Arrhenius equation), a multiphase reaction model, etc.; Set up mass and energy balance equations to simulate the reaction kinetics during the pyrolysis process; Calibrate the pyrolysis model through the corresponding characteristic parameters to ensure that the pyrolysis model accurately reflects the pyrolysis behavior of the waste; The mass and energy balance equations are set up by those skilled in the art according to the actual pyrolysis experiment situation.

[0060] The construction process of the combustion model is as follows: Determine the reaction mechanism of the combustion model. The reaction mechanism of the combustion model is to describe the main combustion reactions, such as the oxidation reaction of organic substances, the combustion of gaseous and solid fuels, the formation of nitrogen oxides, the oxidation of sulfides, etc.; Select a combustion model. Combustion models include, for example, fluid mechanics, reaction kinetics models (such as chemical reaction networks, reaction kinetic equations), etc.; Establish the mass, energy, and momentum balance equations of the combustion process to simulate the conversion of reactants and products; Calibrate the combustion model through corresponding characteristic parameters to ensure that the combustion model accurately reflects the combustion process of garbage; The mass, energy, and momentum balance equations of the combustion process are established by those skilled in the art according to the actual combustion experiment situation.

[0061] The construction process of the flue gas generation model is as follows: Determine the reaction mechanism of the flue gas generation model. The reaction mechanism of the flue gas generation model is to describe the generation and conversion processes of various components in the flue gas, such as the chemical reactions after combustion; Select a flue gas generation model. Flue gas generation models include, for example, flue gas flow models, chemical reaction models (such as reaction kinetics models, reaction mechanism models), etc.; Establish the mass balance and chemical reaction equations of flue gas generation, taking gas diffusion into account; Calibrate the flue gas generation model through corresponding characteristic parameters to ensure that the flue gas generation model accurately reflects the flue gas generation process of garbage; The mass balance and chemical reaction equations of flue gas generation are established by those skilled in the art according to the actual flue gas generation experiment situation.

[0062] The construction process of the pollutant generation model is as follows: Determine the reaction mechanism of the pollutant generation model. The reaction mechanism of the pollutant generation model is to describe the chemical reaction paths of pollutant generation, such as the formation of chlorides, the release of heavy metals, etc.; Select a pollutant generation model. Pollutant generation models include, for example, chemical reaction networks, poison release models, etc.; Establish the chemical reaction equations of pollutant generation, taking the sources and conversion paths of pollutant generation into account; Calibrate the pollutant generation model through corresponding characteristic parameters to ensure that the pollutant generation model accurately reflects the pollutant generation process of garbage; The chemical reaction equations of pollutant generation are established by those skilled in the art according to the actual pollutant generation experiment situation.

[0063] Step S103: Use numerical calculation methods to simulate the garbage combustion process and optimize the combustion mechanism model according to the simulation results.

[0064] The equations of each model in the combustion mechanism model are all transformed into discrete mathematical forms so that numerical solutions can be carried out; set initial conditions, such as combustion temperature, combustion pressure, feed rate, etc., and the initial conditions are the initial data of the waste combustion experiment when obtaining the above characteristic parameters; use numerical calculation methods to solve the discrete equations and simulate the evolution process of each stage during the waste combustion process. Numerical calculation methods include, for example, the finite element difference method, the finite element method, the finite volume method, etc.; post-process the simulation results. Post-processing includes, for example, calculating thermal conductivity, calorific value, ash content, etc.; optimize the model parameters of the combustion record model according to the differences between the simulation results after post-processing and the characteristic parameters. Model parameters include, for example, reaction kinetic constants, reaction rate constants, etc.

[0065] The combustion data model adopts an online learning neural network structure, and the input layer embeds the output features of the mechanism model to dynamically correct the strategy deviation of the output of the combustion mechanism model.

[0066] It should be noted that the combustion mechanism model can simulate stages such as pyrolysis, combustion, flue gas generation, and pollutant generation during the waste combustion process by analyzing waste characteristics and combustion reaction mechanisms, and can deeply reveal the physical and chemical essence of the combustion process, providing a theoretical basis for control strategies; the combustion data model is based on a large amount of actual operation data and can reflect the complex nonlinear relationships and uncertainties during the combustion process; the dual-drive model formed by the cooperation of the two combines the interpretability of the mechanism model and the adaptability of the data model, providing a more accurate simulation environment for designing mechanism-constrained deep reinforcement learning agents.

[0067] Refer to Figure 2 , S2: Use the dual-drive model to simulate the waste combustion process, and design a mechanism-constrained deep reinforcement learning agent according to the dual-drive model; the mechanism-constrained deep reinforcement learning agent can generate effective control strategies according to the dual-drive model, and perform hierarchical progressive optimization control considering multiple objectives such as maximizing combustion efficiency, minimizing carbon dioxide emissions, and meeting pollutant control standards, so as to achieve precise adjustment of the combustion equipment, thereby improving the comprehensive performance of the entire combustion carbon dioxide system.

[0068] The construction method of the mechanism-constrained deep reinforcement learning agent includes:

[0069] Define the state space as the vector of combustion parameters collected in real time, including environmental parameters and waste characteristic parameters;

[0070] Define the action space as the set of adjustable parameters of the combustion equipment, including continuous control quantities such as oxygen supply adjustment amount, grate speed increment, and secondary air damper opening degree;

[0071] Construct an Actor-Critic network architecture that includes a physical feasibility verification gate, where the Actor network is used to generate continuous control actions, and the Critic network is used to design a multi-objective reward function and evaluate multi-objective rewards.

[0072] By defining the combustion parameter vector collected in real time as the state space, the agent can fully perceive the actual situation of the combustion process and master the basic information of the system operation; taking the set of adjustable parameters of the combustion equipment as the action space enables the agent to make effective adjustment actions for different states; construct an Actor-Critic network architecture that includes a physical feasibility verification gate, where the Actor network generates continuous control actions to ensure that reasonable control instructions can be output according to the state; the Critic network designs a multi-objective reward function to evaluate multi-objective rewards, which can guide the agent to learn and optimize in the direction of maximizing combustion efficiency, minimizing carbon dioxide emissions, and meeting pollutant control standards, etc.; overall, the mechanism-constrained deep reinforcement learning agent construction method can improve the intelligence, effectiveness, and rationality of system control, and achieve precise and efficient multi-objective optimal control of the carbon dioxide production system by combustion.

[0073] The implementation method of the physical feasibility verification gate includes:

[0074] Add a constraint verification module to the output layer of the Critic network. The constraint verification module performs non-linear fusion of the predicted Q value and the physical feasibility coefficient through a neural network layer, and outputs the action value evaluation result corrected by physical constraints. For example ; where is the action value evaluation result corrected by physical constraints; is the predicted Q value; is the verification parameter, which can be obtained by optimizing through a natural inspiration optimization algorithm; is the physical feasibility coefficient calculated in real time by the combustion mechanism model; is the trainable parameter matrix, which automatically learns the physical constraint weights through backpropagation;

[0075] In the multi-objective optimal intelligent control of the carbon dioxide production system by combustion, the above formula fuses the predicted Q value and the physical feasibility coefficient calculated in real time by the mechanism model, and learns the physical constraint weights with the help of the trainable parameter matrix, and finally outputs the action value evaluation corrected by physical constraints; the above formula enables the agent to consider both the optimization objectives and strictly follow the physical laws of the combustion process when evaluating the value of control actions, avoiding generating invalid actions that do not conform to physical feasibility, and ensuring that the control strategy can still effectively promote multi-objective optimization under actual physical constraints, improving the reliability and stability of system control, and making the optimization process more in line with the actual needs of the project.

[0076] The design method of the multi-objective reward function includes:

[0077] By combining the combustion efficiency, carbon emissions, and pollutant concentration with dynamic weight coefficients, an efficiency improvement term, a carbon emission penalty term, and a pollutant over-standard gradient penalty term are constructed. A multi-objective reward function is obtained based on the efficiency improvement term, the carbon emission penalty term, and the pollutant over-standard gradient penalty term. Among them, the carbon emission penalty intensity increases non-linearly as the temperature deviates from the optimal range. As the multi-objective reward function at time ; among them, corresponds to the dynamic weight coefficient of the combustion efficiency, that is, the physical feasibility coefficient calculated in real time by the combustion mechanism model; is the combustion efficiency at time is the maximum combustion efficiency; is the dynamic weight coefficient of the carbon emissions; is the carbon emissions at time is the carbon emissions threshold; is the temperature influence factor; is a mathematical constant; is the thermal stress coefficient of the combustion furnace material; is the temperature at time is the temperature threshold; is the dynamic weight coefficient of the concentration of the th pollutant; is the concentration of the th pollutant at time ;

[0078] Referring to Figure 3 the adjustment method of the dynamic weight coefficient includes:

[0079] Take the operation data within a preset time period as the input of the stage recognition model to obtain the current stage, where the current stage includes a startup stage, a steady-state stage, and a transient stage; in the startup stage, increase the dynamic weight coefficient corresponding to the combustion efficiency to a preset weight threshold, and randomly allocate the dynamic weight coefficients corresponding to the carbon emission penalty term and the pollutant over-standard gradient penalty term according to the constraint that the sum of all dynamic weight coefficients is 1; in the steady-state stage, the dynamic weight coefficients corresponding to the carbon emission penalty term and the pollutant over-standard gradient penalty term are dynamically allocated according to the pollutant emission quota; in the transient stage, introduce an emergency adjustment factor to adjust the dynamic weight coefficient corresponding to the pollutant over-standard gradient penalty term.

[0080] By inputting the operation data of the preset time period into the stage recognition model (such as the hidden Markov model and the support vector machine, etc.), it is possible to accurately determine whether the system is in the startup, steady-state, or transient stage. In the startup stage, increasing the weight of the combustion efficiency to the preset threshold and randomly allocating the weights of the carbon emission penalty term and the pollutant over-standard gradient penalty term can quickly improve the combustion efficiency of the system to reach a stable operation state; in the steady-state stage, dynamically allocating the relevant weights according to the pollutant emission quota can enable the system to continuously optimize the carbon emission and pollutant emission control during stable operation; in the transient stage, introducing an emergency adjustment factor to adjust the weight of the pollutant over-standard gradient penalty term can enable the system to respond to emergencies in a timely manner and quickly control the pollutant emissions. This method of dynamically adjusting the weight coefficients according to different stages enables the system to flexibly balance multiple objectives such as combustion efficiency, carbon emissions, and pollutant emissions under different working conditions, and achieve the optimal control of the overall performance.

[0081] The calibration method of the carbon emission penalty term includes:

[0082] Conduct a stepped heating experiment in different temperature ranges of the combustion furnace, collect the material expansion coefficients at each temperature point, calibrate the temperature influence factor through the material expansion coefficients at each temperature point, establish a mapping relationship between the temperature deviation degree and the carbon emission penalty gain coefficient, and calibrate the carbon emission penalty term through the mapping relationship between the temperature deviation degree and the carbon emission penalty gain coefficient.

[0083] By conducting a stepped heating experiment in different temperature ranges of the combustion furnace, collecting the material expansion coefficients at each temperature point and calibrating the temperature influence factor, and establishing a mapping relationship between the temperature deviation degree and the carbon emission penalty gain coefficient, the carbon emission penalty term can accurately reflect the non-linear influence of temperature changes on carbon emissions. This calibration method can guide the intelligent agent to dynamically adjust the carbon emission penalty intensity according to the degree of the real-time temperature deviating from the optimal range during the control process, and prompt the control strategy to strictly constrain carbon emissions while pursuing combustion efficiency. By deeply coupling the physical characteristics of temperature with the carbon emission penalty mechanism, it is ensured that the system can not only operate efficiently but also achieve the goal of minimizing carbon dioxide emissions under complex working conditions, and improve the accuracy and adaptability of multi-objective optimal control.

[0084] Reference Figure 4 S3: Optimize the control strategy according to the optimization objectives, perform hierarchical progressive optimization control, and adjust the combustion equipment.

[0085] The optimization objectives include maximizing combustion efficiency, minimizing carbon dioxide emissions, and meeting pollutant control standards. Maximizing combustion efficiency can improve energy utilization efficiency, reduce costs, and enhance the economy of the system; minimizing carbon dioxide emissions helps to respond to the environmental protection requirements of energy conservation and emission reduction and conforms to the concept of sustainable development; meeting pollutant control standards can effectively reduce the harm to the environment and human health and comply with environmental protection regulations. These objectives are interrelated and restrictive, prompting the combustion-produced carbon dioxide system to comprehensively weigh various factors and dynamically adjust the control strategy through means such as data collection, model construction, intelligent agent design, and hierarchical progressive optimization control during the complex combustion process to achieve the optimal comprehensive performance of the system and achieve the unity of economic, environmental, and social benefits.

[0086] The methods for performing hierarchical progressive optimization control include:

[0087] Conduct pre-training under the constraint of the mechanism model in the virtual environment to generate an initial control strategy library;

[0088] Achieve data-driven strategy fine-tuning through a dual experience replay pool;

[0089] Dynamically relax the action space constraint according to the operation stability index;

[0090] Conducting pre-training in the virtual environment under the constraint of the mechanism model and generating an initial control strategy library enables the system to simulate different working conditions in advance and plan a reasonable control strategy at the theoretical level, providing a basic solution for actual control. Using the dual experience replay pool for data-driven strategy fine-tuning can optimize the initial strategy based on actual operation data and make the control strategy fit the actual operation conditions of the system. Dynamically relaxing the action space constraint according to the operation stability index allows the system to appropriately relax the control range during stable operation to more flexibly explore better control strategies. Generally speaking, this hierarchical progressive optimization control method gradually improves and flexibly adjusts the system control strategy from theory to practice, helping the system efficiently achieve multiple objectives such as maximizing combustion efficiency, minimizing carbon dioxide emissions, and meeting pollutant control standards.

[0091] The dual experience replay pool includes a physical compliance pool and an actual optimization pool; the physical compliance pool is used to store experience samples that meet the preset constraint conditions; the actual optimization pool is used to store samples that are met during the operation process; among them, the sampling priority of the dual experience replay pool is dynamically determined by the weighted fusion result of the physical compliance score and the reward value.

[0092] The physical compliance pool stores empirical samples that meet the preset constraint conditions, ensuring that the control strategy complies with the combustion physical rules and avoiding infeasible operations; the actual optimization pool stores operation process samples, focusing on the optimization requirements in real operation scenarios. At the same time, based on the weighted fusion of the physical compliance score and the reward value, the sampling priority is dynamically determined, enabling the system to follow physical constraints (such as equipment operation limits, combustion mechanism rules) during the learning process and taking into account multi-objective optimization such as combustion efficiency and carbon emissions (guided by the reward value), promoting the dynamic balance between "compliance" and "optimality" of the control strategy, enhancing the adaptability of the agent to complex combustion conditions, and helping the system to achieve multi-objective optimal control more efficiently.

[0093] The method for dynamically relaxing the action space constraint includes:

[0094] When the thermal efficiency volatility in N consecutive control cycles is not higher than the preset volatility threshold, update the constraint parameters;

[0095] The method for updating the constraint parameters includes:

[0096] According to the operation stability index, adaptively relax and adjust the action space constraint parameters, and when the performance index reaches the preset index threshold, proportionally expand the allowable control quantity change range.

[0097] Updating the constraint parameters when the thermal efficiency volatility in N consecutive control cycles is not higher than the preset volatility threshold can determine whether there are conditions to relax the control range based on the system operation stability. By adaptively relaxing and adjusting the constraint parameters according to the operation stability index, the control strategy can better adapt to different operation states of the system. And when the performance index reaches the preset index threshold, proportionally expanding the allowable control quantity change range can enable the system to more flexibly explore the optimization space on the basis of stable operation, further tap the potential of improving combustion efficiency, reducing carbon dioxide emissions and controlling pollutants while ensuring the system stability, helping the system to achieve a better balance among multiple objectives and achieve better optimal control effects.

[0098] Embodiment 2

[0099] Please refer to Figure 3As shown, this embodiment provides a multi-index joint relaxation rule for the refined adjustment of dynamic constraints for relaxation constraints; a multi-index joint relaxation rule is constructed to dynamically adjust the constraint parameters by comprehensively scoring combustion efficiency, carbon emissions, pollutant concentration, etc., and their corresponding weights; the multi-index joint relaxation rule breaks through the relaxation limitation of a single thermal efficiency index, and calculates the constraint adjustment range based on the weighted calculation of multi-objective comprehensive scores, making the dynamic constraint relaxation more comprehensive and scientific, and avoiding the failure of local optimization; at the same time, combined with the safety boundary prediction model to predict the relaxation risk, on the basis of ensuring the safe operation of the system, the coordination and optimization ability of the control strategy for multiple objectives is improved, and the dynamic balance and overall optimization of objectives such as combustion efficiency, carbon emissions, and pollutant control are realized.

[0100] A safety boundary prediction model is also introduced to predict the risk threshold after constraint relaxation; introducing a safety boundary prediction model to predict the risk threshold after constraint relaxation can estimate in advance the risks that the system may face in terms of combustion efficiency, carbon emissions, pollutant emissions, etc. after relaxing the control parameter range (constraint relaxation). By accurately predicting the risk threshold, problems such as excessive carbon emissions and out-of-control pollutant emissions caused by over-relaxing constraints can be avoided; at the same time, it can also prevent missing optimization opportunities due to being too conservative; the safety boundary prediction model enables the system to always maintain within a safe and controllable operating range during the pursuit of multi-objective optimization, balance the relationship between optimization and risk, and help the carbon dioxide combustion system achieve efficient and environmentally friendly multi-objective coordinated optimization control on the premise of safety.

[0101] As described above, this is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, and all should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

[0102] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. The intelligent control method for multi-objective optimization of a carbon dioxide combustion production system, characterized in that, Including: S1: Construct a dual-driven model coordinated by a combustion mechanism model and a combustion data model; S2: Use the dual-driven model to simulate the waste combustion process, and design a mechanism-constrained deep reinforcement learning agent according to the dual-driven model; S3: Optimize the control strategy according to the optimization objective, perform hierarchical progressive optimization control, and adjust the combustion equipment; The construction method of the mechanism-constrained deep reinforcement learning agent includes: Define the state space as a vector of combustion parameters collected in real time, including environmental parameters and waste characteristic parameters; Define the action space as a set of adjustable parameters of the combustion equipment, including the oxygen supply adjustment amount, the grate speed increment, and the continuous control amount of the secondary air damper opening; Construct an Actor-Critic network architecture containing a physical feasibility verification gate, where the Actor network is used to generate continuous control actions, and the Critic network is used to design a multi-objective reward function and evaluate multi-objective rewards; The design method of the multi-objective reward function includes: Construct an efficiency improvement term, a carbon emission penalty term, and a pollutant over-standard gradient penalty term by combining the combustion efficiency, carbon emissions, and pollutant concentration through dynamic weight coefficients, and obtain the multi-objective reward function according to the efficiency improvement term, the carbon emission penalty term, and the pollutant over-standard gradient penalty term. Among them, the carbon emission penalty intensity increases non-linearly with the temperature deviation from the optimal interval; The method of performing hierarchical progressive optimization control includes: Perform pre-training under the constraint of the mechanism model in a virtual environment to generate an initial control strategy library; Implement data-driven policy fine-tuning through a dual experience replay pool; Dynamically relax the action space constraint according to the operation stability index; The dual experience replay pool includes a physical compliance pool and an actual optimization pool; the physical compliance pool is used to store experience samples that meet the preset constraint conditions; the actual optimization pool is used to store samples that are met during the operation process.

2. The multi-objective optimization intelligent control method for a carbon dioxide combustion production system according to claim 1, wherein The construction steps of the combustion mechanism model include: Step S101: Analyze the waste characteristics and determine the characteristic parameters; The characteristic parameters include physical characteristic parameters and chemical characteristic parameters; the physical characteristic parameters include moisture content and thermal conductivity, and the chemical characteristic parameters include calorific value, ash content, combustible content, fixed carbon, and volatile matter; Step S102: Analyze the reaction mechanism of waste combustion and construct a combustion mechanism model; The combustion mechanism model consists of a pyrolysis model, a combustion model, a flue gas generation model, and a pollutant generation model; Step S103: Use numerical calculation methods to simulate the waste combustion process, and optimize the combustion mechanism model according to the simulation results; Convert the equations of each model in the combustion mechanism model into discrete mathematical forms; set the initial conditions; use numerical calculation methods to solve the discrete equations, simulate the evolution process of each stage during the waste combustion process; perform post-processing on the simulation results; optimize the model parameters of the combustion mechanism model according to the difference between the post-processed simulation results and the characteristic parameters.

3. The multi-objective optimization intelligent control method for a combustion carbon dioxide production system according to claim 2, characterized in that The construction process of the pyrolysis model is as follows: Determine the reaction mechanism of the pyrolysis model, and the reaction mechanism of the pyrolysis model is to describe the decomposition reaction of garbage at different temperatures; Select the pyrolysis model; Establish mass and energy balance equations to simulate the reaction kinetics during pyrolysis; Calibrate the pyrolysis model through corresponding characteristic parameters; The construction process of the combustion model is as follows: Determine the reaction mechanism of the combustion model, and the reaction mechanism of the combustion model is to describe the combustion reaction; Select the combustion model; Establish mass, energy, and momentum balance equations for the combustion process to simulate the conversion of reactants and products; Calibrate the combustion model through corresponding characteristic parameters; The construction process of the flue gas generation model is as follows: Determine the reaction mechanism of the flue gas generation model, and the reaction mechanism of the flue gas generation model is to describe the generation and conversion process of each component in the flue gas; Select the flue gas generation model; Establish mass balance and chemical reaction equations for flue gas generation; Calibrate the flue gas generation model through corresponding characteristic parameters; The construction process of the pollutant generation model is as follows: Determine the reaction mechanism of the pollutant generation model, and the reaction mechanism of the pollutant generation model is to describe the chemical reaction path of pollutant generation; Select the pollutant generation model; Establish chemical reaction equations for pollutant generation; Calibrate the pollutant generation model through corresponding characteristic parameters.

4. The multi-objective optimization intelligent control method for a carbon dioxide combustion production system according to claim 1, wherein The implementation method of the physical feasibility verification gate includes: adding a constraint verification module to the output layer of the Critic network, and the constraint verification module non-linearly fuses the predicted Q value and the physical feasibility coefficient through a neural network layer to output an action value evaluation result corrected by physical constraints.

5. The multi-objective optimization intelligent control method for a combustion-produced carbon dioxide system according to claim 1, characterized in that The adjustment method of the dynamic weight coefficient includes: Taking the operation data within a preset time period as the input of the stage identification model to obtain the current stage, and the current stage includes a startup stage, a steady state stage, and a transient stage; In the startup stage, the dynamic weight coefficient corresponding to the combustion efficiency is increased to a preset weight threshold, and according to the constraint condition that the sum of all dynamic weight coefficients is 1, the dynamic weight coefficients corresponding to the carbon emission penalty term and the pollutant over-standard gradient penalty term are randomly allocated; In the steady state stage, the dynamic weight coefficients corresponding to the carbon emission penalty term and the pollutant over-standard gradient penalty term are dynamically allocated according to the pollutant emission quota; In the transient stage, an emergency adjustment factor is introduced to adjust the dynamic weight coefficient corresponding to the pollutant over-standard gradient penalty term.

6. The multi-objective optimization intelligent control method for a carbon dioxide combustion production system according to claim 1, wherein The calibration method of the carbon emission penalty term includes: Conduct a stepwise heating experiment in different temperature intervals of the combustion furnace, collect the material expansion coefficients at each temperature point, calibrate the temperature influence factor through the material expansion coefficients at each temperature point, establish a mapping relationship between the temperature deviation degree and the carbon emission penalty gain coefficient, and calibrate the carbon emission penalty term through the mapping relationship between the temperature deviation degree and the carbon emission penalty gain coefficient.

7. The intelligent control method for multi-objective optimization of a combustion-produced carbon dioxide system according to claim 6, wherein The optimization objectives of the optimization control strategy for the optimization objective include maximizing the combustion efficiency, minimizing the carbon dioxide emissions, and meeting the pollutant control standards.

8. The intelligent control method for multi-objective optimization of a carbon dioxide combustion production system according to claim 1, wherein The sampling priority of the dual experience replay pool is dynamically determined by the weighted fusion result of the physical compliance score and the reward value.

9. The intelligent control method for multi-objective optimization of a carbon dioxide combustion production system according to claim 1, characterized in that, The method of dynamically relaxing the action space constraint includes: When the thermal efficiency volatility over N consecutive control cycles is not higher than the preset volatility threshold, update the constraint parameters; The method for updating the constraint parameters includes: Adaptive relaxation adjustment is performed on the action space constraint parameters according to the operation stability index, and when the performance index reaches the preset index threshold, the allowable control quantity change range is proportionally expanded.

Citation Information

Patent Citations

  • System and method for real-time prediction of calorific value of circulating fluidized bed domestic waste incineration boiler

    CN105864797B

  • Intelligent control system of waste incineration power plant and waste incineration system

    CN118089034A

  • Intelligent boiler combustion optimization method based on module-level mechanism and mathematical hybrid model

    CN118066563A

  • Multi-agent synergistic compatibility method based on Copula thought

    CN119444146A