Multi-objective optimization intelligent control method for combustion carbon dioxide production system

By building a dual-drive model and designing mechanism constraint deep reinforcement learning agent, combining multi-objective reward function and dynamic weight coefficient, multi-objective optimization intelligent control of combustion carbon dioxide production systems is achieved, solving the problems of incomplete combustion and instability of traditional systems, and improving the control accuracy and stability of the system.

CN119983287AActive Publication Date: 2025-05-13CHN ENERGY JIANGSU POWER CO LTD +2

Patent Information

Application Number
CN202510480882.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-13
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

When traditional garbage combustion control systems deal with complex garbage components and calorific value fluctuations, the prediction accuracy is not high, making it difficult to achieve optimal control of the system, resulting in incomplete combustion, unstable emissions and low energy conversion rate.

Method used

A multi-objective optimization intelligent control method for combustion carbon dioxide production systems is proposed. By constructing a dual-drive model that is coordinated by combustion mechanism model and combustion data model, a mechanism-constrained deep reinforcement learning agent is designed, and a multi-objective reward function and dynamic weight coefficient are combined to perform hierarchical progressive optimization control.

Benefits of technology

The combustion efficiency is maximized, carbon dioxide emissions are minimized and pollutant control meets standards, which improves the accuracy, intelligence and stability of system control, and solves the problems of incomplete combustion and instability of traditional systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119983287A_ABST
    Figure CN119983287A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of equipment intelligent control, and discloses a multi-objective optimization intelligent control method for a combustion carbon dioxide production system. The method comprises the following steps: constructing a combustion mechanism model; training a combustion data model; fusing the combustion mechanism model and the combustion data model to obtain a dual-drive model, and determining an optimization target; a dual-drive model is adopted to simulate the garbage combustion process, the control strategy is optimized according to the optimization target, and combustion equipment is adjusted; the advantages of the two models are effectively fused, advantage complementation between the models is achieved, and therefore the physical interpretability and adaptability of the models are improved; in the control strategy optimization process, not only is the combustion efficiency considered, but also carbon emission and pollutant emission are considered, the combustion efficiency is improved, and the carbon emission and the pollutant emission are reduced, so that multi-target optimization of the combustion carbon dioxide production system is achieved, and the control strategy optimization efficiency is improved. And the environment-friendly performance and the economic performance of the garbage combustion carbon dioxide production system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of equipment intelligent control, and more specifically, to a multi-objective optimization intelligent control method for a combustion carbon dioxide production system. Background Art

[0002] As an effective treatment method, the incineration of domestic waste is widely used because it can significantly reduce the volume of waste, eliminate harmful substances and recover energy. However, the incineration process of waste will produce a large amount of pollutants, such as carbon dioxide (CO2), nitrogen oxides (NO x ), dioxins, etc. The emission of these pollutants not only puts pressure on the environment, but also poses a potential threat to human health. In addition, carbon emissions during the combustion process also have a direct impact on global climate change. Therefore, how to ensure efficient combustion of garbage while achieving effective control of pollutant emissions and coordinated reduction of carbon emissions has become a key issue that needs to be solved urgently.

[0003] Traditional waste combustion control systems mainly rely on physical and chemical mechanism models to predict and adjust key parameters in the combustion process, such as the real-time prediction system and method for the calorific value of circulating fluidized bed municipal solid waste incineration boilers disclosed in the Chinese patent with authorization announcement number CN105864797B; however, due to the complex composition of municipal solid waste and large fluctuations in calorific value, these mechanism models often face challenges such as insufficient adaptability and low prediction accuracy in practical applications.

[0004] With the development of information technology and big data technology, data-driven modeling methods are gradually widely used. By learning from a large amount of historical data, they can better handle the nonlinear behavior of the system and adapt to changing operating conditions, such as the intelligent control system of a waste incineration power plant and a waste incineration carbon dioxide production system disclosed in the Chinese patent application with publication number CN118089034A; however, the composition of waste is complex and changeable, which makes it difficult to accurately grasp the physical and chemical characteristics of the combustion process and to achieve optimal control of the system; dynamic constraint relaxation based only on thermal efficiency indicators may lead to local optimization failure and cannot fully guarantee the stable operation and multi-objective optimization of the system under different operating conditions; therefore, there is an urgent need for an intelligent control method that combines a mechanism model and a data-driven model.

[0005] In view of this, the present invention proposes a multi-objective optimization intelligent control method for a combustion carbon dioxide production system to solve the above problems. Summary of the invention

[0006] In order to overcome the above-mentioned defects of the prior art and achieve the above-mentioned purpose, the present invention provides the following technical solution: a multi-objective optimization intelligent control method for a combustion carbon dioxide production system, comprising:

[0007] S1: Construct a dual-driven model that is coordinated by the combustion mechanism model and the combustion data model;

[0008] S2: Use a dual-drive model to simulate the garbage combustion process and design a mechanism-constrained deep reinforcement learning agent based on the dual-drive model;

[0009] S3: Optimize the control strategy according to the optimization target, perform hierarchical progressive optimization control, and adjust the combustion equipment.

[0010] Furthermore, the step of constructing the combustion mechanism model includes:

[0011] Step S101: Analyze garbage characteristics and determine characteristic parameters;

[0012] The characteristic parameters include physical characteristic parameters and chemical characteristic parameters; the physical characteristic parameters include moisture content and thermal conductivity, and the chemical characteristic parameters include calorific value, ash content, combustible content, fixed carbon and volatile matter;

[0013] Step S102: analyzing the reaction mechanism of garbage combustion and constructing a combustion mechanism model;

[0014] The combustion mechanism model consists of pyrolysis model, combustion model, flue gas generation model, and pollutant generation model;

[0015] Step S103: using a numerical calculation method to simulate the garbage combustion process, and optimizing the combustion mechanism model according to the simulation results;

[0016] The equations of each model in the combustion mechanism model are converted into discrete mathematical forms; initial conditions are set; numerical calculation methods are used to solve the discretized equations to simulate the evolution of each stage in the garbage combustion process; the simulation results are post-processed; and the model parameters of the combustion mechanism model are optimized based on the difference between the simulation results after post-processing operations and the characteristic parameters.

[0017] Furthermore, the construction process of the pyrolysis model is as follows: determining the reaction mechanism of the pyrolysis model, the reaction mechanism of the pyrolysis model is to describe the decomposition reaction of garbage at different temperatures; selecting the pyrolysis model; establishing the mass and energy balance equations to simulate the reaction kinetics during the pyrolysis process; calibrating the pyrolysis model through the corresponding characteristic parameters;

[0018] The construction process of the combustion model is as follows: determining the reaction mechanism of the combustion model, the reaction mechanism of the combustion model is to describe the main combustion reaction; selecting the combustion model; establishing the mass, energy and momentum balance equations of the combustion process to simulate the conversion of reactants and products; calibrating the combustion model through corresponding characteristic parameters;

[0019] The smoke generation model construction process is as follows: determining the reaction mechanism of the smoke generation model, the reaction mechanism of the smoke generation model is to describe the generation and transformation process of each component in the smoke; selecting the smoke generation model; establishing the mass balance and chemical reaction equation of the smoke generation; calibrating the smoke generation model through the corresponding characteristic parameters;

[0020] The construction process of the pollutant generation model is: determining the reaction mechanism of the pollutant generation model, the reaction mechanism of the pollutant generation model is to describe the chemical reaction path of pollutant generation; selecting the pollutant generation model; establishing the chemical reaction equation of pollutant generation; calibrating the pollutant generation model through corresponding characteristic parameters.

[0021] Furthermore, the method for constructing the mechanism-constrained deep reinforcement learning agent includes:

[0022] The state space is defined as the combustion parameter vector collected in real time, including environmental parameters and characteristic parameters of garbage;

[0023] The action space is defined as a set of adjustable parameters of the combustion equipment, including the oxygen supply adjustment amount, the grate speed increment and the continuous control amount of the secondary air door opening;

[0024] Construct an Actor-Critic network architecture with a physical feasibility check gate, where the Actor network is used to generate continuous control actions, and the Critic network is used to design multi-objective reward functions and evaluate multi-objective rewards;

[0025] The design method of the multi-objective reward function includes:

[0026] The efficiency improvement item, carbon emission penalty item and pollutant exceeding gradient penalty item are constructed by combining combustion efficiency, carbon emissions and pollutant concentration through dynamic weight coefficients. The multi-objective reward function is obtained based on the efficiency improvement item, carbon emission penalty item and pollutant exceeding gradient penalty item. Among them, the carbon emission penalty intensity increases nonlinearly as the temperature deviates from the optimal range.

[0027] Furthermore, the implementation method of the physical feasibility verification gate includes: adding a constraint verification module to the output layer of the Critic network, the constraint verification module performs nonlinear fusion of the predicted Q value and the physical feasibility coefficient through the neural network layer, and outputs the action value evaluation result corrected by the physical constraint.

[0028] Furthermore, the method for adjusting the dynamic weight coefficient includes:

[0029] The operating data within a preset time period is used as the input of the stage identification model to obtain the current stage, which includes the startup stage, the steady-state stage and the transient stage; in the startup stage, the dynamic weight coefficient corresponding to the combustion efficiency is increased to the preset weight threshold, and the dynamic weight coefficients corresponding to the carbon emission penalty item and the gradient penalty item for exceeding the pollutant standard are randomly allocated according to the constraint that the sum of all dynamic weight coefficients is 1; in the steady-state stage, the dynamic weight coefficients corresponding to the carbon emission penalty item and the gradient penalty item for exceeding the pollutant standard are dynamically allocated according to the pollutant emission quota; in the transient stage, an emergency adjustment factor is introduced to adjust the dynamic weight coefficient corresponding to the gradient penalty item for exceeding the pollutant standard.

[0030] Furthermore, the calibration method of the carbon emission penalty item includes:

[0031] A step-by-step heating experiment was carried out in different temperature ranges of the combustion furnace. The material expansion coefficient at each temperature point was collected. The temperature influence factor was calibrated through the material expansion coefficient at each temperature point. The mapping relationship between the temperature deviation and the carbon emission penalty gain coefficient was established. The carbon emission penalty item was obtained by calibrating the mapping relationship between the temperature deviation and the carbon emission penalty gain coefficient.

[0032] Furthermore, the optimization objectives of the optimization target optimization control strategy include maximizing combustion efficiency, minimizing carbon dioxide emissions and achieving pollutant control standards.

[0033] Furthermore, the method for performing hierarchical progressive optimization control includes:

[0034] Conduct pre-training under the constraints of the mechanism model in a virtual environment to generate an initial control strategy library;

[0035] Data-driven strategy fine-tuning through dual experience replay pools;

[0036] Dynamically relax action space constraints according to operational stability indicators;

[0037] The dual experience replay pool includes a physical compliance pool and an actual optimization pool; the physical compliance pool is used to store experience samples that meet preset constraints; the actual optimization pool is used to store samples that meet the constraints during operation; wherein the sampling priority of the dual experience replay pool is dynamically determined by the weighted fusion result of the physical compliance score and the reward value.

[0038] Furthermore, the method for dynamically relaxing the action space constraints includes:

[0039] When the thermal efficiency fluctuation rate for N consecutive control cycles is not higher than the preset fluctuation rate threshold, the constraint parameters are updated;

[0040] Methods for updating constraint parameters include:

[0041] The action space constraint parameters are adaptively relaxed and adjusted according to the operation stability index. When the performance index reaches the preset index threshold, the allowable control amount change range is proportionally expanded.

[0042] The technical effects and advantages of the multi-objective optimization intelligent control method of the combustion carbon dioxide production system of the present invention are as follows:

[0043] By constructing a dual-drive model to accurately simulate the combustion process, the foundation for intelligent control is laid; the mechanism-constrained deep reinforcement learning agent based on this design can generate reasonable control strategies based on multi-objective reward functions to maximize combustion efficiency, minimize carbon dioxide emissions, and achieve pollutant control standards. The physical feasibility verification gate ensures that the control action conforms to physical laws, and the dynamic weight coefficient is adjusted according to different stages to make the multi-objective optimization more targeted. The carbon emission penalty item is accurately calibrated to strictly constrain carbon emissions. In the hierarchical progressive optimization control, the virtual environment pre-training generates the initial strategy library, the dual experience replay pool fine-tunes the strategy, and the dynamic relaxation of the action space constraint enables the system to flexibly explore the optimization space. The combined effect of these technical means has improved the accuracy, intelligence and stability of the system control, realized multi-objective collaborative optimization, and effectively solved the problems of incomplete combustion, unstable emissions, and low energy conversion rate in traditional combustion systems, promoting efficient and green operation of combustion and carbon dioxide production systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 This is a flow chart of a multi-objective optimization intelligent control method for a combustion carbon dioxide production system according to Example 1 of the present invention;

[0045] Figure 2 This is a schematic diagram of the process flow of the deep reinforcement learning agent design method according to Embodiment 1 of the present invention;

[0046] Figure 3 This is a schematic diagram of a flow chart of a method for adjusting a dynamic weight coefficient according to Embodiment 1 of the present invention;

[0047] Figure 4 Schematic diagram of the hierarchical progressive optimization control flow of Example 1 of the present invention. DETAILED DESCRIPTION

[0048] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0049] Example 1

[0050] See also Figure 1As shown, the multi-objective optimization intelligent control method for the combustion carbon dioxide production system described in this embodiment includes:

[0051] S1: Construct a dual-driven model that is coordinated by the combustion mechanism model and the combustion data model.

[0052] The combustion mechanism model is a mathematical model that describes the combustion process based on the physical and chemical laws of the garbage combustion process and uses basic theories such as thermodynamics, fluid mechanics, and chemical reaction kinetics.

[0053] The steps to build a combustion mechanism model include:

[0054] Step S101: Analyze garbage characteristics, determine characteristic parameters, and provide necessary input data for the combustion mechanism model.

[0055] The characteristic parameters include physical characteristic parameters and chemical characteristic parameters; the physical characteristic parameters include at least moisture content, thermal conductivity, etc., and the chemical characteristic parameters include at least calorific value, ash content, combustible content, fixed carbon, volatile matter, etc.; the characteristic parameters of different garbage are obtained by technicians in this field through garbage combustion experiments;

[0056] The moisture content is the proportion of moisture in the garbage; the thermal conductivity is the thermal conductivity of the garbage; the calorific value is the heat released when the garbage is completely burned; the ash content is the proportion of inorganic residues after the garbage is completely burned; the combustible content is the proportion of components that can be burned in the garbage; the fixed carbon is the remaining carbon content after removing moisture and volatile organic matter from the garbage; and the volatile component is the gas component released by the garbage during the combustion process.

[0057] Step S102: Analyze the reaction mechanism of garbage combustion and construct a combustion mechanism model.

[0058] The combustion mechanism model is composed of at least a pyrolysis model, a combustion model, a flue gas generation model, and a pollutant generation model.

[0059] The construction process of the pyrolysis model is as follows: determine the reaction mechanism of the pyrolysis model, which describes the decomposition reaction of garbage at different temperatures, such as volatile matter release, solid carbonization process, etc.; select a pyrolysis model, such as an empirical model (such as TGA data), a chemical kinetic model (such as the Arrhenius equation), a multiphase reaction model, etc.; establish a mass and energy balance equation to simulate the reaction kinetics during the pyrolysis process; calibrate the pyrolysis model through corresponding characteristic parameters to ensure that the pyrolysis model accurately reflects the pyrolysis behavior of the garbage; the mass and energy balance equation is established by technical personnel in this field according to the actual pyrolysis experimental conditions.

[0060] The construction process of the combustion model is as follows: determine the reaction mechanism of the combustion model, which describes the main combustion reactions, such as the oxidation reaction of organic matter, the combustion of gas and solid fuels, the generation of nitrogen oxides, the oxidation of sulfides, etc.; select a combustion model, such as fluid mechanics, reaction kinetics model (such as chemical reaction network, reaction kinetics equation), etc.; establish the mass, energy and momentum balance equations of the combustion process to simulate the transformation of reactants and products; calibrate the combustion model through the corresponding characteristic parameters to ensure that the combustion model accurately reflects the combustion process of garbage; the mass, energy and momentum balance equations of the combustion process are established by technical personnel in this field according to the actual combustion experiment conditions.

[0061] The construction process of the smoke generation model is as follows: determine the reaction mechanism of the smoke generation model, which describes the generation and transformation process of each component in the smoke, such as the chemical reaction after combustion; select the smoke generation model, such as the smoke flow model, the chemical reaction model (such as the reaction kinetics model, the reaction mechanism model), etc.; establish the mass balance and chemical reaction equation of smoke generation, taking into account gas diffusion; calibrate the smoke generation model through the corresponding characteristic parameters to ensure that the smoke generation model accurately reflects the smoke generation process of garbage; the mass balance and chemical reaction equation of smoke generation are established by technical personnel in this field according to the actual smoke generation experiment conditions.

[0062] The construction process of the pollutant generation model is as follows: determine the reaction mechanism of the pollutant generation model, which is to describe the chemical reaction path of pollutant generation, such as the formation of chlorides, the release of heavy metals, etc.; select the pollutant generation model, such as the chemical reaction network, the poison release model, etc.; establish the chemical reaction equation for pollutant generation, which takes into account the source and transformation path of pollutant generation; calibrate the pollutant generation model through corresponding characteristic parameters to ensure that the pollutant generation model accurately reflects the pollutant generation process of garbage; the chemical reaction equation for pollutant generation is established by technical personnel in this field according to the actual pollutant generation experiment.

[0063] Step S103: using numerical calculation methods to simulate the garbage combustion process, and optimizing the combustion mechanism model according to the simulation results.

[0064] The equations of each model in the combustion mechanism model are converted into discrete mathematical forms so that they can be numerically solved; initial conditions are set, such as combustion temperature, combustion pressure, feed rate, etc. The initial conditions are the initial data of the garbage combustion experiment when the characteristic parameters are obtained as mentioned above; numerical calculation methods are used to solve the discretized equations to simulate the evolution process of each stage in the garbage combustion process, and numerical calculation methods such as finite element difference method, finite element method, finite volume method, etc.; the simulation results are post-processed, such as calculating thermal conductivity, calorific value, ash content, etc.; according to the difference between the simulation results after the post-processing operation and the characteristic parameters, the model parameters of the combustion record model are optimized, such as reaction kinetic constants, reaction rate constants, etc.

[0065] The combustion data model adopts an online learning neural network structure. The input layer embeds the output characteristics of the mechanism model and dynamically corrects the strategy deviation of the combustion mechanism model output.

[0066] It should be noted that the combustion mechanism model simulates the pyrolysis, combustion, flue gas generation and pollutant generation stages in the garbage combustion process by analyzing the garbage characteristics and combustion reaction mechanism, which can deeply reveal the physical and chemical nature of the combustion process and provide a theoretical basis for the control strategy; the combustion data model is based on a large amount of actual operation data and can reflect the complex nonlinear relationships and uncertainties in the combustion process; the collaborative dual-drive model of the two combines the interpretability of the mechanism model and the adaptability of the data model, providing a more accurate simulation environment for designing mechanism-constrained deep reinforcement learning agents.

[0067] Reference Figure 2 , S2: A dual-drive model is used to simulate the garbage combustion process, and a mechanism-constrained deep reinforcement learning agent is designed based on the dual-drive model; the mechanism-constrained deep reinforcement learning agent can generate an effective control strategy based on the dual-drive model, and perform hierarchical progressive optimization control while considering multiple objectives such as maximizing combustion efficiency, minimizing carbon dioxide emissions, and meeting pollutant control standards, so as to achieve precise adjustment of the combustion equipment, thereby improving the overall performance of the entire combustion and carbon dioxide production system.

[0068] The construction method of mechanism-constrained deep reinforcement learning agent includes:

[0069] The state space is defined as the combustion parameter vector collected in real time, including environmental parameters and characteristic parameters of garbage;

[0070] The action space is defined as a set of adjustable parameters of the combustion equipment, including the oxygen supply adjustment amount, the grate speed increment and the continuous control amount of the secondary air door opening;

[0071] Construct an Actor-Critic network architecture with a physical feasibility check gate, where the Actor network is used to generate continuous control actions, and the Critic network is used to design multi-objective reward functions and evaluate multi-objective rewards;

[0072] By defining the combustion parameter vector collected in real time as the state space, the intelligent agent can fully perceive the actual situation of the combustion process and master the basic information of the system operation; the adjustable parameter set of the combustion equipment is used as the action space, so that the intelligent agent can make effective adjustment actions for different states; an Actor-Critic network architecture with a physical feasibility verification gate is constructed, in which the Actor network generates continuous control actions to ensure that reasonable control instructions can be output according to the state; the Critic network designs a multi-objective reward function to evaluate multi-objective rewards, which can guide the intelligent agent to learn and optimize in the direction of maximizing combustion efficiency, minimizing carbon dioxide emissions, and meeting pollutant control standards; overall, the mechanism-constrained deep reinforcement learning agent construction method can improve the intelligence, effectiveness and rationality of system control, and realize accurate and efficient multi-objective optimization control of the combustion carbon dioxide production system.

[0073] The implementation method of the physical feasibility verification gate includes:

[0074] Add a constraint verification module to the output layer of the Critic network. The constraint verification module performs nonlinear fusion of the predicted Q value and the physical feasibility coefficient through the neural network layer, and outputs the action value evaluation result corrected by the physical constraint. ;in, is the action value evaluation result after correction by physical constraints; To predict the Q value; To verify the parameters, they can be obtained through the nature-inspired optimization algorithm; Physical feasibility coefficients calculated in real time for combustion mechanism models; It is a trainable parameter matrix that automatically learns the physical constraint weights through back-propagation;

[0075] In the multi-objective optimization intelligent control of the combustion and carbon dioxide production system, the above formula integrates the predicted Q value and the physical feasibility coefficient calculated in real time by the mechanism model, and uses the trainable parameter matrix to learn the physical constraint weights, and finally outputs the action value evaluation corrected by the physical constraints; the above formula enables the intelligent agent to consider the optimization objectives when evaluating the value of the control action, and strictly follow the physical laws of the combustion process, avoiding the generation of invalid actions that do not conform to the physical feasibility, and ensuring that the control strategy can still effectively promote multi-objective optimization under actual physical constraints, thereby improving the reliability and stability of system control and making the optimization process more in line with the actual needs of the project.

[0076] The design methods of multi-objective reward functions include:

[0077] The efficiency improvement item, carbon emission penalty item and pollutant exceeding gradient penalty item are constructed by combining combustion efficiency, carbon emissions and pollutant concentration through dynamic weight coefficients. The multi-objective reward function is obtained based on the efficiency improvement item, carbon emission penalty item and pollutant exceeding gradient penalty item. Among them, the intensity of carbon emission penalty increases nonlinearly as the temperature deviates from the optimal range. Multi-objective reward function at time ;in, The corresponding dynamic weight coefficient of combustion efficiency is the physical feasibility coefficient calculated in real time by the combustion mechanism model; for Combustion efficiency at all times; For maximum combustion efficiency; is the dynamic weight coefficient of carbon emissions; for Carbon emissions at each moment; is the carbon emission threshold; is the temperature influence factor; is a mathematical constant; is the thermal stress coefficient of the combustion furnace material; for Temperature at all times; is the temperature threshold; For the Dynamic weight coefficient of pollutant concentration; for Moment The concentration of pollutants; For the pollutant concentration thresholds; by integrating combustion efficiency improvement items, the system is encouraged to improve combustion efficiency; using carbon emission penalty items containing temperature influencing factors, the intensity of carbon emission penalty increases nonlinearly as the temperature deviates from the optimal range, accurately constraining carbon emissions; combining pollutant excess penalty items to strictly control pollutant emissions. At the same time, the physical feasibility coefficient is incorporated to ensure that the strategy conforms to the combustion mechanism. Overall, this function provides a multi-objective optimization guide for the intelligent agent, prompting the control strategy to strike a balance between maximizing combustion efficiency, minimizing carbon dioxide emissions, and controlling pollutants to meet standards, ensuring that the system achieves comprehensive performance optimization within the physically feasible range.

[0078] Reference Figure 3 , the adjustment methods of dynamic weight coefficient include:

[0079] The operating data within a preset time period is used as the input of the stage identification model to obtain the current stage, which includes the startup stage, the steady-state stage and the transient stage; in the startup stage, the dynamic weight coefficient corresponding to the combustion efficiency is increased to the preset weight threshold, and the dynamic weight coefficients corresponding to the carbon emission penalty item and the gradient penalty item for exceeding the pollutant standard are randomly allocated according to the constraint that the sum of all dynamic weight coefficients is 1; in the steady-state stage, the dynamic weight coefficients corresponding to the carbon emission penalty item and the gradient penalty item for exceeding the pollutant standard are dynamically allocated according to the pollutant emission quota; in the transient stage, an emergency adjustment factor is introduced to adjust the dynamic weight coefficient corresponding to the gradient penalty item for exceeding the pollutant standard.

[0080] By inputting the operation data of the preset period into the stage identification model (such as hidden Markov model and support vector machine, etc.), it can accurately determine whether the system is in the startup, steady state or transient stage. In the startup stage, increasing the combustion efficiency weight to the preset threshold and randomly assigning the carbon emission penalty item and the pollutant exceeding the standard gradient penalty item weight can quickly improve the system combustion efficiency to achieve a stable operating state; in the steady-state stage, dynamically assigning relevant weights according to the pollutant emission quota can enable the system to continuously optimize carbon emissions and pollutant emission control during stable operation; in the transient stage, introducing emergency adjustment factors to adjust the weight of the pollutant exceeding the standard gradient penalty item can allow the system to respond to emergencies in a timely manner and quickly control pollutant emissions. This method of dynamically adjusting the weight coefficient according to different stages enables the system to flexibly balance multiple objectives such as combustion efficiency, carbon emissions and pollutant emissions under different operating conditions, and achieve optimal control of overall performance.

[0081] The calibration methods for carbon emission penalty items include:

[0082] A step-by-step heating experiment was carried out in different temperature ranges of the combustion furnace. The material expansion coefficient at each temperature point was collected. The temperature influence factor was calibrated through the material expansion coefficient at each temperature point. The mapping relationship between the temperature deviation and the carbon emission penalty gain coefficient was established. The carbon emission penalty item was obtained by calibrating the mapping relationship between the temperature deviation and the carbon emission penalty gain coefficient.

[0083] By conducting step-by-step heating experiments in different temperature ranges of the combustion furnace, collecting the material expansion coefficient at each temperature point and calibrating the temperature influence factor, a mapping relationship between the temperature deviation and the carbon emission penalty gain coefficient is established, so that the carbon emission penalty term can accurately reflect the nonlinear impact of temperature changes on carbon emissions. This calibration method can guide the intelligent agent to dynamically adjust the intensity of carbon emission penalties according to the degree of deviation of the real-time temperature from the optimal range during the control process, prompting the control strategy to strictly constrain carbon emissions while pursuing combustion efficiency. By deeply coupling the physical properties of temperature with the carbon emission penalty mechanism, it is ensured that the system can operate efficiently and achieve the goal of minimizing carbon dioxide emissions under complex working conditions, thereby improving the accuracy and adaptability of multi-objective optimization control.

[0084] Reference Figure 4 , S3: Optimize the control strategy according to the optimization target, perform hierarchical progressive optimization control, and adjust the combustion equipment.

[0085] The optimization goals include maximizing combustion efficiency, minimizing carbon dioxide emissions, and achieving pollutant control standards. Maximizing combustion efficiency can improve energy efficiency, reduce costs, and enhance the economy of the system; minimizing carbon dioxide emissions helps to respond to environmental protection requirements for energy conservation and emission reduction, and is in line with the concept of sustainable development; achieving pollutant control standards can effectively reduce harm to the environment and human health and comply with environmental regulations. These goals are interrelated and mutually constrained, prompting the combustion and carbon dioxide production system to comprehensively weigh various factors and dynamically adjust the control strategy in the complex combustion process through data collection, model construction, intelligent agent design, and hierarchical progressive optimization control, so as to achieve the best overall system performance and achieve the unity of economic, environmental and social benefits.

[0086] The method of performing hierarchical progressive optimization control includes:

[0087] Conduct pre-training under the constraints of the mechanism model in a virtual environment to generate an initial control strategy library;

[0088] Data-driven strategy fine-tuning through dual experience replay pools;

[0089] Dynamically relax action space constraints according to operational stability indicators;

[0090] Carrying out pre-training based on mechanism model constraints in a virtual environment and generating an initial control strategy library can enable the system to simulate different working conditions in advance, plan reasonable control strategies at the theoretical level, and provide a basic solution for actual control. Using a dual experience replay pool for data-driven strategy fine-tuning, the initial strategy can be optimized based on actual operating data to make the control strategy fit the actual operating conditions of the system. Dynamically relaxing the action space constraints according to the operating stability index allows the system to appropriately relax the control range during stable operation to more flexibly explore better control strategies. Overall, this hierarchical progressive optimization control method enables the system control strategy to be gradually improved and flexibly adjusted from theory to practice, helping the system to efficiently achieve multiple goals such as maximizing combustion efficiency, minimizing carbon dioxide emissions, and meeting pollutant control standards.

[0091] The dual experience replay pool includes a physical compliance pool and an actual optimization pool; the physical compliance pool is used to store experience samples that meet the preset constraints; the actual optimization pool is used to store samples that meet the constraints during operation; among them, the sampling priority of the dual experience replay pool is dynamically determined by the weighted fusion result of the physical compliance score and the reward value.

[0092] The physical compliance pool stores experience samples that meet preset constraints to ensure that the control strategy complies with the physical rules of combustion and avoids infeasible operations; the actual optimization pool stores operation process samples to focus on optimization needs in real operation scenarios. At the same time, based on the weighted fusion of physical compliance scores and reward values, the sampling priority is dynamically determined, so that the system not only follows physical constraints (such as equipment operation restrictions and combustion mechanism rules) during the learning process, but also takes into account multi-objective optimization such as combustion efficiency and carbon emissions (guided by reward values), promotes the dynamic balance between "compliance" and "optimization" of the control strategy, improves the adaptability of the intelligent body to complex combustion conditions, and helps the system achieve multi-objective optimization control more efficiently.

[0093] Methods for dynamically relaxing action space constraints include:

[0094] When the thermal efficiency fluctuation rate for N consecutive control cycles is not higher than the preset fluctuation rate threshold, the constraint parameters are updated;

[0095] Methods for updating constraint parameters include:

[0096] The action space constraint parameters are adaptively relaxed and adjusted according to the operation stability index. When the performance index reaches the preset index threshold, the allowable control amount change range is proportionally expanded.

[0097] When the thermal efficiency fluctuation rate for N consecutive control cycles is not higher than the preset fluctuation rate threshold, the constraint parameters are updated, and it is possible to determine whether there are conditions for relaxing the control range based on the system operation stability. By adaptively relaxing and adjusting the constraint parameters according to the operation stability index, the control strategy can better adapt to different operating states of the system. When the performance index reaches the preset index threshold, the allowable control amount variation range is proportionally expanded, which allows the system to explore the optimization space more flexibly on the basis of stable operation. On the premise of ensuring system stability, it further explores the potential for improving combustion efficiency, reducing carbon dioxide emissions and controlling pollutants, which helps the system achieve a better balance between multiple objectives and achieve better optimization control effects.

[0098] Example 2

[0099] See also Figure 3As shown, this embodiment provides a multi-indicator joint relaxation rule for fine-tuning the dynamic constraint relaxation constraint; constructs a multi-indicator joint relaxation rule, and dynamically adjusts the constraint parameters by comprehensively evaluating the combustion efficiency, carbon emissions, pollutant concentration, and other scores, as well as the corresponding weights; the multi-indicator joint relaxation rule breaks through the relaxation limitation of a single thermal efficiency indicator, and calculates the constraint adjustment range based on a weighted calculation of the multi-objective comprehensive score, making the dynamic constraint relaxation more comprehensive and scientific, and avoiding local optimization failure; at the same time, combined with the safety boundary prediction model to predict the relaxation risk, on the basis of ensuring the safe operation of the system, improves the control strategy's ability to coordinate and optimize multiple objectives, and achieves dynamic balance and overall optimization of objectives such as combustion efficiency, carbon emissions, and pollutant control.

[0100] A safety boundary prediction model is also introduced to predict the risk threshold after constraint relaxation. The introduction of a safety boundary prediction model can predict the risk threshold after constraint relaxation, and can predict in advance the risks that the system may face in terms of combustion efficiency, carbon emissions, and pollutant emissions after relaxing the control parameter range (constraint relaxation). By accurately predicting the risk threshold, it is possible to avoid system performance deterioration due to excessive relaxation of constraints, such as excessive carbon emissions and uncontrolled pollutant emissions; at the same time, it can also prevent optimization opportunities from being missed due to being too conservative; the safety boundary prediction model can enable the system to always remain within a safe and controllable operating range in the pursuit of multi-objective optimization, balance the relationship between optimization and risk, and help the combustion and carbon dioxide production system achieve efficient and environmentally friendly multi-objective collaborative optimization control under the premise of safety.

[0101] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.

[0102] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A multi-objective optimization intelligent control method for a combustion carbon dioxide production system, characterized in that: include: S1: Construct a dual-driven model that is coordinated by the combustion mechanism model and the combustion data model; S2: Use a dual-drive model to simulate the garbage combustion process and design a mechanism-constrained deep reinforcement learning agent based on the dual-drive model; S3: Optimize the control strategy according to the optimization target, perform hierarchical progressive optimization control, and adjust the combustion equipment.

2. The multi-objective optimization intelligent control method for a combustion carbon dioxide production system according to claim 1 is characterized in that: The steps of constructing the combustion mechanism model include: Step S101: Analyze garbage characteristics and determine characteristic parameters; The characteristic parameters include physical characteristic parameters and chemical characteristic parameters; the physical characteristic parameters include moisture content and thermal conductivity, and the chemical characteristic parameters include calorific value, ash content, combustible content, fixed carbon and volatile matter; Step S102: analyzing the reaction mechanism of garbage combustion and constructing a combustion mechanism model; The combustion mechanism model consists of pyrolysis model, combustion model, flue gas generation model, and pollutant generation model; Step S103: using a numerical calculation method to simulate the garbage combustion process, and optimizing the combustion mechanism model according to the simulation results; The equations of each model in the combustion mechanism model are converted into discrete mathematical forms; initial conditions are set; numerical calculation methods are used to solve the discretized equations to simulate the evolution of each stage in the garbage combustion process; the simulation results are post-processed; and the model parameters of the combustion mechanism model are optimized based on the difference between the simulation results after post-processing operations and the characteristic parameters.

3. The multi-objective optimization intelligent control method for a combustion carbon dioxide production system according to claim 2 is characterized in that: The construction process of the pyrolysis model is as follows: determining the reaction mechanism of the pyrolysis model, which describes the decomposition reaction of garbage at different temperatures; selecting the pyrolysis model; establishing mass and energy balance equations to simulate the reaction kinetics during the pyrolysis process; and calibrating the pyrolysis model through corresponding characteristic parameters; The construction process of the combustion model is as follows: determining the reaction mechanism of the combustion model, the reaction mechanism of the combustion model is to describe the combustion reaction; selecting the combustion model; establishing the mass, energy and momentum balance equations of the combustion process to simulate the conversion of reactants and products; calibrating the combustion model through corresponding characteristic parameters; The smoke generation model construction process is as follows: determining the reaction mechanism of the smoke generation model, the reaction mechanism of the smoke generation model is to describe the generation and transformation process of each component in the smoke; selecting the smoke generation model; establishing the mass balance and chemical reaction equation of the smoke generation; calibrating the smoke generation model through the corresponding characteristic parameters; The construction process of the pollutant generation model is: determining the reaction mechanism of the pollutant generation model, the reaction mechanism of the pollutant generation model is to describe the chemical reaction path of pollutant generation; selecting the pollutant generation model; establishing the chemical reaction equation of pollutant generation; calibrating the pollutant generation model through corresponding characteristic parameters.

4. The multi-objective optimization intelligent control method for a combustion carbon dioxide production system according to claim 3 is characterized in that: The method for constructing the mechanism-constrained deep reinforcement learning agent includes: The state space is defined as the combustion parameter vector collected in real time, including environmental parameters and characteristic parameters of garbage; The action space is defined as a set of adjustable parameters of the combustion equipment, including the oxygen supply adjustment amount, the grate speed increment and the continuous control amount of the secondary air door opening; Construct an Actor-Critic network architecture with a physical feasibility check gate, where the Actor network is used to generate continuous control actions, and the Critic network is used to design multi-objective reward functions and evaluate multi-objective rewards; The design method of the multi-objective reward function includes: The efficiency improvement item, carbon emission penalty item and pollutant exceeding gradient penalty item are constructed by combining combustion efficiency, carbon emissions and pollutant concentration through dynamic weight coefficients. The multi-objective reward function is obtained based on the efficiency improvement item, carbon emission penalty item and pollutant exceeding gradient penalty item. Among them, the carbon emission penalty intensity increases nonlinearly as the temperature deviates from the optimal range.

5. The multi-objective optimization intelligent control method for a combustion carbon dioxide production system according to claim 4 is characterized in that: The implementation method of the physical feasibility check gate includes: adding a constraint check module to the output layer of the Critic network, the constraint check module nonlinearly fuses the predicted Q value and the physical feasibility coefficient through the neural network layer, and outputs the action value evaluation result corrected by the physical constraint.

6. The multi-objective optimization intelligent control method for a combustion carbon dioxide production system according to claim 4 is characterized in that: The method for adjusting the dynamic weight coefficient includes: The operating data within a preset time period is used as the input of the stage identification model to obtain the current stage, which includes the startup stage, the steady-state stage and the transient stage; in the startup stage, the dynamic weight coefficient corresponding to the combustion efficiency is increased to the preset weight threshold, and the dynamic weight coefficients corresponding to the carbon emission penalty item and the gradient penalty item for exceeding the pollutant standard are randomly allocated according to the constraint that the sum of all dynamic weight coefficients is 1; in the steady-state stage, the dynamic weight coefficients corresponding to the carbon emission penalty item and the gradient penalty item for exceeding the pollutant standard are dynamically allocated according to the pollutant emission quota; in the transient stage, an emergency adjustment factor is introduced to adjust the dynamic weight coefficient corresponding to the gradient penalty item for exceeding the pollutant standard.

7. The multi-objective optimization intelligent control method for a combustion carbon dioxide production system according to claim 4 is characterized in that: The calibration method of the carbon emission penalty item includes: A step-by-step heating experiment was carried out in different temperature ranges of the combustion furnace. The material expansion coefficient at each temperature point was collected. The temperature influence factor was calibrated through the material expansion coefficient at each temperature point. The mapping relationship between the temperature deviation and the carbon emission penalty gain coefficient was established. The carbon emission penalty item was obtained by calibrating the mapping relationship between the temperature deviation and the carbon emission penalty gain coefficient.

8. The multi-objective optimization intelligent control method for a combustion carbon dioxide production system according to claim 7 is characterized in that: The optimization objectives of the optimization control strategy include maximizing combustion efficiency, minimizing carbon dioxide emissions, and achieving pollutant control standards.

9. The multi-objective optimization intelligent control method for a combustion carbon dioxide production system according to claim 8, characterized in that: The method for performing hierarchical progressive optimization control comprises: Conduct pre-training under the constraints of the mechanism model in a virtual environment to generate an initial control strategy library; Data-driven strategy fine-tuning through dual experience replay pools; Dynamically relax action space constraints according to operational stability indicators; The dual experience replay pool includes a physical compliance pool and an actual optimization pool; the physical compliance pool is used to store experience samples that meet preset constraints; the actual optimization pool is used to store samples that meet the constraints during operation; wherein the sampling priority of the dual experience replay pool is dynamically determined by the weighted fusion result of the physical compliance score and the reward value.

10. The multi-objective optimization intelligent control method for a combustion carbon dioxide production system according to claim 9, characterized in that: The method for dynamically relaxing action space constraints includes: When the thermal efficiency fluctuation rate for N consecutive control cycles is not higher than the preset fluctuation rate threshold, the constraint parameters are updated; Methods for updating constraint parameters include: The action space constraint parameters are adaptively relaxed and adjusted according to the operation stability index. When the performance index reaches the preset index threshold, the allowable control amount change range is proportionally expanded.

Citation Information

Patent Citations

  • System and method for real-time prediction of calorific value of circulating fluidized bed domestic waste incineration boiler

    CN105864797B

  • Intelligent control system of waste incineration power plant and waste incineration system

    CN118089034A

  • Heating furnace combustion intelligent control method and device based on big data cloud platform

    CN115307452A

  • Combustion system optimization design method based on physically-driven parameterized proxy model

    CN116542164A

  • Multi-time-scale detection method for dioxin emission concentration in urban solid waste incineration process

    CN117610401A

Cited By

  • Incineration pollution cooperative control method based on multi-objective reinforcement learning

    CN120650717A

  • Composite atmosphere control system and method for gasification treatment of waste circuit board

    CN121254713A

  • Heating furnace combustion control method, device and equipment and storage medium

    CN121430345A

  • Heating furnace combustion control method, device, equipment and storage medium

    CN121430345B