Boiler flue gas temperature dynamic compensation method based on CFD-reinforcement learning

By employing a dynamic compensation method for boiler flue gas temperature based on CFD-reinforcement learning, and combining a multi-scale coupled 3D CFD model with an intelligent agent, the real-time performance and accuracy of boiler flue gas temperature control under complex operating conditions are solved, thereby improving the stability and economy of boiler operation.

CN121365626APending Publication Date: 2026-01-20SHANGHAI WANGTE ENERGY RESOURCE SICENCE & TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511567546.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

Existing boiler flue gas temperature control methods are difficult to achieve real-time performance, accuracy, and adaptability to complex operating conditions. Traditional control methods suffer from problems such as response lag, insufficient accuracy, and weak generalization ability.

Method used

A dynamic compensation method for boiler flue gas temperature based on CFD-reinforcement learning is adopted. By constructing a multi-scale coupled three-dimensional CFD model and an intelligent agent, and combining a macroscopic gas-solid two-phase flow-combustion reaction-radiative heat transfer coupled algorithm with a microscopic pulverized coal particle combustion dynamics algorithm, boiler operating parameters are collected in real time, flue gas temperature distribution data is output, and real-time adjustment is performed through an echo state network-predictive control collaborative algorithm.

Benefits of technology

It achieves precise dynamic compensation of boiler flue gas temperature, reduces temperature deviation, avoids equipment corrosion and heat exchange efficiency decline, improves the stability and economy of boiler operation, shortens the training cycle, and reduces data costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121365626A_ABST
    Figure CN121365626A_ABST
Patent Text Reader

Abstract

The invention relates to a boiler flue gas temperature dynamic compensation method based on CFD-reinforcement learning, and belongs to the technical field of boiler operation control. The method comprises the steps that boiler operation parameters and flue gas temperature data are collected in real time, a multi-scale coupling three-dimensional CFD model is constructed, a macroscopic gas-solid two-phase flow-combustion reaction-radiation heat transfer coupling algorithm and a microcosmic pulverized coal particle combustion dynamics algorithm are fused, and hearth and flue temperature distribution is simulated and output; building a CFD-reinforcement learning fusion agent based on a CFD model, taking similar working condition agent parameters as initial values, combining historical and virtual data staged training, and reinforcing the weight of a key temperature measurement point through a focusing regulation and control strategy; an echo state network-predictive control collaborative algorithm is adopted, real-time data are input to obtain an initial adjustment amount, and after security constraint optimization, a compensation execution mechanism is controlled to act. The flue gas temperature is accurately and dynamically compensated, the adaptability to complex working conditions is improved, and equipment safety and operation efficiency are both considered.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of boiler operation control, and particularly relates to a boiler flue gas temperature dynamic compensation method based on CFD-reinforcement learning. BACKGROUND

[0002] As a core equipment of energy conversion and industrial production, the stable control of the flue gas temperature of a boiler is directly related to the heat exchange efficiency, equipment safety and environmental protection emission indicators. The flue gas temperature at the furnace outlet and the tail flue is too high, which is easy to cause the over-temperature aging of the heating surface, shorten the service life of the equipment; and the temperature is too low, which will reduce the thermal efficiency of the boiler, increase the energy consumption, and may cause low-temperature corrosion, ash deposition and other problems, and even affect the normal operation of the subsequent denitration and desulfurization system. Therefore, the precise dynamic compensation of the flue gas temperature of the boiler is a key technical requirement for guaranteeing the safe, efficient and environmentally friendly operation of the boiler. The current flue gas temperature compensation of the boiler mainly relies on traditional control methods and conventional intelligent control technologies, but their adaptability and precision under complex working conditions still have significant limitations. Early, PID control or empirical manual adjustment was mostly used. The PID control is based on the linear system assumption design, while there are strong nonlinear disturbances such as coal supply fluctuation, coal quality change and load switching in the operation process of the boiler, and the flue gas temperature is affected by the coupling of multiple physical fields such as gas-solid two-phase flow, combustion reaction and radiation heat transfer. It is difficult for the PID parameters to match the working condition changes in real time, and overshoot and lag phenomena are easy to occur. The empirical manual adjustment relies on the judgment of the working condition of the operator, and not only has slow response speed, but also the compensation precision is restricted by the experience level of the personnel, and it is difficult to cope with the complex scene of multivariable coupling. In order to improve the defects of the traditional method, single neural network, fuzzy control and other intelligent technologies have been developed subsequently. Although the single neural network can fit the nonlinear relationship, it needs to rely on a large amount of historical operation data for training, and can only make decisions based on local temperature data, and cannot perceive the global distribution of the temperature field in the furnace and the flue, which is easy to cause compensation deviation due to “local data one-sidedness”; although the fuzzy control can handle uncertain information, the rule base construction relies on expert experience, and lacks generalization ability for extreme working conditions not covered (such as low load stable combustion and high-sulfur coal combustion), and the compensation effect is easy to be limited by the working condition range. Although the computational fluid dynamics (CFD) technology can simulate the coupling process of multiple physical fields in the boiler and output the global temperature distribution data, the numerical calculation process is complex and time-consuming, and it is difficult to meet the real-time requirements of temperature compensation when it is used alone; although the reinforcement learning can optimize the control strategy through trial and error learning, the traditional reinforcement learning needs to start training from zero, the training cycle is long due to the changeable working conditions of the boiler, the data demand is large, and the combination of the temperature field physical law is lacking, which is easy to cause the problem of “disconnection between the strategy and the actual physical process”, leading to unreasonable action of the compensation execution mechanism. The prior art is difficult to meet the comprehensive needs of boiler flue gas temperature compensation for "real-time, accuracy, and adaptability to complex working conditions" at the same time, and a new compensation method that can integrate the advantages of multi-physical field simulation and intelligent decision-making and take into account global perception of the temperature field and dynamic adjustment optimization is needed to solve the problems of insufficient accuracy, response lag, and weak generalization ability of traditional technology under complex working conditions. SUMMARY

[0003] To solve the above problems in the prior art, the application provides a boiler flue gas temperature dynamic compensation method based on CFD-reinforcement learning. The object of the application can be achieved by the following technical solutions. A boiler flue gas temperature dynamic compensation method based on CFD-reinforcement learning, characterized by comprising the following steps: S1: Real-time acquisition of boiler operating parameters and actual flue gas temperature data; based on the actual structure size of the boiler, a multi-scale coupled three-dimensional CFD model covering the furnace, horizontal flue, and tail flue is constructed, the collected operating parameters are input as real-time boundary conditions of the multi-scale coupled three-dimensional CFD model based on the macroscopic gas-solid two-phase flow-combustion reaction-radiation heat coupling algorithm and the microscopic coal particle combustion dynamics algorithm, and based on numerical calculation, the flue gas temperature distribution data in the furnace and flue are output; S2: Based on the multi-scale coupled three-dimensional CFD model, define the state space, action space, and reward rules of the CFD-reinforcement learning fusion agent, use the agent parameters trained under the boiler working condition as the initial parameters, combine the current boiler historical operation data with the virtual simulation data generated by the multi-scale coupled three-dimensional CFD model to construct a hybrid training data set, use a phased training strategy to train the agent, and at the same time, incorporate a focus control strategy in the state feature extraction link to strengthen the weight of temperature measurement point data until the reward value converges to a preset stable threshold; S3: The echo state network-predictive control collaborative algorithm inputs the collected boiler operating parameters into the multi-scale coupled three-dimensional CFD model, obtains real-time flue gas temperature distribution data, and calculates the deviation of the actual temperature from the set value; the time series data composed of real-time temperature data, temperature deviation, and last time adjustment amount are input, the boiler temperature dynamic change law is captured through a reservoir pool, and the initial compensation adjustment amount under the current working condition is output; the initial adjustment amount is optimized and corrected in real time under the constraints of boiler equipment safety constraints and temperature regulation response speed to generate the final compensation adjustment amount, and the corresponding compensation actuator is controlled to act according to the final adjustment amount.

[0004] Specifically, the macroscopic gas-solid two-phase flow-combustion reaction-radiation heat transfer coupling algorithm adopts the Euler-Lagrange method to describe the gas-solid two-phase flow, the gas phase is modeled in the Euler coordinate system, and the flow velocity and turbulence intensity distribution are calculated through the turbulence model; the solid phase is tracked in the Lagrange coordinate system to calculate the particle trajectory, concentration distribution and momentum exchange with the gas phase; the combustion reaction adopts the vortex dissipation concept model to quantify the mixing reaction rate of the gas phase fuel and oxygen, and combines the heterogeneous combustion reaction rate of coke to calculate the total heat release in the combustion process; the radiation heat transfer adopts the discrete coordinate method to calculate the regional radiation heat transfer based on the radiation characteristics of the radiation gas and coal powder particles in the flue gas.

[0005] Specifically, the microscopic coal particle combustion kinetics algorithm is for the whole life cycle of single particle coal combustion. In the coal pyrolysis stage, a double competition reaction model is adopted to establish the pyrolysis kinetics equation of volatile components and non-volatile components respectively, and the volatile release rate and composition ratio at different temperatures are calculated; in the coke combustion stage, an Arrhenius type reaction rate equation is adopted, and a particle pore structure factor is introduced to modify the reaction area, so as to quantify the reaction rate of coke and oxygen, carbon dioxide and generate the coke combustion heat release law; in the sulfur release stage, a segmented kinetics model is adopted to distinguish the release characteristics of organic sulfur and inorganic sulfur, and the influence of sulfur release amount on the specific heat capacity of the particle is calculated; the pyrolysis rate of the particle, the coke reaction heat release rate and the microscopic parameters of the change of the particle temperature are mapped to the macroscopic calculation grid of the multi-scale coupled three-dimensional CFD model through the volume average method.

[0006] Specifically, the state space design process of the intelligent agent is as follows: taking the temperature field key information output by the multi-scale coupled three-dimensional CFD model as the core, selecting the predicted temperature at the furnace outlet, the predicted temperature at the economizer inlet, the predicted temperature at the air preheater inlet, and the deviation value between the actual flue gas temperature and the preset set value, and supplementing the compensation execution mechanism adjustment amount at the last moment, to jointly constitute the state space.

[0007] Specifically, the action space of the intelligent agent is based on the actual operable means of boiler flue gas temperature regulation, and the secondary air damper opening adjustment amount, the economizer bypass damper opening adjustment amount and the burner swing angle adjustment amount are selected as the core dimensions of the action space; the secondary air damper opening adjustment changes the air distribution ratio in the combustion area to adjust the flame center temperature, the economizer bypass damper opening adjustment changes the flue gas flow through the economizer to adjust the tail flue temperature, and the burner swing angle adjustment directly adjusts the furnace outlet temperature by changing the flame center height; the value range of the adjustment amount in the action space matches the safety operation limit of the corresponding execution mechanism, and the step size setting of adjacent adjustment amounts is based on the adjustment accuracy and the equipment response speed.

[0008] Specifically, the design logic of the reward rule is as follows: a multi-objective weighted reward function design idea is adopted, a first objective is a temperature compensation effect, the reward value monotonically increases with the decrease of the deviation of the actual temperature from the set value; a second objective is a device running stability, the reward value monotonically increases with the decrease of the difference between the current adjustment amount and the adjustment amount at the last moment; a third objective is a boiler running economy, the reward value monotonically increases with the increase of the real-time combustion efficiency of the boiler, and the combustion efficiency is calculated based on the oxygen content and the exhaust gas temperature; the priority of the three objectives is balanced by a preset weight coefficient, and the weight coefficient is dynamically adjusted according to the running demand of the boiler.

[0009] Specifically, the specific implementation process of taking the agent parameters trained under the boiler working condition as the initial parameters is as follows: first, a reference boiler similar to the current boiler is selected, the similarity is judged according to the rated evaporation capacity, the volatile characteristic of the coal used, the consistency of the burner type, and the overlapping of the load regulation range, and the operation law of the reference boiler is comparable to that of the current boiler; then the core parameters of the agent trained under the reference boiler are extracted, including the convolution layer weight and the full connection layer bias of the deep Q network, and the core parameters are taken as the initial parameters of the agent.

[0010] Specifically, the detailed implementation process of the staged training strategy is as follows: in the first stage, the agent is trained offline to master the basic law of boiler temperature regulation, and the training data set is constructed by integrating the historical operation data of the current boiler accumulated for a long time and the virtual working condition data generated by the multi-scale coupled three-dimensional CFD model, and the agent is iteratively trained by using a fixed learning rate; in the second stage, the agent is fine-tuned online to adapt to the actual operation characteristics of the current boiler, and the network parameters of the agent are updated at a preset time by real-time acquisition of the operation data and the adjustment effect data of the current boiler, and a decay learning rate is used in the updating process until the reward mean value converges to a preset stable threshold.

[0011] Specifically, the focus regulation strategy is integrated into the state feature extraction link by calculating the attention weight of the temperature measurement point data in the state space, the weight calculation process introduces the temperature field spatial gradient information output by the multi-scale coupled three-dimensional CFD model, establishes the association mapping between the attention weight and the boiler working condition, and the deviation contribution degree of the temperature measurement point data is calculated based on a sliding window.

[0012] Specifically, the predictive control is specifically a model predictive control, by constructing a prediction model of boiler temperature adjustment, based on historical operation data and simulation data of the multi-scale coupled three-dimensional CFD model, the dynamic relationship between operation parameters, adjustment amount and temperature change is quantified; the prediction time domain and the control time domain are set, the prediction time domain is used to predict the temperature change trend, and the control time domain is used to determine the current adjustment amount sequence to be optimized; the constraint conditions include boiler equipment safety constraints, lower limit of secondary air pressure, and upper limit of burner swing angle adjustment rate; by solving the optimization problem with constraints, the initial adjustment amount output by the echo state network is corrected to generate the final adjustment amount.

[0013] Specifically, the construction process of the echo state network is as follows: the core component is a reservoir, the reservoir is composed of neurons connected to each other, the connection weights between neurons are randomly generated and remain fixed during the training process, and the output layer weights are adjusted; by setting the spectral radius of the reservoir and the sparsity of the input weight matrix, the reservoir can capture the dynamic change rule of the time series data in the state space.

[0014] Specifically, the compensation actuator action is based on the response characteristic difference of the compensation actuator to formulate a cooperative rule: the secondary air door is used as rapid adjustment, the coal economizer bypass damper is used as transition adjustment, and the burner swing angle is used as deep adjustment; a real-time temperature feedback correction mechanism is introduced, after the actuator action, the temperature change data of the corresponding area is collected in real time, if the temperature deviation reduction rate is lower than the preset threshold, the adjustment amplitude is immediately adjusted.

[0015] The beneficial effects of the present application are: The method can accurately simulate the global flue gas temperature distribution in the furnace and flue by the multi-scale coupled three-dimensional CFD model, which combines the macro gas-solid two-phase flow-combustion reaction-radiation heat coupling algorithm and the micro coal particle combustion dynamics algorithm, avoiding the problem of "one-sided temperature field cognition" caused by traditional dependence on local measurement point data; combined with the decision-making ability of the CFD-strengthened learning integrated agent, the optimal adjustment strategy can be output according to the root cause of temperature deviation, so that the deviation amplitude of flue gas temperature from the set value is significantly reduced, effectively avoiding the corrosion and wear of heat exchange equipment (such as coal economizer and air preheater) caused by temperature overshoot, or the problem of heat exchange efficiency reduction caused by too low temperature, prolonging the service life of the core equipment of the boiler.

[0016] To overcome the defects of traditional compensation methods (such as PID control and empirical adjustment) in complex scenarios such as boiler variable load operation, coal type volatile matter fluctuation, and seasonal operating condition switching, the method improves adaptability through two technical means: first, the CFD-enhanced learning agent uses a "similar condition agent parameter initialization + phased training" mode to reduce dependence on current boiler historical data and quickly adapt to new conditions. Second, the state feature extraction section focuses on the control strategy, dynamically strengthens the weight of key temperature measurement points (such as the furnace outlet and the economizer inlet), and accurately captures the temperature fluctuation pattern under changing conditions. The echo state network-predictive control collaborative algorithm can track temperature changes in real time, avoid regulation lag caused by sudden changes in operating conditions, and ensure stable compensation effects throughout the operating cycle.

[0017] The method explicitly includes both "equipment operation stability" and "boiler operation economy" in the agent reward rules, guiding the agent to eliminate temperature deviations while avoiding sudden changes in the adjustment amount (such as frequent opening and closing of the secondary air damper and sharp adjustment of the burner angle) during the training process, reducing mechanical wear of the actuator, and lowering equipment maintenance costs. Precise temperature compensation ensures that the boiler combustion zone temperature is within the optimal range, avoiding insufficient combustion (increasing exhaust gas heat loss) or excessive combustion (increasing fuel consumption) caused by temperature deviation. Combined with the efficient heat exchange of the economizer and air preheater, the method can significantly reduce the energy consumption per unit of boiler evaporation and improve overall operational efficiency.

[0018] Traditional reinforcement learning agents require massive historical operating data for zero training, which has the engineering pain points of "long data collection period and slow training convergence." The method uses transfer learning to use the parameters of agents trained under similar boiler operating conditions as initial parameters, significantly reducing the amount of training data required for the current boiler. It also uses virtual simulation data generated by the CFD model (which can cover extreme and rare operating conditions) to build a hybrid training data set, avoiding the problem of insufficient training due to the difficulty of obtaining extreme operating condition data in actual operation. This shortens the agent training period, quickly meets the needs of engineering applications, and reduces the time and data costs of technology implementation.

[0019] Compared to traditional compensation methods based on fixed control logic, the echo state network (ESN) in this method quickly captures the temperature time series variation pattern (such as the lag characteristic of temperature deviation with adjustment amount) through the reservoir, which can output the initial adjustment amount in real time. Combined with the model predictive control (MPC) for rapid optimization and correction of the initial adjustment amount, it can start precise adjustment as soon as the temperature deviation appears, avoiding the accumulation and expansion of the deviation. The collaborative control rules of the actuator (fast adjustment unit responds first, and deep adjustment unit starts as needed) further accelerate the adjustment response speed, reduce long-term temperature deviation caused by adjustment lag, and ensure stable boiler operating parameters. BRIEF DESCRIPTION OF DRAWINGS

[0020] For the convenience of those skilled in the art to understand, the present application is further described below in conjunction with the drawings.

[0021] Fig. 1 The overall architecture and workflow diagram of a CFD-reinforcement learning-based boiler flue gas temperature dynamic compensation method of the present application; Fig. 2 The CFD-reinforcement learning fusion agent architecture diagram of a CFD-reinforcement learning-based boiler flue gas temperature dynamic compensation method of the present application. DETAILED DESCRIPTION

[0022] In order to further illustrate the technical means and effects adopted by the present application to achieve the predetermined object, the specific embodiments, structures, features and effects according to the present application are described in detail below in conjunction with the drawings and preferred embodiments.

[0023] Please refer to Figs. 1-2 A CFD-reinforcement learning-based boiler flue gas temperature dynamic compensation method, characterized in that it comprises the following steps: S1: Real-time acquisition of boiler operating parameters and actual flue gas temperature data; based on the actual structural size of the boiler, a multi-scale coupled three-dimensional CFD model covering the furnace, horizontal flue and tail flue is constructed, based on the macroscopic gas-solid two-phase flow-combustion reaction-radiation heat coupling algorithm and the microscopic coal particle combustion dynamics algorithm, the acquired operating parameters are input as the real-time boundary conditions of the multi-scale coupled three-dimensional CFD model, based on numerical calculation, the flue gas temperature distribution data in the furnace and flue are output; S2: Defining the state space, action space and reward rules of the CFD-reinforcement learning fusion agent based on the multi-scale coupled three-dimensional CFD model, taking the trained agent parameters under the boiler operating conditions as the initial parameters, combining the current boiler historical operation data with the virtual simulation data generated by the multi-scale coupled three-dimensional CFD model to construct a hybrid training data set, using a phased training strategy to train the agent, and at the same time, incorporating a focus control strategy in the state feature extraction link to strengthen the weight of the temperature measurement point data until the reward value converges to a preset stable threshold; S3: The echo state network-predictive control collaborative algorithm inputs the collected boiler operating parameters into the multi-scale coupled three-dimensional CFD model in real time, obtains real-time flue gas temperature distribution data and calculates the deviation of the actual temperature from the set value; the time series data composed of real-time temperature data, temperature deviation and last time adjustment amount are input, the reservoir is used to capture the dynamic change law of the boiler temperature, and the initial compensation adjustment amount under the current working condition is output; the initial adjustment amount is optimized and corrected in real time under the constraint conditions of the safety constraints of the boiler equipment and the temperature regulation response speed, and the final compensation adjustment amount is generated, and the corresponding compensation actuator is controlled to act according to the final adjustment amount.

[0024] Specifically, the macroscopic gas-solid two-phase flow-combustion reaction-radiation heat transfer coupling algorithm adopts the Euler-Lagrange method to describe the gas-solid two-phase flow, the gas phase is modeled in the Euler coordinate system, and the flow velocity and turbulence intensity distribution are calculated through the turbulence model; the solid phase is tracked in the Lagrangian coordinate system, and the particle motion trajectory, concentration distribution and momentum exchange with the gas phase are calculated; the combustion reaction adopts the vortex dissipation concept model, quantifies the mixing reaction rate of the gas phase fuel and oxygen, and calculates the total heat release in the combustion process combined with the heterogeneous combustion reaction rate of coke; the radiation heat transfer adopts the discrete coordinate method, and the radiation characteristics of the radiation gas and coal powder particles in the flue gas are considered to calculate the regional radiation heat transfer.

[0025] The macroscopic gas-solid two-phase flow-combustion reaction-radiation heat transfer coupling algorithm adopts the Euler-Lagrange method to describe the gas-solid two-phase flow, the gas phase (flue gas) is modeled in the Euler coordinate system, and the flow velocity and turbulence intensity distribution are calculated through the k-ε turbulence model; the solid phase (coal powder particles) is tracked in the Lagrangian coordinate system, and the particle motion trajectory, concentration distribution and momentum exchange with the gas phase are calculated; the combustion reaction adopts the vortex dissipation concept (EDC) model, quantifies the mixing reaction rate of the gas phase fuel and oxygen, and calculates the total heat release in the combustion process combined with the heterogeneous combustion reaction rate of coke; the radiation heat transfer adopts the discrete coordinate (DO) method, considers the radiation characteristics of the radiation gas such as CO2 and H2O and coal powder particles in the flue gas, and calculates the radiation heat transfer of each region; the three parameters are coupled through the data interaction module—the gas phase flow rate affects the combustion reaction mixing efficiency, the combustion heat release provides heat source for the temperature field, and the radiation heat transfer corrects the local temperature distribution, which jointly supports the temperature field simulation accuracy of the CFD model.

[0026] Specifically, the micro-pulverized coal particle combustion dynamics algorithm is designed for the whole life cycle of single-particle coal combustion. In the pyrolysis stage, a double-competition reaction model is used to establish the pyrolysis kinetics equation of volatile components and non-volatile components, respectively, to calculate the volatile release rate and composition ratio at different temperatures. In the coke combustion stage, an Arrhenius-type reaction rate equation is used to introduce a particle pore structure factor to correct the reaction area, quantify the reaction rate of coke and oxygen and carbon dioxide, and generate the coke combustion heat release law. In the sulfur release stage, a segmented kinetics model is used to distinguish the release characteristics of organic sulfur and inorganic sulfur, and to calculate the influence of sulfur release on the specific heat capacity of the particle. The particle pyrolysis rate, coke reaction heat release rate, and particle temperature change micro-parameters are mapped to the macroscopic calculation grid of the multi-scale coupled three-dimensional CFD model through the volume average method.

[0027] Specifically, the state space design process of the intelligent agent is as follows: taking the temperature field key information output by the multi-scale coupled three-dimensional CFD model as the core, selecting the predicted temperature at the furnace outlet, the predicted temperature at the economizer inlet, and the predicted temperature at the air preheater inlet, and incorporating the deviation value between the actual flue gas temperature and the preset set value, and supplementing the last time compensation actuator adjustment amount, to form the state space.

[0028] Specifically, the action space of the intelligent agent is based on the actual operable means of boiler flue gas temperature regulation, and the secondary air damper opening adjustment, the economizer bypass damper opening adjustment, and the burner swing angle adjustment are selected as the core dimensions of the action space. The secondary air damper opening adjustment adjusts the flame center temperature by changing the air distribution ratio in the combustion zone, the economizer bypass damper opening adjustment adjusts the tail flue temperature by changing the flue gas flow through the economizer, and the burner swing angle adjustment directly adjusts the furnace outlet temperature by changing the flame center height. The value range of the adjustment amount in the action space matches the safety operation limit of the corresponding actuator, and the step size of adjacent adjustment amounts is set based on the adjustment accuracy and the equipment response speed.

[0029] Specifically, the design logic of the reward rule is as follows: a multi-objective weighted reward function design idea is adopted. The first target is the temperature compensation effect, and the reward value monotonically increases with the decrease of the deviation between the actual temperature and the set value. The second target is the equipment operation stability, and the reward value monotonically increases with the decrease of the difference between the current adjustment amount and the last time adjustment amount. The third target is the boiler operation economy, and the reward value monotonically increases with the increase of the real-time combustion efficiency of the boiler, which is calculated based on the oxygen content and the exhaust gas temperature. The priority of the three targets is balanced by a preset weight coefficient, and the weight coefficient is dynamically adjusted according to the boiler operation demand.

[0030] Specifically, the specific implementation process of taking the agent parameters trained under the boiler working condition as the initial parameters is as follows: first, a reference boiler similar to the current boiler is selected, and the similarity is determined according to the rated evaporation capacity, the volatile characteristics of the coal used, the consistency of the burner type, and the overlapping of the load regulation range, and the operation law of the reference boiler is comparable to that of the current boiler; then, the core parameters of the agent trained under the reference boiler are extracted, including the convolution layer weight and the full connection layer bias of the deep Q network, and the core parameters are taken as the initial parameters of the agent.

[0031] In the agent parameter extraction stage, for the selected reference boiler, the parameter sensitivity analysis method is used to identify the core parameters of the agent that significantly affect the performance of the flue gas temperature compensation control. The convolution layer part of the deep Q network (DQN) focuses on extracting the weight matrix containing 3x3 and 5x5 convolution kernels, where the shallow convolution layer weight is used to capture the local detail features of the temperature field, and the deep convolution layer weight is used to extract the global structural features; the full connection layer extracts the bias vector, which, after being processed by the ReLU activation function, maps the feature tensor output by the convolution layer to the decision signal of the corresponding action space (such as water injection amount adjustment, burner damper opening adjustment, etc.).

[0032] In the parameter migration application stage, the adversarial training mechanism in transfer learning is introduced. When migrating the parameters of the reference boiler to the current agent, a discriminator network is constructed to distinguish the source of the parameters (reference boiler or current boiler), and the gradient inversion training is performed on the agent generator, so that the migrated parameters can adapt to the new working condition environment. At the same time, an incremental parameter fusion strategy is adopted. Initially, 80% of the reference parameters and 20% of the randomly initialized parameters are mixed, and the proportion of the current boiler parameters is gradually increased as the training progresses. This reduces the demand and cycle of training data, effectively alleviates the negative transfer problem caused by the difference in data distribution, and avoids the training convergence difficulty caused by the lack of historical data of the current boiler.

[0033] Specifically, the detailed implementation process of the phased training strategy is as follows: in the first phase, the agent is trained offline to master the basic rules of boiler temperature regulation. Specifically, the training data set is constructed by integrating the historical operation data of the current boiler accumulated for a long time and the virtual working condition data generated by the multi-scale coupled three-dimensional CFD model, and the agent is iteratively trained with a fixed learning rate; in the second phase, the agent is fine-tuned online to adapt to the actual operation characteristics of the current boiler. Specifically, the operation data and regulation effect data of the current boiler are collected in real time, and the network parameters of the agent are updated at a predetermined time. During the updating process, a decaying learning rate is used, and the reward mean is completely converged to the preset stable threshold.

[0034] Specifically, the focus regulation strategy is integrated into the state feature extraction link to calculate the attention weight of the temperature measurement point data in the state space. The weight calculation process introduces the temperature field spatial gradient information output by the multi-scale coupled three-dimensional CFD model, establishes the association mapping between the attention weight and the boiler operating condition, and statistically calculates the deviation contribution of the temperature measurement point data based on a sliding window.

[0035] Specifically, the predictive control is specifically a model predictive control. A prediction model of boiler temperature regulation is constructed to quantify the dynamic relationship between operating parameters, regulation amount and temperature change based on historical operation data and simulation data of the multi-scale coupled three-dimensional CFD model. A prediction time domain and a control time domain are set. The prediction time domain is used to predict the temperature change trend, and the control time domain is used to determine the current regulation amount sequence to be optimized. The constraint conditions include boiler equipment safety constraints, lower limit of secondary air pressure, and upper limit of burner swing angle adjustment rate. The initial regulation amount output by the echo state network is corrected by solving the constrained optimization problem to generate the final regulation amount.

[0036] Specifically, the construction process of the echo state network is as follows: the core component is a reservoir, which is composed of interconnected neurons. The connection weights between neurons are randomly generated and remain fixed during the training process, and the output layer weights are adjusted. By setting the spectral radius of the reservoir and the sparsity of the input weight matrix, the reservoir can capture the dynamic change law of the time series data in the state space.

[0037] Specifically, the compensation actuator action is based on the response characteristic difference of the compensation actuators to develop a coordination rule: the secondary air door is used for rapid regulation, the coal economizer bypass damper is used for transition regulation, and the burner swing angle is used for deep regulation. A real-time temperature feedback correction mechanism is introduced. After the actuator action, the temperature change data of the corresponding area are collected in real time. If the temperature deviation reduction rate is lower than the preset threshold, the adjustment amplitude is immediately adjusted.

[0038] First, different regulation levels are divided according to the absolute value of the deviation between the actual temperature and the set value. When the absolute value of the deviation is small, only the secondary air door opening is adjusted to achieve temperature compensation. The reason is that the secondary air door has fast response speed and small system disturbance, which is suitable for fine tuning. When the absolute value of the deviation is at a moderate level, the secondary air door opening and the coal economizer bypass damper opening are adjusted simultaneously to accelerate the temperature deviation elimination speed and avoid excessive adjustment of a single actuator. When the absolute value of the deviation is large, the secondary air door opening, the coal economizer bypass damper opening and the burner swing angle are adjusted simultaneously to fully utilize all available means to quickly pull the temperature back to the set value range and avoid equipment damage due to excessive deviation. The switching of each regulation level is determined dynamically based on the real-time deviation to ensure the flexibility and adaptability of the regulation strategy.

[0039] In this embodiment, the CFD-reinforced learning-based boiler flue gas temperature dynamic compensation method embodiment This embodiment takes a certain power subcritical pulverized coal boiler in a power plant as the application object. The boiler is designed with a rated evaporation capacity, and burns bituminous coal (the received basis volatile matter meets the design coal type requirement, and the received basis low heat value meets the boiler heat load demand). The main heating surfaces include the furnace (the height and cross-sectional size are adapted to the boiler capacity), the horizontal flue (the length matches the overall structure of the boiler), and the tail flue (the economizer and air preheater are arranged according to the design). The compensation actuators include the secondary air dampers corresponding to the number of layers, the single economizer bypass damper (the flow area is adapted to the flue gas flow rate demand), and the burner (the swing angle adjustment range meets the flame center adjustment demand). The specific implementation process is as follows: CFD model construction and temperature field simulation Data acquisition Through the boiler DCS system, the operating parameters are collected in real time at a preset sampling frequency: the coal supply amount (the measurement range and accuracy meet the boiler load regulation demand), the primary air volume (the measurement range and accuracy match the combustion air distribution requirement), the secondary air volume (the measurement range and accuracy adapt to the furnace combustion condition), the furnace pressure (the measurement range and accuracy meet the furnace safe operation monitoring), and the feedwater temperature (the measurement range and accuracy meet the steam-water system design). The actual flue gas temperature is collected by using high-temperature-resistant K-type thermocouples (the range and accuracy are adapted to the flue gas temperature measurement scene). The measurement points are arranged according to the temperature monitoring key areas: the furnace outlet (the number of measurement points meets the cross-sectional temperature uniformity representation), the economizer inlet (the number of measurement points adapts to the flue temperature monitoring demand), and the air preheater inlet (the number of measurement points meets the tail flue temperature collection requirement).

[0040] Multi-scale coupled three-dimensional CFD model construction Based on the actual structure size of the boiler, a multi-scale coupled three-dimensional CFD model covering the furnace (including the designed number of burners), the horizontal flue (including the screen superheater), and the tail flue (including the economizer and the air preheater) is constructed. The structured grid division is adopted. The total number of grids matches the model accuracy demand. The grids in the key areas such as the burner area and the economizer tube bundle area are densified (the size meets the local calculation accuracy requirement). The grid size in other areas is set according to the calculation efficiency and accuracy balance principle.

[0041] Macroscopic gas-solid two-phase flow-combustion reaction-radiation heat transfer coupling algorithm Gas-solid two-phase flow calculation: The Euler-Lagrange method is adopted. The gas phase is modeled in the Euler coordinate system. A turbulence model adapted to the turbulence characteristics in the furnace (with the corresponding turbulence Prandtl number and wall function) is selected to calculate the flow velocity and turbulence intensity distribution. The solid phase is tracked in the Lagrangian coordinate system to calculate the coal particle motion trajectory, concentration distribution, and momentum exchange with the gas phase.

[0042] Combustion reaction calculation: the EDC model is used, the reaction mechanism is selected to adapt to the combustion characteristics of the coal, the heterogeneous combustion reaction rate coefficient of coke is set according to the coke reaction characteristics of the coal, and the mixing reaction rate of the gaseous fuel and oxygen and the total heat release of coke combustion are quantified.

[0043] Radiation heat transfer calculation: the DO method is used, the radiation characteristics of the radiation gas in the flue gas (determined according to the actual combustion product composition) and the emissivity of the coal particles (adapted to the physical characteristics of the coal particles) are considered, and the radiation heat transfer in each region is calculated.

[0044] Microscopic coal particle combustion kinetics algorithm Coal pyrolysis stage: a double-competition reaction model is used, the activation energy and frequency factor of the volatile component and the non-volatile component are set according to the coal pyrolysis experimental data, and the pyrolysis kinetics equation is established to calculate the volatile release rate and composition ratio at different temperatures.

[0045] Coke combustion stage: the Arrhenius-type reaction rate equation is used, the activation energy and frequency factor are adapted to the coke combustion characteristics, the particle pore structure factor (determined according to the pore characteristics of the coal particles) is introduced to correct the reaction area, and the reaction rate and heat release law of coke and O2, CO2 are quantified.

[0046] Sulfur release stage in coal: a segmented kinetics model is used, the release temperature interval of organic sulfur and inorganic sulfur is determined according to the sulfur existence form of the coal, and the influence of sulfur release amount on the specific heat capacity of the particle is calculated.

[0047] Micro-macro mapping: through the volume average method, the particle pyrolysis rate, coke reaction heat release rate, and particle temperature change are mapped to the macroscopic calculation grid of the CFD model.

[0048] Numerical calculation and temperature output The real-time collected operating parameters are input as boundary conditions into the CFD model, the algorithm adapted to the flow and combustion coupling calculation is used for solving, the iteration step is set to the residual convergence to the preset precision threshold, and the predicted temperatures of the furnace outlet, the economizer inlet, and the air preheater inlet are output, and the deviation of the actual temperature of the corresponding measuring point is controlled within the preset allowable range.

[0049] CFD-reinforced learning fusion agent training Agent state space design Taking the key information of the temperature field output by the CFD model as the core, selecting the predicted temperatures of the furnace outlet, the economizer inlet, and the air preheater inlet, including the deviation of the actual temperature of the above measuring points from the preset set value, and supplementing the adjustment amount of each compensation actuator at the previous time, together constitute a state space adapted to the temperature regulation demand, and all state quantities are standardized to a preset interval.

[0050] Agent action space design Based on the actual operable means of boiler flue gas temperature regulation, the secondary air door opening degree, the economizer bypass damper opening degree, and the burner swing angle are selected as the core dimensions of the action space: Secondary air door opening degree: The value range matches the safe operation threshold of the air door, and the step size is adapted to the regulation accuracy and device response speed (the flame center temperature is adjusted by changing the air distribution ratio in the combustion area); Economizer bypass damper opening degree: The value range meets the safety operation requirements of the damper, and the step size balances the regulation accuracy and response efficiency (the tail flue temperature is adjusted by changing the flue gas flow through the economizer); Burner swing angle: The value range is adapted to the mechanical limit of the swing angle, and the step size meets the fine adjustment requirements of the flame center height (the furnace outlet temperature is adjusted by changing the flame center height).

[0051] Reward rule design A multi-objective weighted reward function design is adopted, and the formula is as follows: Where: (Temperature compensation effect): The reward value monotonically increases with the decrease of the deviation between the actual temperature and the set value, and the value range is set according to the deviation allowed interval; (Device operation stability): The reward value monotonically increases with the decrease of the difference between the current regulation quantity and the last time regulation quantity, and the value range is set according to the allowed range of regulation quantity fluctuation; (Boiler operation economy): The reward value monotonically increases with the increase of the real-time combustion efficiency of the boiler, the combustion efficiency is calculated based on the oxygen content and the exhaust gas temperature, and the value range is set according to the efficiency design interval; is the weight coefficient, which is dynamically adjusted according to the priority of boiler operation (such as variable load, steady-state operation).

[0052] Migration learning initial parameter selection First, select a reference boiler similar to the current boiler operating condition: the similarity judgment basis includes rated parameters (evaporation capacity, power), coal type characteristics (volatile range), burner type, and load regulation range overlap (to ensure the comparability of operating rules); then extract the core parameters of the agent trained by the reference boiler (including the convolution layer weights and full connection layer biases of the neural network) as the initial parameters of the current agent.

[0053] Construction of mixed training data set Integrate the historical operation data accumulated by the current boiler for a long time (covering the conventional load and coal type conditions) and the virtual simulation data generated by the CFD model (including low load, high load and extreme coal type conditions) to build a mixed training data set covering all conditions and ensure that the data quantity meets the training needs of the intelligent agent.

[0054] Stage training strategy Offline pre-training stage: Set the learning rate and batch size for the adaptive intelligent agent basic learning, and set the iteration number to the basic law of the intelligent agent mastering the boiler temperature regulation. The ε-greedy strategy (ε is linearly decayed according to the training progress) is adopted, and the reward mean after training is stable in the preset basic threshold range.

[0055] Online fine-tuning stage: Set the real-time data update interval (update the network parameters once every preset number of real-time data), and decay the learning rate according to the training convergence requirements. The iteration number is set to adapt to the actual operation characteristics of the current boiler, and the final reward mean converges to the preset stable threshold. At this time, the intelligent agent's response time to the key temperature deviation meets the design requirements.

[0056] Focus on control strategy integration In the state feature extraction link, the attention weight of the temperature measurement point data in the state space is calculated by the Softmax function: The weight calculation process introduces the temperature field spatial gradient information (the weight of the area with larger gradient is higher) output by the CFD model, and establishes the association mapping between the attention weight and the boiler operating condition; Based on the sliding window, the deviation contribution degree of the temperature measurement point data is calculated in real time. The initial weight is set according to the importance of the temperature measurement point. When the boiler operating condition (such as load fluctuation) reaches the preset threshold, the weight of the key temperature measurement point (such as the furnace outlet temperature) is automatically increased. If the deviation contribution of a certain temperature measurement point exceeds the preset value for a continuous number of steps, the weight of the key measurement point data is temporarily increased to strengthen the influence of the key measurement point data on the intelligent agent's decision-making.

[0057] Echo state network-predictive control collaborative compensation Echo state network (ESN) construction and initial adjustment output ESN model construction: The core component is the reservoir (composed of interconnected neurons), the connection weights between neurons are randomly generated and remain fixed during the training process, and only the output layer weights are adjusted; By setting the spectral radius of the reservoir and the sparsity of the input weight matrix, it is ensured that the reservoir can capture the dynamic change law of the time series data in the state space.

[0058] Time series data input: Collect time series data of CFD temperature data (furnace outlet, economizer inlet, air preheater inlet temperature), temperature deviation (difference between actual temperature and set value), last time adjustment amount to adapt to the demand of dynamic capture of time window length.

[0059] Initial adjustment amount output: ESN model outputs the initial compensation adjustment amount (including secondary air damper, economizer bypass damper, and burner swing angle adjustment amount) under the current working condition based on the dynamic law of time series data.

[0060] Model predictive control (MPC) optimization correction Prediction model construction: By integrating historical operation data and CFD simulation data, the dynamic relationship between operation parameters, adjustment amount and temperature change is quantified, and a prediction model for boiler temperature regulation is constructed.

[0061] Time domain and constraint setting: Set the prediction time domain (used to predict the temperature change trend, length adapted to the temperature response characteristics) and the control time domain (used to determine the current adjustment amount sequence to be optimized, length balanced between adjustment precision and calculation efficiency); constraint conditions include boiler equipment safety constraints (actuator opening range), operation safety constraints (such as lower limit of secondary air pressure, upper limit of burner swing angle adjustment rate).

[0062] Optimization solution and adjustment amount correction: By solving the optimization problem with constraints (the objective function is to minimize the temperature deviation and adjustment amount fluctuation), the initial adjustment amount output by ESN is optimized and corrected in real time to generate the final compensation adjustment amount.

[0063] Actuator action and feedback correction Actuator cooperative action: Based on the response characteristic differences of each compensation actuator, cooperative rules are formulated: secondary air damper as a fast adjustment component (fast response speed, used for short-term temperature fluctuation compensation), economizer bypass damper as a transition adjustment component (moderate response speed, used for medium-term temperature trend adjustment), and burner swing angle as a deep adjustment component (slow response speed, used for long-term temperature baseline correction).

[0064] Real-time temperature feedback correction: After the actuator action, the corresponding region temperature change data is collected in real time, and the temperature deviation reduction rate is calculated; if the deviation reduction rate is lower than the preset threshold, the adjustment amplitude is immediately adjusted (such as increasing the damper opening adjustment amount, adjusting the swing angle change angle), to ensure that the temperature deviation converges quickly to the allowed range.

[0065] Variable load condition compensation: when the boiler appears variable load condition (load rises or falls according to preset rate), the CFD model predicts the temperature change trend of key measuring points in real time, the ESN outputs initial adjustment amount based on time series data under new condition, and drives the actuator to act after being corrected by the MPC; the temperature change is continuously monitored, the adjustment strategy is dynamically adjusted according to the convergence of deviation, and the temperature compensation under variable load condition is realized.

[0066] The above is only the preferred embodiment of the present application, and does not limit the present application in any form. Although the present application has been disclosed as above with the preferred embodiment, it is not intended to limit the present application. Any person skilled in the art can make slight changes or modifications to the above disclosed technical content to obtain equivalent embodiments with equivalent changes, without departing from the technical solution range of the present application. Any simple modification, equivalent change and modification of the above embodiments according to the technical essence of the present application still belong to the technical solution range of the present application.

Claims

1. A dynamic compensation method for boiler flue gas temperature based on CFD-reinforcement learning, characterized in that, Includes the following steps: S1: Real-time acquisition of boiler operating parameters and actual flue gas temperature data; Based on the actual structural dimensions of the boiler, a multi-scale coupled three-dimensional CFD model covering the furnace, horizontal flue, and tail flue is constructed. Based on the macroscopic gas-solid two-phase flow-combustion reaction-radiative heat transfer coupled algorithm and the microscopic pulverized coal particle combustion dynamics algorithm, the collected operating parameters are used as the real-time boundary condition input of the multi-scale coupled three-dimensional CFD model. Based on numerical calculation, the flue gas temperature distribution data in the furnace and flue are output. S2: Based on the multi-scale coupled three-dimensional CFD model, define the state space, action space, and reward rules of the CFD-reinforcement learning fusion agent. Use the parameters of the agent trained under boiler operating conditions as initial parameters. Combine the current historical boiler operating data with the virtual simulation data generated by the multi-scale coupled three-dimensional CFD model to construct a hybrid training dataset. Use a phased training strategy to train the agent. At the same time, incorporate a focused control strategy in the state feature extraction stage to strengthen the weight of temperature measurement data until the reward value converges to a preset stable threshold. S3: The echo state network-predictive control collaborative algorithm inputs the collected boiler operating parameters into the multi-scale coupled three-dimensional CFD model in real time, obtains real-time flue gas temperature distribution data, and calculates the deviation between the actual temperature and the set value. Using time-series data composed of real-time temperature data, temperature deviation, and the previous adjustment amount as input, the algorithm captures the dynamic change law of boiler temperature through the storage tank and outputs the initial compensation adjustment amount under the current operating condition. Using boiler equipment safety constraints and temperature regulation response speed as constraints, the algorithm optimizes and corrects the output initial adjustment amount in real time to generate the final compensation adjustment amount. The algorithm then controls the corresponding compensation actuator to operate according to the final adjustment amount.

2. The method according to claim 1, characterized in that, In S1, the macroscopic gas-solid two-phase flow-combustion reaction-radiative heat transfer coupling algorithm uses the Eulerian-Lagrange method to describe the gas-solid two-phase flow. The gas phase is modeled in Eulerian coordinates, and the flow velocity and turbulence intensity distribution are calculated through a turbulence model. The solid phase is tracked in Lagrange coordinates, and the particle motion trajectory, concentration distribution, and momentum exchange with the gas phase are calculated. The combustion reaction adopts the eddy dissipation concept model to quantify the mixing reaction rate of gas phase fuel and oxygen. Combined with the heterogeneous combustion reaction rate of coke, the total heat release during the combustion process is calculated. The radiative heat transfer adopts the discrete coordinate method, and the regional radiative heat transfer is calculated based on the radiative characteristics of radiative gases and coal particles in the flue gas.

3. The method according to claim 1, characterized in that, In S1, the microscopic coal powder combustion kinetics algorithm is designed for the entire life cycle of single-particle coal powder combustion. In the coal powder pyrolysis stage, a dual-competitive reaction model is adopted to establish pyrolysis kinetic equations for easily volatile components and less volatile components, respectively, and to calculate the release rate and composition ratio of volatiles at different temperatures. In the coke combustion stage, an Arrhenius-type reaction rate equation is adopted, and a particle pore structure factor is introduced to correct the reaction area, quantify the reaction rate of coke with oxygen and carbon dioxide, and generate the coke combustion exothermic law. In the sulfur release stage of coal, a segmented kinetic model is adopted to distinguish the release characteristics of organic sulfur and inorganic sulfur, and to calculate the influence of sulfur release on particle specific heat capacity. The microscopic parameters of particle pyrolysis rate, coke reaction exothermic rate, and particle temperature change are mapped to the macroscopic calculation grid of the multi-scale coupled three-dimensional CFD model through the volume averaging method.

4. The method according to claim 1, characterized in that, In S2, the state space design process of the intelligent agent is as follows: taking the key information of the temperature field output by the multi-scale coupled three-dimensional CFD model as the core, the predicted temperature of the furnace outlet, the predicted temperature of the economizer inlet, and the predicted temperature of the air preheater inlet are selected, and the deviation value between the actual flue gas temperature and the preset set value is included. The adjustment amount of the compensation actuator at the previous moment is also added to form the state space.

5. The method according to claim 1, characterized in that, In S2, the action space of the intelligent agent is based on the actual operable means of boiler flue gas temperature regulation, and the adjustment amount of the secondary damper opening, the adjustment amount of the economizer bypass damper opening, and the adjustment amount of the burner swing angle are selected as the core dimensions of the action space. The adjustment of the secondary damper opening regulates the flame center temperature by changing the air distribution ratio in the combustion zone, the adjustment of the economizer bypass damper opening regulates the tail flue temperature by changing the flow rate of flue gas through the economizer, and the adjustment of the burner swing angle directly adjusts the furnace outlet temperature by changing the flame center height. The value range of the adjustment amount in the action space matches the safe operating threshold of the corresponding actuator, and the step size setting of adjacent adjustment amounts is based on the adjustment accuracy and equipment response speed.

6. The method according to claim 1, characterized in that, In S2, the design logic of the reward rule is as follows: a multi-objective weighted reward function is adopted, the first objective is the temperature compensation effect, and the reward value increases monotonically as the deviation between the actual temperature and the set value decreases. The second objective is equipment operational stability, and the reward value increases monotonically as the difference between the current adjustment and the previous adjustment decreases. The third objective is boiler operational economy, and the reward value increases monotonically as the real-time combustion efficiency of the boiler increases. The combustion efficiency is calculated based on the oxygen content of the flue gas and the exhaust gas temperature. The priorities of the three objectives are balanced by preset weighting coefficients, which are dynamically adjusted according to the boiler's operational needs.

7. The method according to claim 1, characterized in that, In S2, the specific implementation process of using the trained agent parameters under the boiler operating condition as initial parameters is as follows: First, a reference boiler similar to the current boiler operating condition is selected. The similarity judgment criteria include rated evaporation capacity, volatile matter characteristics of the coal type, consistent burner type, and overlapping load adjustment range. The operating rules of the reference boiler are comparable to those of the current boiler. Then, the core parameters of the trained agent of the reference boiler are extracted, including the convolutional layer weights and fully connected layer biases of the deep Q network, and the core parameters are used as the initial parameters of the agent.

8. The method according to claim 1, characterized in that, In S2, the detailed implementation process of the phased training strategy is as follows: In the first phase, the agent is trained offline to master the basic rules of boiler temperature regulation. Specifically, the training dataset is constructed by integrating the historical operating data accumulated by the boiler over a long period of time with the virtual operating condition data generated by the multi-scale coupled three-dimensional CFD model. The agent is iteratively trained using a fixed learning rate. The second stage involves online fine-tuning to adapt the agent to the actual operating characteristics of the current boiler. Specifically, this is achieved by collecting real-time operating data and adjustment effect data of the current boiler and updating the network parameters of the agent at preset intervals. During the update process, a decaying learning rate is used until the mean reward completely converges to a preset stable threshold.

9. The method according to claim 1, characterized in that, In S2, a focused control strategy is incorporated into the state feature extraction stage. By calculating the attention weight of the temperature measurement point data in the state space, the weight calculation process introduces the temperature field spatial gradient information output by the multi-scale coupled three-dimensional CFD model, and establishes the correlation mapping between the attention weight and the boiler operating condition. Based on the sliding window, the deviation contribution of the temperature measurement point data is statistically analyzed in real time.

10. The method according to claim 1, characterized in that, In S3, the predictive control is specifically model predictive control. By constructing a predictive model for boiler temperature regulation, and based on historical operating data and simulation data from the multi-scale coupled three-dimensional CFD model, the dynamic relationship between operating parameters, regulation amounts, and temperature changes is quantified. A prediction time domain and a control time domain are set. The prediction time domain is used to predict the temperature change trend, and the control time domain is used to determine the current sequence of regulation amounts to be optimized. The constraints include boiler equipment safety constraints, secondary air pressure lower limit, and burner swivel angle adjustment rate upper limit. By solving a constrained optimization problem, the initial adjustment value output by the echo state network is corrected to generate the final adjustment value.

11. The method according to claim 1, characterized in that, In S3, the construction process of the echo state network is as follows: the core component is a reservoir, which is composed of interconnected neurons. The connection weights between neurons are randomly generated and remain fixed during training. The weights of the output layer are adjusted. By setting the spectral radius of the reservoir and the sparsity of the input weight matrix, the reservoir can capture the dynamic change patterns of time-series data in the state space.

12. The method according to claim 1, characterized in that, In S3, the action of the compensation actuator is based on the difference in response characteristics of the compensation actuator to formulate a coordination rule: the secondary air damper is used for rapid adjustment, the economizer bypass damper is used for transitional adjustment, and the burner swing angle is used for deep adjustment; a real-time temperature feedback correction mechanism is introduced, and after the actuator is activated, the temperature change data of the corresponding area is collected in real time. If the rate of decrease in temperature deviation is lower than the preset threshold, the adjustment range is adjusted immediately.

Citation Information

Cited By

  • Heat storage control method and control system based on flue gas parameter fluctuation

    CN121857281A

  • Thermal energy storage control method and control system based on flue gas parameter fluctuations

    CN121857281B

  • Coke oven whole furnace heat flow assignment method based on computational fluid dynamics method

    CN122433622B