Boiler combustion multi-target cooperative control method based on PINN and reinforcement learning
By combining PINN and RL multi-objective self-optimization control methods, the problems of uneven coal powder distribution, combustion deviation and insufficient prediction accuracy in boiler combustion systems are solved, realizing high-precision prediction, fast response and intelligent control of boiler combustion systems, and improving the system's multi-objective collaborative optimization capability.
Patent Information
- Application Number
- CN202511721288.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2045-11-21
AI Technical Summary
Existing boiler combustion systems suffer from problems such as uneven coal powder distribution, severe combustion deviation, insufficient prediction accuracy, sluggish control system response and limitations of single-objective optimization, and a lack of multi-source information fusion and global coordination mechanisms.
A multi-objective self-optimization control method combining physical information neural network (PINN) and reinforcement learning (RL) is adopted. By embedding physical constraints such as energy conservation, fluid continuity and heat transfer equations, a multi-source sensor network is constructed for data fusion, and a reinforcement learning model is established for dynamic policy updates and multi-objective collaborative optimization.
It achieves high-precision prediction, rapid response and intelligent optimal control, improves the multi-objective collaborative control capability of boiler combustion system, enhances physical interpretability and self-learning adaptability, and realizes Pareto optimal trade-offs and real-time response for multiple objectives.
Smart Images

Figure CN121557514A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary field of thermal energy engineering and industrial artificial intelligence, and in particular to a multi-objective collaborative control method for boiler combustion based on PINN and reinforcement learning. Background Technology
[0002] Currently, coal-fired power plants still occupy a significant proportion of the power structure in various countries, and the operational performance of their boiler combustion systems directly affects the thermal efficiency, fuel consumption, and pollutant emissions of the generating units. As a crucial material conveying link in the boiler combustion system, the pulverized coal conveying network is responsible for transporting the ground pulverized coal to each burner via a gas-solid two-phase flow, determining the pulverized coal concentration and distribution at each burner, thus directly affecting the local combustion conditions, combustion efficiency, and NOx formation mechanism within the furnace. However, current coal-fired boilers still commonly face the following technical challenges during operation:
[0003] (1) In the actual coal conveying process, due to the asymmetrical structure of the branch pipes, the difference in elbow resistance, the fluctuation of coal particle size distribution, and the uneven primary air velocity, there are significant differences in the coal concentration at the inlet of each burner, forming "rich" and "lean" zones. This phenomenon not only causes the furnace flame to deflect and the temperature difference of the heating surface to increase, but also induces hidden dangers such as coking of the water-cooled wall, excessive thermal deviation, and local overheating. At present, it is difficult to maintain the uniformity of coal conveying under dynamic load changes by manually adjusting the dampers or using static air distribution curves. Uneven coal distribution remains an important bottleneck restricting the improvement of boiler combustion efficiency.
[0004] (2) Complex temporal coupling relationships exist among variables such as temperature field, pressure field, and oxygen distribution in the boiler combustion system. Traditional DCS systems mainly rely on empirical models or single-point linear extrapolation to estimate the furnace state, making it difficult to accurately capture instantaneous changes under unsteady-state processes. Although deep learning models such as LSTM and GRU can improve prediction performance, due to the lack of physical constraints such as energy conservation and fluid continuity, their results often show accumulated deviations when coal quality changes abruptly or operating conditions switch, affecting the effectiveness of the control strategy. Inaccurate predictions directly lead to a mismatch between pulverized coal feeding and actual combustion demand, resulting in long-term oversupply or undersupply in some branch pipes.
[0005] (3) Current pulverized coal feeding control systems generally adopt fixed-parameter PID or cascade PID structures. Their regulation strategies are based on a single variable (such as air volume or pulverized coal feed rate), which makes it difficult to handle multi-variable and strongly coupled boiler combustion systems. When the load changes rapidly, the PID parameters cannot self-tune in real time, resulting in system response lag, oscillation, or overcompensation, which leads to dynamic imbalance in pulverized coal distribution. At the same time, traditional control logic focuses on single-objective optimization (such as maintaining stable steam pressure or constant oxygen content), neglecting the multi-objective balance between thermal efficiency, emissions, and safety.
[0006] (4) The boiler combustion process is affected by multiple parameters such as temperature, pressure, flow rate, oxygen content, and pulverized coal concentration. However, the data of each sub-link in the current pulverized coal conveying control system are usually processed independently, resulting in weak information coupling. Asynchronous sampling of sensors, clock drift, and data delay make it impossible for the real-time control system to achieve high-dimensional state reconstruction and rapid decision-making. Especially in large units, the number of pulverized coal conveying pipes is large and the response time varies significantly. The control system lacks a unified adaptive coordination mechanism, making it difficult to form an intelligent closed loop from the furnace state to the pulverized coal conveying command.
[0007] In summary, the key technical challenges in this field can be summarized as follows: uneven coal powder distribution and severe combustion deviation; insufficient accuracy in predicting key parameters and lag in combustion regulation; slow response and limitations of single-objective optimization in the coal conveying control system; and lack of multi-source information fusion and global coordination mechanisms.
[0008] Therefore, there is an urgent need for a new control framework that effectively integrates physical mechanisms and data-driven methods. Summary of the Invention
[0009] The purpose of this invention is to address the shortcomings of the aforementioned background technology by providing a multi-objective self-optimizing control method that integrates Physical Information Neural Network (PINN) and Reinforcement Learning (RL). This method fully combines the advantages of physical mechanism models and data-driven intelligent algorithms. By embedding physical constraints such as energy conservation, fluid continuity, and heat transfer equations into the neural network, it ensures that the prediction process conforms to the thermodynamic laws of the boiler combustion system. This improves the prediction accuracy and interpretability of key parameters such as furnace temperature field, steam drum pressure, and pulverized coal flow rate, ultimately achieving high-precision prediction, rapid response, and intelligent optimal control of the entire system.
[0010] To achieve the above objectives, this invention provides a multi-objective cooperative control method for boiler combustion based on PINN and reinforcement learning, comprising the following steps:
[0011] S1 collects boiler combustion-related data through a multi-source sensor network and fuses edge data.
[0012] S2, based on physical information neural network, performs high-fidelity modeling and distributed training of furnace temperature field-pressure distribution, and predicts key physical fields inside the boiler through physical information neural network;
[0013] S3, based on the state features provided by the physical information neural network, constructs and trains a reinforcement learning model based on PINN state embedding, models the action space and execution mechanism of the reinforcement learning model, and sets a comprehensive reward function so that the reinforcement learning model can autonomously update the policy according to the real-time state and dynamically balance the multi-objective optimization requirements.
[0014] S4 uses a reinforcement learning model for multi-objective collaborative optimization control, constructs a multi-objective cost function, continuously evaluates the multi-objective cost function through the reinforcement learning model, and automatically adjusts the weights to achieve multi-objective collaborative control of boiler combustion.
[0015] Furthermore, the multi-source sensor network in S1 includes an infrared thermal imager, a pressure transmitter, a guided wave radar level gauge, a flue gas analyzer, and a high-precision flow meter to cover multi-dimensional operating data of boiler combustion, including furnace temperature field, steam drum parameters, flue gas composition, and pulverized coal conveying pipe flow. The data establishes a physical link with the edge computing node through an industrial Ethernet, and adopts a PTP-based clock synchronization mechanism. The collected raw data is initially verified, denoised, and validated for integrity in real time at the edge node. After being uniformly timestamped, it is written into the local time series database to form a timestamped multimodal time series dataset.
[0016] Furthermore, the three-dimensional transient temperature field of the furnace is selected in S2. Pressure distribution of primary air-pulverized coal two-phase flow As a collaborative prediction object, the spatial domain Covering the flue gas duct from the burner outlet to the furnace outlet, in the time domain It covers typical load disturbance cycles and uses spatiotemporal sampling data as supervision labels to establish a strongly supervised and weakly constrained model.
[0017] Furthermore, during model training in S2, the loss function includes data error and physical residual. Data error is the difference between the predicted value and the sensor's measured value; physical residual is the degree to which the predicted value does not satisfy the physical equation when substituted into the physical equation; and an adaptive weight strategy is used to dynamically update the loss function.
[0018] Furthermore, the loss function in S2 is:
[0019]
[0020] in, To monitor data errors, For the residuals of the governing equations, , These correspond to the boundary and initial conditions, respectively. These are the weighting coefficients. .
[0021] Furthermore, in S3, the three-dimensional transient temperature field of the furnace obtained in S2 will be... With the pressure field of primary air-pulverized coal two phases KL decomposition is performed on the edge side to extract intrinsic orthogonal mode coefficients, which, together with key measurable variables, constitute a multidimensional compact state vector. It is used for training reinforcement learning models.
[0022] Furthermore, continuous motion vectors are set in S3. Including the speed of the variable frequency powder feeder Secondary air damper opening on the second floor , Burnout damper opening induced draft fan guide vane angle Primary air main pipe pressure setting deviation The dynamics of each actuator are described using first-order inertia and rate saturation.
[0023] Furthermore, the comprehensive reward function set in S3 is as follows:
[0024]
[0025] in, To calculate the real-time combustion efficiency based on coal elemental analysis and flue gas heat loss, The NOx concentration at the chimney inlet. This represents the absolute value of the pressure deviation in the steam drum. The standard deviation of furnace cross-sectional temperature is used to quantify combustion stability. The total power consumption of the blower and powder feeder is used for economic penalties, and the weight vector consists of weighting coefficients. Adaptive scalarization is employed, and automatic updates are based on Pareto front curvature and operating procedures.
[0026] Furthermore, the multi-objective cost function in S4 is:
[0027]
[0028] in, For real-time combustion efficiency based on ASME PTC4-2008 reverse balance, Directly quantify energy loss. The NOx concentration in flue gas, converted to 6% O2, is used to characterize environmental constraints. The temperature variance of the furnace cross-section is used to reflect combustion stability and coking tendency. The steam drum pressure deviation characterizes the energy imbalance between the turbine and boiler. This represents the total auxiliary power consumption of the blower and powder feeder, reflecting the economic efficiency of operation; This is the weight vector.
[0029] Furthermore, in S4, the reinforcement learning agent is configured to interact with the weight vector. Implementing online adaptation includes: state expansion, action space output, and reward shaping.
[0030] The above-described solution of the present invention has the following beneficial effects:
[0031] The multi-objective cooperative control method for boiler combustion based on PINN and reinforcement learning provided in this invention achieves significant breakthroughs in four dimensions compared to traditional PID cascade control frameworks: physical interpretability, self-learning adaptability, real-time response, and multi-objective equilibrium capability. Specifically:
[0032] Significantly enhanced physical interpretability: By embedding energy conservation, gas-solid momentum equations, and radiation heat transfer models into the PINN loss function in residual form, the prediction results no longer depend solely on data distribution but strictly satisfy the first principles of thermodynamics and fluid mechanics. Under extreme conditions (low load, sudden changes in coal quality), the extrapolation error is significantly reduced compared to pure data-driven LSTM, and the source of temperature / pressure deviations can be quantitatively traced, providing operators with interpretable decision-making basis.
[0033] Superior self-learning and cross-domain transfer capabilities: Based on the weight adaptive mechanism of reinforcement learning, the agent can evolve its strategy online without offline retraining within a certain range of coal quality low calorific value changes and load instructions; after running continuously for 30 days, the average episode reward is significantly improved, verifying that its long-term adaptability is significantly better than that of fixed gain PID.
[0034] Multi-objective Pareto optimal trade-off: The dynamic weighted CBF-QP layer ensures that multiple objectives such as NOx, combustion efficiency, temperature variance, and auxiliary power consumption are at the feasible Pareto front in real time; 24-hour continuous operation data shows that compared with fixed weighted MPC, the Hypervolume index is improved by 19%, achieving a four-dimensional balance of "high efficiency, low emissions, stable combustion, and energy saving".
[0035] Improved real-time response: The edge-cloud collaborative architecture compresses the entire inference-optimization-execution latency to 12ms (a 41% reduction from PID's 20.3ms), significantly less than the pure lag in fuel delivery, effectively increasing the system phase margin and maintaining a steam pressure deviation of <±0.25MPa even at a variable load rate of 5%Pe / min.
[0036] Other beneficial effects of the present invention will be described in detail in the following detailed description section. Attached Figure Description
[0037] Figure 1 This is a flowchart of the steps of the present invention;
[0038] Figure 2 This is an industrial process diagram of boiler combustion involved in an embodiment of the present invention. Detailed Implementation
[0039] The following specific examples illustrate the implementation of this disclosure. Those skilled in the art can easily understand other advantages and effects of this disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. This disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0040] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using structures and / or functionalities other than one or more of the aspects set forth herein.
[0041] It should also be noted that the illustrations provided in the following embodiments are merely schematic representations of the basic concept of this disclosure. The illustrations only show components relevant to this disclosure and are not drawn according to the actual number, shape, and size of components in implementation. In actual implementation, the type, quantity, and proportion of each component can be arbitrarily changed, and the component layout may be more complex. Furthermore, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0042] like Figure 1 As shown, an embodiment of the present invention provides a multi-objective cooperative control method for boiler combustion based on PINN and reinforcement learning, comprising the following steps:
[0043] S1 collects boiler combustion-related data through a multi-source sensor network and simultaneously fuses edge data.
[0044] In this embodiment, the multi-source sensor network includes infrared thermal imagers, pressure transmitters, guided wave radar level gauges, flue gas analyzers, and high-precision flow meters, which are respectively arranged in corresponding positions to provide real-time coverage of multi-dimensional operational data such as furnace temperature field, steam drum parameters, flue gas composition, and pulverized coal conveying pipe flow rate. Simultaneously, these sensors establish high-bandwidth, low-latency physical links with edge computing nodes via industrial Ethernet, and employ a PTP-based clock synchronization mechanism to ensure precise alignment of the sampling times of each sensor, thereby guaranteeing the consistency and comparability of multimodal data across the time axis.
[0045] Meanwhile, in this embodiment, the collected raw data is initially verified, denoised, and validated for integrity at the edge nodes in real time. After being uniformly timestamped, it is written into the local time series database to form a structured and scalable timestamped multimodal time series dataset, providing a reliable data foundation for subsequent feature extraction, physical constraint modeling, online prediction, and closed-loop control.
[0046] S2 uses a Physical Information Neural Network (PINN) to perform high-fidelity modeling and distributed training of the furnace temperature field-pressure distribution, and predicts key physical fields inside the boiler through the PINN.
[0047] In this embodiment, the three-dimensional transient temperature field of the furnace is selected. Pressure distribution of primary air-pulverized coal two-phase flow As a collaborative prediction object, the spatial domain Covering the flue gas duct from the burner outlet to the furnace outlet, in the time domain Covering typical load disturbance cycles Spatiotemporal sampling data from an infrared thermal imager (sampling frequency 50Hz, 640×512 pixels) and a high-frequency pressure array (sampling frequency 1kHz, 128 measurement points) were used as supervisory labels to establish a strongly supervised-weakly constrained hybrid learning model.
[0048] During model training, the loss function consists of two parts: data error and physical residual. The data error is the difference between the predicted value and the sensor's measured value; the physical residual is the degree to which the predicted value does not satisfy the physical equation (residual).
[0049] Therefore, this embodiment sets up a PINN loss function, which explicitly introduces the residual term of the following conservation equation to ensure that the prediction results satisfy the first principles of thermodynamics and fluid mechanics. Specifically, the energy conservation (transient heat conduction-convection-radiation coupling) equation is:
[0050]
[0051] in, Where is the thermal diffusivity, Internal energy per unit mass Using the P1-approximate radiation model, This is a heat source for pulverized coal combustion.
[0052] The continuity equation for compressible fluids (low Mach number approximation) is as follows:
[0053]
[0054] in, The fluid density is given.
[0055] The momentum conservation equation (Navier-Stokes, transient-incompressible) is as follows:
[0056]
[0057] in, The momentum exchange source term for the particle phase is described by a two-way coupled Eulerian-Lagrange term.
[0058] Combining the above equations, the residuals are calculated at the sampling points and collocation points using automatic differentiation (AD). And embed the total loss in mean square form:
[0059]
[0060] in, To monitor data errors, For the residuals of the governing equations, , These correspond to the boundary and initial conditions, respectively. The weights are dynamically updated using an adaptive weighting strategy to ensure that the physical consistency error and the data fitting error converge synchronously during the training process.
[0061] It should be noted that in this embodiment, PINN uses a deep residual fully connected network with 8 hidden layers, each containing 256 neurons, and the activation function is... The output layer is set to linear. A parallel branching structure is also implemented, where the temperature and pressure branches share the first three layers, and the last five layers are independent, to reduce the number of parameters and improve convergence speed. A spatiotemporal point allocation strategy is employed. 1.2 × 10⁻⁶ sequences were generated using Sobol low-difference sequences within the specified range. 5 Multiple sampling points are used to ensure local density in high-gradient regions (burner outlet, flame peak), increasing the sampling point density by 3.6 times compared to uniform sampling. Encoder-decoder feature extraction is employed, and Fourier Feature Mapping is introduced to enhance the network's ability to represent high-frequency temperature fluctuations.
[0062] It should be noted that in this embodiment, model training is completed on an edge-cloud collaborative platform. The edge side is responsible for real-time data caching and initial cleaning, and then streaming the data to the cloud GPU cluster via gRPC. Distributed training is performed using the Horovod distributed framework, employing data parallelism and gradient compression, with a batch size set to 4.096 × 10⁻⁶. 4 The learning rate is set to 1×10. -3 Cosine annealing was also employed. Simultaneously, mixed-precision training was introduced, reducing the training time per epoch from 420 seconds to 95 seconds, with a total training time of approximately 16 hours for 600 epochs. The prediction error (RMSE) and physical residual were also improved. Converged to 3.8K and 1×10, respectively. -3 To meet the requirements of industrial control for consistency between model accuracy and mechanism (residual) () dual indicators.
[0063] S3. Construct and train a reinforcement learning model based on PINN state embeddings.
[0064] In this embodiment, the state space of the PINN prediction field → low-dimensional manifold representation is first constructed. Specifically, the three-dimensional transient temperature field of the furnace obtained in S2 is... With the pressure field of primary air-pulverized coal two phases KL decomposition was performed on the edge side to extract the coefficients of the first 16 intrinsic orthogonal modes (PODs). , The key measurable variables (steam drum pressure, flue gas oxygen content, NOx concentration, and furnace negative pressure) together constitute a 32-dimensional compact state vector. Therefore, this vector not only preserves the high-dimensional dynamic features of the physical field, but also significantly reduces the input dimension, which can improve the convergence speed of subsequent model training by about 40%.
[0065] Then, the action space and execution mechanism of the reinforcement learning model are modeled. Specifically, continuous action vectors are set. Includes: Variable frequency powder feeder speed (Adjustment accuracy ±0.1Hz); Secondary air damper opening on the second floor ; Burnout damper opening ; Guide vane angle of induced draft fan Primary air main pipe pressure setting offset The dynamics of each actuator are described using first-order inertia plus rate saturation:
[0066]
[0067] Among them, Rate limiting absolute value This ensures that the output can be directly mapped to the AO channel of the DCS system.
[0068] Set up a comprehensive reward function (multi-objective Pareto-guided):
[0069]
[0070] in, This is the real-time combustion efficiency (%) calculated based on coal elemental analysis and flue gas heat loss. NOx concentration at the chimney inlet (mg / m³) -3 ), This represents the absolute value of the steam drum pressure deviation (MPa). The standard deviation (K) of furnace cross-sectional temperature is used to quantify combustion stability. The total power consumption (kW) of the blower-feeder is used for economic penalties. The weight vector consists of weighting coefficients. Adaptive scalarization is employed, with automatic updates every 10 episodes based on the Pareto front curvature and runtime rules, ensuring long-term cumulative rewards. It balances the four dimensions of "high efficiency, low emissions, stable combustion, and energy saving" to enable reinforcement learning models to learn how to maximize long-term cumulative rewards.
[0071] Therefore, the reinforcement learning model using the distributed deep reinforcement learning algorithm in this embodiment possesses self-learning and online tuning capabilities, enabling it to autonomously update the policy based on real-time states and dynamically balance the optimization requirements of multiple objectives such as thermal efficiency, emissions, and stability. It should be noted that the network architecture of the distributed deep reinforcement learning algorithm is set as a dual Actor-Critic network, both using a 6-layer, 256-neuron ResNet-FC. The Actor output layer uses Tanh activation and is linearly mapped to the action boundary, while the Critic embeds a dual-channel attention mechanism after state-action concatenation, improving sensitivity to key coupling features. Simultaneously, the distributed deep reinforcement learning algorithm adopts a distributed parallel architecture based on the RayRLlib framework, with 8 workers sampling synchronously and a batch size B=512. Policy updates are run on the GPU, with a single step time of 18ms, meeting real-time requirements. Furthermore, the distributed deep reinforcement learning algorithm introduces safe reinforcement learning, using a CBF constraint layer to project the Actor output in real time, ensuring... Always meet the upper and lower limits of the air-to-coal ratio, the safe zone of the furnace negative pressure, and the saturation constraint of the feeder rate to achieve safe operation with "zero violations".
[0072] S4 uses a reinforcement learning model for multi-objective collaborative control.
[0073] In this step, the multi-objective cost function of the reinforcement learning model is first constructed, and the instantaneous generalized dimensionless cost function is defined:
[0074]
[0075] in, For real-time combustion efficiency (%) based on ASME PTC4-2008 reverse equilibrium, Directly quantify energy loss. NOx concentration in flue gas converted to 6% O2 (mg•Nm³) -3 ), used to characterize environmental constraints, The temperature variance of the furnace cross section (K) 2 This is used to reflect combustion stability and coking tendency. The pressure deviation in the steam drum (MPa) characterizes the energy imbalance between the boiler and turbine. This represents the total auxiliary power consumption (kW) of the blower and powder feeder, reflecting the economic efficiency of operation. All items are normalized and mapped to [0,1] to ensure dimensional consistency and gradient comparability.
[0076] As a preferred implementation, this embodiment further employs a reinforcement learning agent (Meta-RL, DDPG- For the weight vector Implementing online adaptation includes: state expansion, action space output, and reward shaping. State expansion involves adding load commands to the existing 32-dimensional physical field features. The inputs for the strategy are: coal quality (lower calorific value) and environmental assessment indicators (35 dimensions in total). The outputs are then used to define the action space. ,satisfy and convex combinations are ensured through softmax projection. Reward shaping is performed using... The exponential moving average decline rate as the meta-reward Guide the intelligent agent to improve under 30% BMCR (low-load stable combustion) conditions. , Improvement under 100% BMCR (high load assessment) conditions , This enables automatic matching of "working conditions and targets".
[0077] Therefore, the model continuously evaluates the multi-objective cost function during operation. Meta-RL is used to automatically adjust weights, achieving the effect of "automatically adjusting the optimization focus under different operating conditions," thus realizing multi-objective coordinated control of boiler combustion.
[0078] In this embodiment, a hard constraint layer based on the control barrier function (CBF) is introduced to ensure NOx 50mg•Nm -3 , 0.35MPa 15K. Real-time correction when any constraint approaches the boundary. This increases the corresponding weight index, ensuring the solution remains at the feasible Pareto front. Monitoring with the Hypervolume index shows a 27% improvement after 600 episodes, validating the improvement in multi-objective equilibrium.
[0079] As a preferred implementation, the system in this embodiment adopts an edge-cloud collaborative inference and millisecond-level closed-loop control execution mode. First, a two-level heterogeneous computing framework of "edge real-time layer - cloud acceleration layer" is constructed: the lightweight PINN after INT8 quantization is deployed and run on the edge side; the cloud is responsible for the policy update, rolling optimization and model distillation of the global reinforcement learning model, and maintains bidirectional communication with the edge nodes through a time-sensitive network. The single-hop ring network latency is <0.25ms, which meets the strict timing requirements of the sampled value message.
[0080] Specifically, edge nodes can receive local sensor data (32-dimensional POD coefficients + 5-dimensional actuator feedback) at 1ms intervals. After INT8 quantization and PINN forward propagation by the TensorRT engine, the low-dimensional modal coefficients of the furnace temperature and pressure fields in the next 200ms time domain are output. At the same time, UKF is used to correct the model-measurement deviation online with a correction step size of 10ms to ensure the mean square error of the prediction. Every second, the cloud collects compressed features (4.0kB total) from all 126 edge nodes in the plant, reconstructs the global state, and runs a multi-objective MPC-RL two-layer optimizer: the upper MPC solves the finite-time optimal control problem in 500ms steps, with constraints including the wind-to-coal ratio safety envelope, NOx emission limits, and steam drum pressure change rate limits; the lower RL (DDPG-ensemble) provides an initial feasible solution, shortening the MILP solution time to 18ms (Gurobi 10.0, GPU-accelerated). The optimization results are digitally signed and then sent to the edge via 5G-uRLLC (3GPPRel-17, 99.99% reliability, 4ms air interface latency).
[0081] After receiving the optimization results from the cloud, the edge node performs a secondary verification with the local security projection layer (CBF-QP, solver OSQP, iterations ≤30, time 0.8ms) to generate the final action. Then, it broadcasts the result to each frequency converter and servo driver via the EtherCAT bus at a period of 250µs, achieving a total closed-loop latency of ≤12ms for the entire "sensing-inference-decision-execution" chain, which is much smaller than the pure delay of fuel delivery (≈35ms), meeting the stringent requirements of boiler combustion control for phase margin.
[0082] A dual-loop system of "cloud distillation-edge increment" is adopted: every 1000 training steps completed in the cloud, a lightweight student network (halved in width and reduced by 2 layers in depth) is generated, and the accuracy is verified with KL divergence <0.01 as the threshold; it is then sent to the edge via OTA differential upgrade package (<180kB), and the hot update process is completed within 200ms. During the switching period, the output retains the value of the previous cycle to ensure no disturbance.
[0083] Simultaneously, a fault degradation mechanism is set up, defining three fault modes: L1 (communication packet loss > 3 frames): the edge automatically switches to the local backup PID to maintain a safe operating point; L2 (cloud optimization timeout > 50ms): edge MPC is activated (shortening the prediction time domain to 50ms) to ensure transient stability; L3 (dual network redundancy failure): the unit RB (Runback) logic is triggered, and the load rolls back to 70% BMCR at a rate of 5% Pe / min. All degradation events are reported to the DCS through the IEC62439-3PRP zero-switching redundancy bus, enabling fault traceability.
[0084] The following specific case further illustrates the effectiveness of this method. A system corresponding to this invention was independently deployed on the DCS cabinet side of a 600MW coal-fired power unit, running in parallel with the existing DCS via OPCUA PubSub, achieving "transparent bypass" verification. The demonstration period was from June 5th to July 4th, 2025, accumulating 720 hours of continuous operation, covering peak shaving across the entire range from 30% THA to 100% BMCR and three 50% step disturbance tests. The industrial process diagram is as follows. Figure 2 As shown.
[0085] The multi-source sensor network is deployed as follows: For the furnace combustion zone, 20 infrared thermal imagers are deployed, achieving continuous 24-hour temperature measurement through water-cooled jackets and compressed air purging, with a spatial resolution of 0.8 mrad and a temperature measurement uncertainty of ±1.5℃; For the primary air-coal duct, 10 fiber optic pressure sensors are installed, using white oil filling to isolate high-temperature particles, with a frequency response ≥5kHz; For the steam drum and downcomer, 4 guided wave radar level gauges are used for density correction after boiler drum pressure compensation to ensure the identification of false water levels under low load; For the flue gas side, a laser scattering NOx / O2 analyzer and a Venturi differential pressure flow meter (±1%RD) together form a mass balance closed loop. All sensor nodes achieve sub-microsecond clock synchronization via IEEE 1588v, and the sampled stream is aggregated to the edge computing node through a TSN switch, with end-to-end jitter <50µs.
[0086] A PINN model was constructed and trained. The RMSE of the predicted final furnace outlet temperature using the PINN model was reduced to 0.31℃, a 58% reduction compared to a pure LSTM model with the same structure; the physical residual converged to 1.3×10⁻⁶. -3 This meets the reliability requirements for engineering extrapolation.
[0087] When building a reinforcement learning model, the state Average temperature of furnace cross section Primary air main pipe pressure NOx concentration Steam drum pressure deviation Oxygen volume fraction (7 dimensions in total). Action. Variable frequency powder feeder speed Secondary air door angle The continuous output is mapped to the AO channel of the DCS system via a 4–20mA input. Sampling period: 200ms; Actuator inertia: 1.2s; Rate limiting ±3°s -1 The reward function is set as follows: The reward value is processed by a non-linear exponential moving average and then used for Critic updates to ensure smoothness and strategy stability.
[0088] When training the reinforcement learning model, the DDPG algorithm, Actor-Critic dual network, 256×2 hidden layers, Tanh activation, and experience replay with a capacity of 1×10⁻⁶ are used. 6 PER=0.6, linearly increasing from 0.4 to 1.0; training 12000 episodes (≈2.4×10⁻⁶). 6 (Step), cloud-based GPU parallel training with 8 workers, training time 36 hours; convergence metrics: average episode reward improvement 65%, action gradient norm <0.05, policy network parameter drift. 1×10 -3 .
[0089] Load disturbance tests showed that the intelligent agent could suppress NOx overshoot to <5% within 45 seconds, restore steam pressure deviation to ±0.15MPa, and reduce combustion efficiency fluctuation to <0.3%. Key performance indicators during the demonstration period were as follows: coal consumption for power generation reduced by 3.7gkWh. -1 (from 295.4 g kWh) -1 Reduced to 291.7 g kWh -1 This translates to an annual saving of approximately 6,500 tons of standard coal and a reduction in CO2 emissions of 1.68 × 10⁻⁶ tons. 4 t; the hourly average NOx concentration increased from 320 mg N / m -3 Reduced to 270 mg / Nm -3Ammonia consumption decreased by 11.2%; the standard deviation of furnace temperature section decreased by 23.7%, and the frequency of local high-temperature hot spots decreased from an average of 42 times per day to 9 times, effectively alleviating coking and high-temperature corrosion; the control response delay was shortened by 41%, the number of times the primary frequency regulation dead zone was triggered decreased from an average of 18 times per day to 3 times, and the unit flexibility index was improved by 22%; the edge-cloud collaborative framework operated without failure for 720 hours, with a data packet loss rate of <0.01%, meeting the requirement of DL / T1712-2017 for control system availability ≥99.9%.
[0090] Based on the same inventive concept, this embodiment also provides a system, including: a memory for storing a computer program; and a processor for executing the computer program to implement the relevant steps of the aforementioned multi-objective cooperative control method for boiler combustion based on PINN and reinforcement learning.
[0091] The processor may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor can be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor may also include a main processor and coprocessors. The main processor, also known as the Central Processing Unit (CPU), is used to process data in the wake-up state; the coprocessors are low-power processors used to process data in the standby state. In some embodiments, the processor may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor may also include an Artificial Intelligence (AI) processor, which handles computational operations related to machine learning.
[0092] The memory may include one or more computer-readable storage media, which may be non-transitory. The memory may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In this embodiment, the memory is used to store at least the following computer program, which, after being loaded and executed by the processor, is capable of implementing the aforementioned software method steps. In addition, the resources stored in the memory may also include operating systems and data, and the storage method may be temporary or permanent storage. The operating system may include Windows, Unix, Linux, etc.
[0093] This embodiment also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the relevant steps of the multi-objective cooperative control method for boiler combustion based on PINN and reinforcement learning as described above.
[0094] The system and computer-readable storage medium provided in this embodiment have the same inventive concept and beneficial effects as the aforementioned method, and will not be repeated here.
[0095] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0096] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A multi-objective cooperative control method for boiler combustion based on PINN and reinforcement learning, characterized in that, Includes the following steps: S1 collects boiler combustion-related data through a multi-source sensor network and fuses edge data. S2, based on physical information neural network, performs high-fidelity modeling and distributed training of furnace temperature field-pressure distribution, and predicts key physical fields inside the boiler through physical information neural network; S3, based on the state features provided by the physical information neural network, constructs and trains a reinforcement learning model based on PINN state embedding, models the action space and execution mechanism of the reinforcement learning model, and sets a comprehensive reward function so that the reinforcement learning model can autonomously update the policy according to the real-time state and dynamically balance the multi-objective optimization requirements. S4 constructs a multi-objective cost function based on a reinforcement learning model. The multi-objective cost function is continuously evaluated through the reinforcement learning model, and the weights are automatically adjusted to achieve multi-objective coordinated control of boiler combustion.
2. The multi-objective cooperative control method for boiler combustion based on PINN and reinforcement learning according to claim 1, characterized in that, The multi-source sensor network in S1 includes an infrared thermal imager, a pressure transmitter, a guided wave radar level gauge, a flue gas analyzer, and a high-precision flow meter to cover multi-dimensional operating data of boiler combustion, including furnace temperature field, steam drum parameters, flue gas composition, and pulverized coal conveying pipe flow. Data establishes a physical link with edge computing nodes via industrial Ethernet, using a PTP-based clock synchronization mechanism; The collected raw data is initially verified, denoised, and validated for integrity at the edge nodes in real time. After being uniformly timestamped, it is written into the local time series database to form a timestamped multimodal time series dataset.
3. The multi-objective cooperative control method for boiler combustion based on PINN and reinforcement learning according to claim 1, characterized in that, S2 selects the three-dimensional transient temperature field of the furnace. Pressure distribution of primary air-pulverized coal two-phase flow As a collaborative prediction object, the spatial domain Covering the flue gas duct from the burner outlet to the furnace outlet, in the time domain It covers typical load disturbance cycles and uses spatiotemporal sampling data as supervision labels to establish a strongly supervised and weakly constrained model.
4. The multi-objective cooperative control method for boiler combustion based on PINN and reinforcement learning according to claim 3, characterized in that, During model training in S2, the loss function includes data error and physical residual. Data error is the difference between the predicted value and the sensor's measured value; physical residual is the degree to which the predicted value does not satisfy the physical equation when substituted into the physical equation; and an adaptive weight strategy is used to dynamically update the loss function.
5. The multi-objective cooperative control method for boiler combustion based on PINN and reinforcement learning according to claim 4, characterized in that, The loss function in S2 is: ; in, To monitor data errors, For the residuals of the governing equations, , These correspond to the boundary and initial conditions, respectively. These are the weighting coefficients. .
6. The multi-objective cooperative control method for boiler combustion based on PINN and reinforcement learning according to claim 1, characterized in that, The three-dimensional transient temperature field of the furnace obtained in S2 will be used in S3. With the pressure field of primary air-pulverized coal two phases KL decomposition is performed on the edge side to extract intrinsic orthogonal mode coefficients, which, together with key measurable variables, constitute a multidimensional compact state vector. It is used for training reinforcement learning models.
7. The multi-objective cooperative control method for boiler combustion based on PINN and reinforcement learning according to claim 6, characterized in that, Set continuous motion vectors in S3 Including the speed of the variable frequency powder feeder Secondary air damper opening on the second floor , Burnout damper opening induced draft fan guide vane angle Primary air main pipe pressure setting deviation The dynamics of each actuator are described using first-order inertia and rate saturation.
8. The multi-objective cooperative control method for boiler combustion based on PINN and reinforcement learning according to claim 6, characterized in that, The comprehensive reward function set in S3 is as follows: ; in, To calculate the real-time combustion efficiency based on coal elemental analysis and flue gas heat loss, The NOx concentration at the chimney inlet. This represents the absolute value of the pressure deviation in the steam drum. The standard deviation of furnace cross-sectional temperature is used to quantify combustion stability. The total power consumption of the blower and powder feeder is used for economic penalties, and the weight vector consists of weighting coefficients. Adaptive scalarization is employed, and automatic updates are based on Pareto front curvature and operating procedures.
9. The multi-objective cooperative control method for boiler combustion based on PINN and reinforcement learning according to claim 1, characterized in that, The multi-objective cost function in S4 is: ; in, For real-time combustion efficiency based on ASME PTC4-2008 reverse balance, Directly quantify energy loss. The NOx concentration in flue gas, converted to 6% O2, is used to characterize environmental constraints. The temperature variance of the furnace cross-section is used to reflect combustion stability and coking tendency. The steam drum pressure deviation characterizes the energy imbalance between the turbine and boiler. This represents the total auxiliary power consumption of the blower and powder feeder, reflecting the economic efficiency of operation; This is the weight vector.
10. The multi-objective cooperative control method for boiler combustion based on PINN and reinforcement learning according to claim 9, characterized in that, In S4, the reinforcement learning agent is configured to assign weight vectors... Implementing online adaptation includes: state expansion, action space output, and reward shaping.
Citation Information
Patent Citations
Method and system for utility boiler combustion subspace modeling and multi-objective optimization
CN103576655A
Combustion system optimization design method based on physically-driven parameterized proxy model
CN116542164A
Intelligent boiler combustion optimization method based on module-level mechanism and mathematical hybrid model
CN118066563A
Boiler combustion strategy optimization method, system, equipment and medium
CN119983323A
Method and system for predicting temperature field of SNCR (selective non-catalytic reduction) reaction area of pulverized coal fired boiler
CN120126589A
Cited By
Aero-engine multi-module adaptive coupling simulation system and method
CN122151554A