Method for preventing main steam pressure fluctuation based on control of fire coal amount

By strengthening the event trigger control method combined with learning agents and deep neural networks, the coal-fired quantity is adjusted in real time, the hysteresis and adaptability problems of main steam pressure control of coal-fired boilers are solved, the accuracy of main steam pressure and multi-objective optimization are achieved, and the safety and economicality of boiler operation are improved.

CN120332748APending Publication Date: 2025-07-18GUIZHOU ELECTRIC POWER DESIGN INST +1
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510664975.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the main steam pressure control response of coal-fired boilers is lagging and has poor adaptability, making it difficult to take into account multiple optimization goals, and the parameters of traditional PID controllers are difficult to adjust. The complexity of the boiler combustion system leads to fluctuations in the main steam pressure and its economy is affected.

Method used

The event trigger control method based on reinforcement learning agents is adopted, combined with the deep neural network prediction model and dynamic threshold function, the coal burning volume is adjusted in real time to prevent main steam pressure fluctuations, and the robustness and adaptability are improved through online model correction, safety constraint checking and feedforward compensation mechanisms.

Benefits of technology

It improves the accuracy and prospectiveness of main steam pressure control, reduces pressure fluctuations, enhances the system's adaptability to disturbances, achieves multi-objective coordinated optimization and safety, and improves the overall benefits of boiler operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120332748A_ABST
    Figure CN120332748A_ABST
Patent Text Reader

Abstract

The invention discloses a method for preventing main steam pressure fluctuation on the basis of controlling the coal combustion amount. The method comprises the following steps: acquiring running state data of the coal-fired boiler in real time; based on the operation state data and a preset boiler dynamic prediction model, predicting a steam pressure sequence of the boiler in a future preset time range, and judging whether to trigger a control action or not in combination with an event triggering strategy obtained by learning of a reinforcement learning agent; if it is judged that the control action is triggered, the reinforcement learning agent outputs an adjustment instruction of the total coal feeding amount according to the current operation state data; and controlling the total coal feeding amount of the boiler according to the adjusting instruction. According to the method, the timing and the force of control intervention can be intelligently determined according to the real-time state and future trend prediction of the boiler, and the accuracy and the foresight of main steam pressure control are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for preventing main steam pressure fluctuations based on controlling the coal combustion amount, belonging to the technical field of boiler control in thermal power plants. Background Art

[0002] In coal-fired thermal power plants, the main steam pressure is a key parameter characterizing the balance between boiler combustion and steam turbine inlet steam, and its stability and economy directly affect the safe and efficient operation of the unit. Traditional main steam pressure control methods mostly adopt feedback control strategies based on PID, combined with feedforward signals for compensation. However, the boiler combustion system is a complex thermal process with large inertia, large lag, nonlinearity, multivariable coupling, and variable operating conditions. It is difficult to tune the parameters of traditional PID controllers, and it is difficult to adapt to wide-range load changes, coal quality fluctuations, and slow drifts in equipment characteristics, often resulting in large fluctuations in the main steam pressure under disturbances, or sacrificing some economy to ensure stability.

[0003] In recent years, event-triggered control has received attention because it can reduce unnecessary control calculations and communications and reduce actuator wear. Traditional event-triggered mechanisms usually rely on fixed, preset state thresholds or error thresholds. In the face of a complex dynamic system such as a boiler, its adaptability and optimization ability are limited, and it is difficult to find the best balance among control performance, energy consumption, and communication efficiency.

[0004] Data-driven methods, especially reinforcement learning, provide a new approach to solving complex control problems. Through interaction and learning with the environment, the RL agent can autonomously discover the optimal control strategy. Combining data-driven learning with event-triggered mechanisms is expected to enable the system to autonomously learn the optimal trigger conditions and control actions, thus transcending the limitations of human design. However, applying general data-driven event-triggered learning to the actual main steam pressure control of coal-fired boilers still faces many challenges, such as how to ensure the safety and robustness of the learning process and the final strategy, how to adapt to the online changes in boiler characteristics, and how to achieve the collaborative optimization of multiple operating objectives. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for preventing main steam pressure fluctuations based on controlling the coal combustion amount, aiming to solve the problems of lagging response, poor adaptability, and difficulty in balancing multiple optimization objectives in the main steam pressure control of coal-fired boilers in the prior art.

[0006] To achieve the above object, the present invention provides the following technical solution: A method for preventing main steam pressure fluctuations based on controlling the coal combustion amount, comprising the following steps:

[0007] S01. Acquire the operation status data of the coal-fired boiler in real time, wherein the operation status data includes the current main steam pressure, the main steam pressure change rate, the current total coal supply and the unit load instruction;

[0008] S02. Based on the operating status data and a preset boiler dynamic prediction model, predict the steam pressure sequence of the boiler within a preset time range in the future, and determine whether to trigger a control action in combination with the event triggering strategy learned by the reinforcement learning agent; the event triggering strategy is used to determine under what predicted steam pressure deviation conditions to perform control intervention;

[0009] S03. If it is determined that a control action is triggered, the reinforcement learning agent outputs an adjustment instruction for the total coal supply according to the current operation status data;

[0010] S04. Control the total coal supply of the boiler according to the adjustment instruction to prevent main steam pressure fluctuations.

[0011] As a preferred solution, the boiler dynamic prediction model is a deep neural network model trained based on historical operating data, such as a long short-term memory network (LSTM) or a gated recurrent unit (GRU) network, to capture the complex nonlinear dynamics of the boiler system.

[0012] As a preferred solution, the event triggering strategy includes a state-dependent dynamic threshold function δ learned by the reinforcement learning agent learned (s t ), when the expected deviation of the predicted steam pressure series is When the dynamic threshold is exceeded, a control action is triggered.

[0013] As a preferred solution, the training of the reinforcement learning agent is based on a carefully designed reward function R t The reward function comprehensively considers the pressure control accuracy (such as the negative value of the square of the deviation from the set point), the coal economy (such as the penalty for the total coal feed or adjustment amount), the adjustment amplitude of the control action (such as the negative value of the square of the adjustment amount to protect the actuator) and the event triggering frequency (such as giving a small fixed penalty for each trigger to avoid unnecessary frequent triggering).

[0014] As a preferred solution, the method further includes the step of online correction of the boiler dynamic prediction model. By continuously monitoring the prediction residual of the prediction model, when the residual exceeds a preset threshold, some parameters of the prediction model are fine-tuned online, or an auxiliary model is trained to predict and compensate for the residual, and the uncertainty of the prediction can be quantified to improve the adaptability of the model to changes in operating conditions.

[0015] As a preferred solution, the method further includes a safety constraint verification step. After the reinforcement learning agent generates an adjustment instruction, the corrected or uncertainty-considered prediction model is used to predict the short-term state trajectory (especially steam pressure and pressure change rate) after the execution of the instruction, and verify whether the preset hard safety constraints (such as maximum / minimum steam pressure, maximum pressure change rate, coal feeding amount limit) will be violated. If a violation is possible, the adjustment instruction is minimally corrected to satisfy the constraints, or a pre-designed and more conservative backup controller (such as a robust PID) is activated to ensure the short-term safety of the system. At the same time, an additional negative reward is given to the RL agent to prompt it to learn to avoid proposing potentially unsafe actions.

[0016] As a preferred solution, the method further includes a step of online identifying key disturbance parameters that affect the boiler combustion efficiency and dynamic characteristics but are difficult to directly measure, such as the effective calorific value (Q ar,eff ) of coal or the combustion characteristic index. The identified disturbance parameters are added to the state space of the reinforcement learning agent . The event-triggered dynamic threshold δ learned (s′ t ) and the policy π(a t |s′ t ) of the reinforcement learning agent are both based on this extended state, and a feedforward compensation coal feeding amount adjustment term ΔF coal,ff,t is calculated according to the identified disturbance parameters and their change rates to actively compensate for the disturbance effect.

[0017] As a preferred solution, the step of predicting based on the operating state data and the boiler dynamic prediction model further includes multi-time scale state prediction, such as short-term prediction (from several seconds to dozens of seconds, for fast parameters such as main steam pressure) and medium-term prediction (from several minutes to dozens of minutes, for medium-speed parameters such as metal temperature and drum water level). Correspondingly, the event-triggered logic also adopts a hierarchical structure, setting a fast event trigger and a medium-term event trigger to coordinate control loops with different dynamic characteristics, such as the coordination between coal feeding adjustment (fast) and feed water and desuperheating water adjustment (medium-term).

[0018] As a preferred solution, the method further includes a step of dynamically adjusting the weight coefficients w price,t of each target (such as pressure stability, coal combustion economy, emission level) in the reward function of the reinforcement learning agent according to real-time working conditions (such as load demand, identified coal quality), external economic factors (such as real-time electricity price E emission,t ) and environmental protection policies (such as emission cost C j,t ). These dynamic weights W t = [wpressure,t , …, w emission,t as an additional input to the reinforcement learning agent's policy network π(a t |s′ t , W t ), enabling the control policy to adapt to the current operating focus and achieve intelligent collaborative optimization among multiple objectives.

[0019] The beneficial effects of the present invention are as follows: Compared with the prior art,

[0020] 1. Through data-driven event-triggered learning, it can intelligently determine the timing and intensity of control intervention based on the real-time state and future trend prediction of the boiler, improving the accuracy and foresight of main steam pressure control and effectively reducing pressure fluctuations.

[0021] 2. By combining online model correction, safety constraint verification, and backup control mechanisms, the robustness and safety of the control system in the model mismatch, unknown disturbance, or RL policy exploration stage are significantly improved, ensuring that the boiler operation does not violate key safety constraints.

[0022] 3. By online identifying key disturbance parameters such as coal quality and introducing feedforward compensation, the adaptive ability of the system to these common and significant disturbances is enhanced, further improving the control accuracy and reducing the lag of relying solely on feedback control.

[0023] 4. By introducing a dynamic target weight adjustment mechanism, the control system can flexibly perform collaborative optimization among multiple objectives such as pressure stability, economic operation, and clean sewage discharge according to external conditions such as real-time electricity prices and environmental protection requirements, improving the comprehensive benefits and intelligent level of boiler operation.

[0024] 5. By adopting hierarchical event triggering and multi-time scale prediction, it can better coordinate the control loops with different dynamic response characteristics in the boiler system, avoid control conflicts, optimize resource allocation, and make the rapid control of the main steam pressure coordinated with the stability of medium-term processes such as water level and wall temperature. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a schematic overall flow chart of a method for preventing main steam pressure fluctuations based on controlling coal combustion amount according to the present invention;

[0026] Figure 2 is a schematic module structure diagram of a device for preventing main steam pressure fluctuations based on controlling coal combustion amount according to the present invention;

[0027] Figure 3 is a schematic implementation flow chart of safety constraint verification and intervention in the method according to the present invention;

[0028] Figure 4This is a schematic diagram of the dynamic target weight adjustment and collaborative optimization concept in the method of the present invention. Detailed implementation manners

[0029] The following further describes the software embodiments of the present invention with reference to the accompanying drawings. However, the present invention is not limited to the following embodiments. The "module" described in this specification, unless otherwise specified, may refer to a hardware entity that can perform a specific function, a software entity (such as a program code segment), or a combination of both. For a software entity, it may be an instruction stored on a computer-readable storage medium and executed by a processor.

[0030] Embodiment 1: A method for preventing main steam pressure fluctuations based on controlling coal combustion amount

[0031] Refer to Figure 1 , a method for preventing main steam pressure fluctuations based on controlling coal combustion amount provided by the present invention may specifically include the following steps:

[0032] Step S01: Real-time obtain the operation status data of the coal-fired boiler. These data are the basis for subsequent prediction and decision-making. Specifically, the operation status data s t may include but are not limited to: the current main steam pressure P steam,t , the main steam pressure change rate P steam,t (which can be obtained by differentiating or filtering P steam,t ), the current total coal feeding amount F coal,t , the unit load instruction L demand,t . According to specific application scenarios, it may also include other auxiliary parameters, such as the primary air pressure, the secondary air volume, the air supply temperature T air,t , the oxygen content in the flue gas O 2,t , etc., to more comprehensively characterize the operation status of the boiler. These data are usually collected in real time from the distributed control system (DCS) of the boiler. Mathematically, the state vector can be expressed as:

[0033]

[0034] where "..." represents other optional state parameters.

[0035] Step S02: Based on the operation status data and a preset boiler dynamic prediction model, predict the future steam pressure sequence, and combine the event trigger strategy learned by the reinforcement learning agent to determine whether to trigger a control action. First, use a pre-trained boiler dynamic prediction model M pred , according to the current state s t and a hypothetical or currently maintained control action (for example, maintaining the current coal feeding amount, or a tentative coal feeding amount adjustment a h ypoth etical,t ), predict the steam pressure sequence of the boiler within the next N time steps

[0036]

[0037] The prediction model M pred is preferably a deep neural network, such as a long short-term memory network (LSTM) or a gated recurrent unit (GRU), which can learn the complex non-linear dynamic characteristics of the boiler from a large amount of historical operation data.

[0038] Then, the event-triggering strategy judges whether control intervention is required. This strategy is not based on a fixed threshold, but is learned by a reinforcement learning agent to achieve the optimal triggering timing. One implementation is that the RL agent learns a state-dependent dynamic threshold function δ learned (s t ). When there exists a certain time point k∈[1, N] in the predicted future steam pressure sequence, and a certain measure of the expected deviation from the target set point P setpoint (for example, the expected value of the absolute deviation, or a risk measure considering prediction uncertainty) exceeds this dynamic threshold δ learned (s t ), then a control action is triggered. For example, the triggering condition can be:

[0039]

[0040] where E[·] represents the expectation. Learning this δ learned (s t ) function is to balance control performance, energy consumption, and actuator wear.

[0041] Step S03: If it is judged that a control action is triggered, the reinforcement learning agent outputs an adjustment instruction for the total coal feeding amount according to the current operation state data. When an event is triggered, a pre-trained reinforcement learning agent (for example, using algorithms suitable for continuous control such as deep deterministic policy gradient DDPG, soft actor-critic SAC, etc.) outputs an optimal coal feeding amount adjustment action a t (or an extended state s′ t , see the subsequent embodiments) according to the current complete state s t .

[0042] a t = π(s t )

[0043] where π is the policy network of the RL agent. This adjustment action a t = ΔF coal,t represents the recommended adjustment value for the current total coal feeding amount F coal,t .

[0044] The training objective of the RL agent is to maximize an accumulated reward. The reward function R t is carefully designed to guide the agent to learn the desired behaviors, usually including the following aspects:

[0045] · Pressure control accuracy R pressure : For example, -(P steam,t -P setpoint ) 2 , punishing the deviation of the steam pressure from the set value.

[0046] · Coal combustion economy R coal : For example, -(F coal,t ) 2 or -|ΔF coal,t |, punishing excessive coal consumption or unnecessary adjustments, and encouraging fuel conservation.

[0047] · Smoothness of control actions R action : For example, -(ΔF coal,t ) 2 , punishing drastic control actions to protect actuators such as coal feeders.

[0048] · Event-triggering cost R event : For example, giving a small fixed negative reward -C trigger for each triggered event to avoid overly frequent and unnecessary triggering and calculations.

[0049] The total reward function can be expressed as a weighted sum of these components:

[0050] R t = w1·R pressur e + w2·R coal + w3·R action + w4·R event

[0051] where w1, w2, w3, w4 are weight coefficients, which reflect the relative importance of different control objectives and need to be debugged or dynamically adjusted according to actual operating requirements (see the subsequent embodiments).

[0052] Step S04: Control the total coal feeding amount of the boiler according to the adjustment instruction. Apply the coal feeding amount adjustment instruction a t = ΔF coal,t output by the RL agent to the coal feeding system of the boiler, that is, the new coal feeding amount set value is F coal,t+1 = F coal,t + ΔF coal,t . In this way, the fluctuation of the main steam pressure can be actively prevented or quickly suppressed.

[0053] The present invention also includes the following improvement points:

[0054] Improvement Point 1: Online Model Calibration and Uncertainty Quantification. For the boiler dynamic prediction model M pred which may become inaccurate due to factors such as changes in operating conditions and equipment aging, an online model calibration mechanism is introduced. Continuously monitor the prediction performance of M pred , for example, calculate the actual steam pressure P steam,t and the previous predicted value to obtain the residual When the cumulative residual or instantaneous residual exceeds a certain threshold, start model calibration. The calibration methods can include:

[0055] · Perform online fine-tuning on some parameters of M pred (such as neural network) using recent (s, a, s′) data.

[0056] · Train a small auxiliary model (such as a simple linear model or a small neural network) to predict the prediction residual of M pred and use this predicted residual to compensate for the output of M : pred

[0057] Meanwhile, an uncertainty measure U(s pred , a t ) can be added to the prediction of M t , for example, by using a Bayesian neural network or integrating multiple models and observing the variance of their predictions. This uncertainty U t can be used to adjust the sensitivity of event triggering (for example, when U t is high, δ learned (s t ) can be reduced or an uncertainty penalty can be directly added to the bias term) and the conservativeness of subsequent safety constraint verification.

[0058] Improvement Point 2: Safety Constraint Verification and Backup Control. Referring to Figure 3 , to ensure the safety of the RL policy during exploration or when the model is inaccurate, after the RL agent 103 generates a candidate action a t , a safety constraint verification module SVM301 is introduced. The SVM uses the calibrated or uncertainty - considered prediction model 302 to predict the future short - term state trajectory 303 after executing a t , especially P steam and Then verify 304 whether there is a violation of the preset hard safety constraints (for example, ). If the SVM detects that the candidate action a t may lead to a violation (305 is judged as "yes"), then:​

[0059] · Try action correction 306: Search for a minimum correction action a' t near a t such that it satisfies the safety constraints.

[0060] · If no safe correction action can be found, or the system uncertainty U t is too high, activate a pre-designed backup controller 307 (such as a robust PID or MPC based on a simplified model), which generates a safe control action F coal,robust .

[0061] The final adjustment of the coal feeding amount to the boiler 308 can be expressed as:

[0062]

[0063] When the action proposed by the RL policy is corrected by the SVM or replaced by the backup controller, give the RL agent an additional negative reward signal, such as R' t = R t - w penalty · I override where I override is an indicator (1 when correction or replacement occurs), prompting it to learn to avoid proposing potentially unsafe actions.

[0064] Improvement point 3: Online identification and feedforward compensation of key disturbance parameters. To cope with slow but significant disturbances such as coal quality changes, an online identification module is introduced to estimate these key disturbance parameters D est that are difficult to measure directly, such as the effective calorific value Q of coal ar,eff . This can be done using easily measurable input-output data (such as coal feeding amount, steam flow rate, steam temperature, flue gas temperature, oxygen content, etc.) combined with a simplified boiler heat balance model (for example, D steam · (h steam - h feedwater ) ≈ η boiler (D est )· F coal · Q ar,eff ) or a data-driven parameter estimator (such as a Kalman filter). The identified disturbance parameters (such as ) are augmented into the state space of RL to form an extended state The policy π(a t |s' t ) and the prediction model M pred (s' t , a t ) of the RL agent are both based on this extended state. The event-triggered dynamic threshold δ learned(s′ t ) also depends on For example, when the coal quality is identified as deteriorating, δ learned can be automatically reduced to make the trigger more sensitive. In addition, according to the identified and its change rate calculate a feedforward compensation coal supply adjustment term ΔF coal,ff,t :

[0065]

[0066] where is the reference disturbance parameter value, K ff and K ff,deriv are the feedforward gains. The final coal supply adjustment is a′ final,t = π(s′ t ) + ΔF coal,ff,t .

[0067] Improvement Point 4: Multi-time scale prediction and hierarchical event triggering. The boiler system contains dynamic processes with different response speeds. Therefore, the prediction model can be extended to provide state predictions on different time scales, such as short-term predictions (in seconds) for fast parameters like main steam pressure and medium-term predictions (in minutes) for main metal temperatures, drum water levels, etc. Accordingly, design hierarchical event triggers: The fast event trigger T S is based on short-term predictions and triggers coal combustion adjustment at high frequencies; the medium-term event trigger T M is based on medium-term predictions and triggers the corresponding medium-term control loops (such as feed water, desuperheating water) at medium and low frequencies. This hierarchical coordination enables fast coal combustion adjustment to take into account the expected impacts of medium-term control actions.

[0068] Improvement Point 5: Dynamic weight adjustment for collaborative optimization of energy efficiency and emissions. Referring to Figure 4 , in order to achieve dynamic collaborative optimization of multiple objectives, introduce a dynamic objective weight vector W t = [w pressure,t , w coal,t , w action,t , w event,t , w emission,t 401. These weights are no longer fixed, but are dynamically generated by a high-level policy module 402 according to external inputs (such as real-time electricity price E price,t 403, emission cost C emission,t 404), grid peak shaving demand D grid,t ) and internal states (such as current load L demand,t , identified coal quality ). For example, when the electricity price is high, increase the weight related to coal combustion economy. The reward function of the RL agent becomes R t (Wt ) and its policy network will also use the current weight vector W t as an additional input π(a t |s′ t ,W t )(406). In this way, the policy can directly adjust its behavior according to the relative importance of current targets and output to the boiler system 407. The sensitivity δ learned (s′ t ,W t ) should also be affected by W t .

[0069] Improvement Point 6: Interpretability analysis and human-machine collaborative optimization. To enhance the transparency of complex control systems, an interpretable AI (XAI) module can be introduced. This module uses techniques such as SHAP and LIME to analyze the decision-making basis of RL policies and presents the results to operators through an enhanced human-machine interface (HMI), explaining why specific weights are selected at specific times, why events are triggered, and why specific actions are taken. Operators can supervise and fine-tune key parameters (such as weight generation logic, trigger threshold range) through the HMI, and their feedback (such as the evaluation of an action) can be quantified as an additional reward signal R uman_feedback or samples for imitation learning to guide the continuous optimization of RL policies: R total,t =R t (W t )+R human_feedback,t .

[0070] Improvement Point 7: Multi-boiler collaborative optimization of federated learning and privacy protection. For power plants with multiple boilers, a federated learning (FL) framework can be introduced. Each boiler acts as a client and trains its control model on local data. The update of model parameters (or the model itself) is sent to the central server for aggregation after being processed by privacy protection technologies such as differential privacy and homomorphic encryption to form a global model. The global model is then distributed to each client as the basis for its local model and can be fine-tuned individually. In this way, the experience of multiple boilers can be pooled without directly sharing sensitive operation data, accelerating the overall performance improvement.

[0071] Improvement Point 8: Digital Twin-driven Policy Reinforcement and Fault Rehearsal. Build a high-fidelity Digital Twin (DT) for the boiler. The main training process of the RL agent can be carried out in the DT, which provides a safe, repeatable, and accelerable environment. The learned RL policy can be stress-tested in the DT (simulating extreme working conditions, sensor or actuator failures) to evaluate its robustness boundary. Before the new policy is deployed to the actual boiler, it can perform "shadow running" on the DT, in parallel with the physical boiler, to verify its effectiveness and safety. The DT can also be used to assist federated learning to generate additional training data.

[0072] Improvement Point 9: Co-optimization of Hyperparameters and Structure Based on Evolutionary Algorithm. The performance of the entire data-driven event-triggered control system depends on the configuration of numerous hyperparameters (such as RL network structure, reward function weight range, trigger threshold function parameters, etc.) and structure parameters. An evolutionary algorithm (EA) can be introduced to perform collaborative and automated optimization of these parameters in the digital twin environment. Encode the system configuration as genotype individuals, and perform iterative evolution through operations such as fitness evaluation (running RL training and testing in the DT), selection, crossover, and mutation, and finally find the configuration scheme with the optimal performance.

[0073] Improvement Point 10: Continuous Learning and Concept Drift Detection. The characteristics of the boiler, the external environment, or the control target may undergo "concept drift". Introduce a concept drift detection module (CDDM) to continuously monitor indicators such as control performance, model prediction error, RL reward trend, and data distribution changes. When a significant drift is detected, start a continuous learning mechanism (such as elastic weight consolidation EWC, experience replay, etc.) to update the affected model components to adapt to the new situation, while reducing catastrophic forgetting of old knowledge and ensuring the long-term effectiveness of the control system in a dynamically changing environment.

[0074] Through the combination of the above steps and optional features, the method of the present invention can achieve intelligent, safe, efficient, and adaptive control of the main steam pressure of a coal-fired boiler.

[0075] Embodiment 2: A Device for Preventing Main Steam Pressure Fluctuations by Controlling Coal Quantity

[0076] Refer to Figure 2 , a device 200 for preventing main steam pressure fluctuations by controlling coal quantity provided by the present invention can be integrated into the boiler control system or used as an independent intelligent control unit. The device 200 may include:

[0077] Data acquisition module 201: It is used to obtain the operation status data s of the coal-fired boiler in real time through a sensor interface or from the distributed control system (DCS) of the boiler t . As described in Embodiment 1, these data include the current main steam pressure P steam,t, main steam pressure change rate Current total coal feed amount F coal,t , unit load command L demand,t , and optional other auxiliary parameters such as primary air pressure, flue gas oxygen content O 2,t and so on.

[0078] Event trigger decision module 202: This module is the core decision-making unit of the device and is internally configured with:

[0079] · Boiler dynamic prediction model M pred 202a: Store and run a pre-trained (or online calibratable) boiler dynamic prediction model. This model receives the current state data s t from the data acquisition module 201 and a hypothetical control action, and is used to predict the steam pressure sequence of the boiler within a preset future time range as described by the formula.

[0080] · Reinforcement learning agent (partial function) 202b: The RL agent here at least includes the event trigger strategy part learned by it, such as the state-dependent dynamic threshold function δ learned (s t ). This module is based on the output of the prediction model 202a and the current state s t , and applies event trigger logic to judge whether a control action needs to be triggered.

[0081] Control instruction generation module 203: This module is also embedded in or tightly coupled to the reinforcement learning agent. When the event trigger decision module 202 determines that a control action needs to be triggered, this module (or the policy network π of the RL agent) outputs an adjustment instruction a t for the total coal feed amount according to the current operating state data s t (or the extended state s') provided by the data acquisition module 201 t =ΔF coal,t , as described by the formula.

[0082] Coal combustion control module 204: This module receives the coal feed amount adjustment instruction a t from the control instruction generation module 203, and converts it into a specific control signal for the boiler coal feed actuator (such as the coal feeder speed controller), so as to adjust the actual total coal feed amount of the boiler.

[0083] Preferably, the device 200 may further include a processor 205 and a memory 206. Computer program instructions are stored in the memory 206, and these instructions include the code for implementing the functions of the above-mentioned modules, such as the boiler dynamic prediction model M predparameters and structures, the policy network parameters of the reinforcement learning agent, the event-triggering logic algorithm, as well as the online model correction algorithm, the safety constraint verification logic, the disturbance parameter identification algorithm, the dynamic weight adjustment logic, etc. described in the subsequent embodiments. When the processor 205 executes these instructions in the memory 206, it implements all or part of the functions of the data acquisition module 201, the event-triggering decision module 202, the control instruction generation module 203, and the coal combustion control module 204.

[0084] Further, the device may integrate the functional modules corresponding to one or more additional features described in Embodiment 1:

[0085] · Online model correction unit (which can be integrated within the event-triggering decision module 202 or be an independent unit): used to online adjust the parameters or structure of the prediction model 202a according to the deviation between the actual operation data and the output of the prediction model.

[0086] · Safety constraint verification and intervention unit 207: After the control instruction generation module 203 outputs an adjustment instruction, this unit intervenes and uses the prediction model 202a or a dedicated safety prediction model to verify whether the instruction will cause a violation of the preset safety constraints. If a violation is possible, the instruction is corrected or a backup controller 208 is activated. The backup controller can be an independent module within the device or call a standard controller in the DCS.

[0087] · Disturbance parameter identification unit 209: used to online estimate key disturbance parameters such as coal quality based on operation data, and provide the identification results to the event-triggering decision module 202 and the control instruction generation module 203 to achieve adaptive triggering and feedforward compensation.

[0088] · Dynamic target weight adjustment unit 210: dynamically adjusts the weights of each target in the reward function used by the RL agents 202b, 203 during learning and decision-making according to external inputs such as electricity prices, emission policies, and internal states.

[0089] · Human-machine interaction interface 211: used for operators to monitor the system operation, view XAI explanations, perform parameter fine-tuning, and provide feedback.

[0090] Data exchange and collaborative work are carried out between these modules through an internal bus or a communication interface.

[0091] The above content is a further detailed description of the present invention in combination with specific preferred implementation manners, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A method for preventing main steam pressure fluctuations based on controlling coal combustion amount, characterized in that, It includes the following steps: S01. Obtain the operation status data of the coal-fired boiler in real time. The operation status data includes the current main steam pressure, the main steam pressure change rate, the current total coal feeding amount, and the unit load command. S02. Based on the operation status data and a preset boiler dynamic prediction model, predict the steam pressure sequence of the boiler within a preset time range in the future, and combine with the event trigger strategy learned by the reinforcement learning agent to determine whether to trigger a control action. The event trigger strategy is used to determine under what predicted steam pressure deviation conditions to perform control intervention. S03. If it is determined to trigger a control action, the reinforcement learning agent outputs an adjustment instruction for the total coal feeding amount according to the current operation status data. S04. Control the total coal feeding amount of the boiler according to the adjustment instruction to prevent the main steam pressure from fluctuating.

2. The method according to claim 1, wherein The boiler dynamic prediction model is a deep neural network model trained based on historical operation data.

3. The method according to claim 1 or 2, characterized in that, The event trigger strategy includes a state-dependent dynamic threshold learned by the reinforcement learning agent. When the expected deviation between the predicted steam pressure sequence and the steam pressure target set point exceeds the dynamic threshold, a control action is triggered.

4. The method according to claim 1, wherein The training of the reinforcement learning agent is based on a reward function, and the reward function comprehensively considers the pressure control accuracy, the coal combustion economy, the adjustment range of the control action, and the event trigger frequency.

5. The method according to claim 1, wherein It also includes the step of online correcting the boiler dynamic prediction model: continuously monitor the prediction performance of the prediction model. When the prediction residual exceeds the preset threshold, perform online fine-tuning on the prediction model or use a residual compensation model for compensation.

6. The method according to claim 1, wherein It also includes a safety constraint verification step: after the reinforcement learning agent generates an adjustment instruction, use the corrected prediction model to verify whether the adjustment instruction will cause the boiler operation parameters to violate the preset safety constraints. If the verification shows possible violations, correct the adjustment instruction or activate a backup controller.

7. The method according to claim 1, wherein It also includes the step of online identifying the key disturbance parameters that affect the boiler combustion efficiency and dynamic characteristics; and supplement the identified disturbance parameters to the state input of the reinforcement learning agent, and use them to adaptively adjust the sensitivity of event triggering and realize dynamic feedforward compensation of the coal feeding amount.

8. The method according to claim 1, wherein The step S02 also includes multi-time scale state prediction and hierarchical event triggering: for the dynamic processes with different response speeds in the boiler system, perform state prediction on different time scales, and set corresponding hierarchical event triggers to coordinate the actions of the coal feeding amount adjustment with fast changes and the control loops such as feed water or desuperheating water with medium-term changes.

9. The method according to claim 1, characterized in that, It also includes the step of dynamically adjusting the target weights in the reward function according to the real-time working conditions, load demands, electricity prices, or emission policies, and using the dynamic weights as the input of the reinforcement learning agent strategy.

Citation Information

Cited By

  • Thermal power plant boiler steam pressure stability control method and system

    CN121050485A

  • A method and system for stabilizing the steam pressure of a boiler in a thermal power plant

    CN121050485B

  • Electrode charging variable flow thermal safety control method and system based on AI multi-objective optimization

    CN121070102A

  • Aerial imaging system attitude error compensation method

    CN121632208A