Boiler heat storage peak regulation optimization control method based on load prediction

By adopting a boiler thermal storage peak-shaving optimization control method based on load forecasting, the problem of the disconnect between economic optimization and physical operation in the thermal system is solved. This method achieves a balance between safety and economy when the system experiences sudden load changes, and improves the system's adaptability and energy utilization efficiency.

CN121323017APending Publication Date: 2026-01-13HUADIAN HUTUBI ENERGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511468730.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing thermal system scheduling methods are disconnected from economic optimization and physical operation, failing to accurately capture the spatiotemporal delay effects and pressure fluctuations in heat energy transfer, leading to operational risks and energy waste.

Method used

A boiler thermal storage peak-shaving optimization control method based on load forecasting is adopted. By combining probabilistic heat load forecasting, heating robustness constraints and flexible state corridor planning with real-time control and online parameter identification, a closed-loop control strategy is formed to cope with the dynamic changes of the system.

Benefits of technology

This ensures the system has sufficient safety margin when facing sudden load changes, reduces operational risks, achieves a balance between economy and safety, and improves the system's adaptability and energy efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121323017A_ABST
    Figure CN121323017A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of thermodynamic system intelligent control, and discloses a boiler heat storage peak regulation optimization control method based on load prediction, and the method comprises the steps: firstly obtaining a probabilistic heat load prediction result in a future planning time domain; then based on a prediction result and a system physical model, an elastic state corridor of the charging state of the heat storage equipment is generated through robust optimization solution so as to deal with load uncertainty, then under the boundary constraint of the corridor, the system executes a real-time control strategy so as to maximize economic operation benefits, and finally, the control strategy is optimized. A strategy-model mismatch signal is calculated by monitoring the execution effect of a real-time control strategy, and when the signal exceeds a preset threshold value, online parameter identification and correction of a system physical model are triggered. According to the method, the problems of load uncertainty and long-term drifting of the model are solved by constructing the self-adaptive closed loop, and the economical efficiency and the self-adaptive capacity of long-term operation of the heat supply system are remarkably improved on the premise that the safety robustness of the heat supply system is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent control of thermal systems, and particularly to a boiler heat storage peak regulation optimization control method based on load prediction. BACKGROUND

[0002] Regional central heating systems are key infrastructure for ensuring urban energy supply. The core goal of their operation and dispatch is to achieve the optimal economy of system operation on the premise of meeting user heat load demand. Existing heating network dispatch methods mostly use economic optimization models based on heat load prediction to develop operation plans for key devices such as heat source output and water pump flow by solving mathematical programming problems with the goal of minimizing fuel and electricity consumption costs.

[0003] In the implementation process of the prior art, there is a significant disconnection between economic optimization dispatch and the physical operation process of the heating network. The models used for economic optimization usually highly simplify the complex hydraulic and thermal dynamic processes of the heating network, and cannot accurately capture the time and space delay effects of heat transmission in the pipe network and the dynamic propagation characteristics of pressure fluctuations. Therefore, the feasibility of the dispatch plan generated solely based on economic indicators in the future actual physical process cannot be guaranteed. Executing such a plan may result in over-limit pressure at specific nodes of the pipe network or substandard heating temperature at the end users, etc., which poses operational risks. To avoid such risks, a wide and fixed static safety margin must be used in actual operation, which limits the depth of the exploration of the system energy-saving potential and results in unnecessary energy consumption. Although high-fidelity simulation technologies such as digital twinning have emerged, their applications are mostly limited to offline analysis or passive verification of already developed plans, and they cannot form an active and forward-looking information closed loop with the optimization decision-making process, so the optimization dispatch is still carried out without accurate prediction of the future physical dynamics. SUMMARY

[0004] To address the deficiencies of the prior art, the present application provides a boiler heat storage peak regulation optimization control method based on load prediction, which solves the problem of disconnection between economic optimization and physical operation safety caused by model simplification, as well as the problem of the inability of the control model to adapt to system dynamic changes autonomously in long-term operation.

[0005] To achieve the above purpose, the present application is implemented by the following technical solution: a boiler heat storage peak regulation optimization control method based on load prediction, comprising the following steps:

[0006] obtaining probabilistic heat load prediction results in a future planning time domain and obtaining the current real-time state of the boiler heat storage system;

[0007] generate an elastic state corridor of the charging state of the thermal storage device in the planning time domain based on the probabilistic heat load prediction result and the system physical model, the elastic state corridor including time-varying upper and lower boundaries;

[0008] execute a real-time control strategy according to the current real-time state and real-time electricity price under the boundary constraint of the elastic state corridor, and generate and issue direct control instructions for the boiler and the thermal storage device;

[0009] monitor the system state after execution of the direct control instructions, and calculate a strategy-model mismatch signal for representing the execution effect of the real-time control strategy;

[0010] trigger online parameter identification and correction of the system physical model when the strategy-model mismatch signal exceeds a preset trigger threshold.

[0011] Preferably, the probabilistic heat load prediction result is specifically a set of heat load quantile prediction curves in a future planning time domain.

[0012] Preferably, in the step of generating the elastic state corridor, the optimization solving is performed under a heating robustness constraint condition, and the heating robustness constraint is:

[0013] ensuring that the total heating capacity of the system is not lower than an extreme high load quantile prediction value in the probabilistic heat load prediction result at all times.

[0014] Preferably, the real-time control strategy is a constraint reinforcement learning strategy, and the goal of the strategy is to maximize economic operation benefits under the boundary constraint of the elastic state corridor.

[0015] Preferably, the strategy-model mismatch signal is calculated in the following manner:

[0016] obtaining a generation value of the constraint reinforcement learning strategy generated due to the system state touching the boundary of the elastic state corridor;

[0017] and calculating a moving average value of the generation value in a preset time window to obtain the strategy-model mismatch signal.

[0018] Preferably, the step of online parameter identification and correction specifically includes:

[0019] locking recent high-frequency process data when the strategy-model mismatch signal exceeds the preset trigger threshold;

[0020] and updating parameters of the system physical model by using the high-frequency process data by means of an online identification algorithm.

[0021] Preferably, in the step of generating the elastic state corridor, the optimization solution is also constrained by the model confidence region;

[0022] The model confidence region is a region in the system state space where the prediction variance of the system physical model is lower than a preset variance threshold.

[0023] The objective function for optimization includes a penalty term, which is proportional to the distance by which the planned state trajectory deviates from the model's confidence region.

[0024] Preferably, the heating robustness constraint further includes an adaptive risk aversion coefficient, the value of which is proportional to the distance of the state trajectory from the model confidence region, and is used to dynamically adjust the required reserved safety margin.

[0025] Preferably, the method further includes:

[0026] Calculate the cumulative value of the martingale used to characterize the historical cumulative forecast bias;

[0027] When the cumulative value of the risk martingale continuously exceeds the preset risk threshold, an information gain term is temporarily introduced into the objective function of the optimization solution to generate an active detection action for actively stimulating the system dynamics, and the data generated by the active detection action is used to target and update the physical model of the system.

[0028] A boiler thermal storage peak-shaving optimization control method based on load forecasting, characterized by comprising:

[0029] The online forecasting unit is used to obtain probabilistic heat load forecast results within the future planning time domain;

[0030] The upper-level robust planning unit, connected to the online prediction unit, is used to generate an elastic state corridor of the charging state of the thermal storage equipment in the planning time domain by optimizing the solution based on the probabilistic heat load prediction results and the system physical model.

[0031] The lower-level real-time execution unit is connected to the upper-level robust planning unit and is used to execute real-time control strategies and generate and issue direct control commands under the boundary constraints of the elastic state corridor.

[0032] The system also includes a mismatch identification and model adaptation unit, which is connected to the lower-level real-time execution unit. This unit monitors the execution effect of the real-time control strategy and calculates the strategy-model mismatch signal. When the strategy-model mismatch signal exceeds a preset trigger threshold, it triggers online parameter identification and correction of the system's physical model.

[0033] This invention provides an optimized control method for boiler thermal storage peak shaving based on load forecasting. It has the following beneficial effects:

[0034] 1. This invention plans a flexible state corridor by employing probabilistic heat load prediction combined with heating robustness constraints. This approach ensures that the planned operating boundary already incorporates extreme uncertainties in the load, guaranteeing that the system still has sufficient safety margin when facing sudden high load conditions and avoiding the risk of heating interruption due to inaccurate predictions.

[0035] 2. This invention uses the cost generated when the lower-level real-time control strategy touches the boundary of the upper-level planned corridor to infer the degree of deviation between the current physical model and the actual dynamics of the system. Once the deviation exceeds the threshold, online parameter identification is triggered, forming a closed loop of planning-execution-monitoring-correction. It can autonomously cope with the drift of operating conditions such as equipment aging and efficiency decline, ensuring the accuracy of the model in long-term operation.

[0036] 3. This invention adopts a two-layer optimization architecture. The upper-layer robust planning unit is responsible for handling complex uncertainties and formulating safe and flexible state corridors, decoupling safety constraints from economic scheduling. The lower-layer real-time execution unit only needs to execute constraint reinforcement learning and other strategies within the given safety boundary to maximize economic benefits based on real-time electricity prices. This reduces the decision-making difficulty of the lower-layer real-time control, enabling it to respond quickly to changes in electricity prices and achieve optimal economic operation while ensuring safety. Attached Figure Description

[0037] Figure 1 This is a flowchart of the optimized control method of the present invention;

[0038] Figure 2 This is a functional block diagram of the optimized control system of the present invention;

[0039] Figure 3 This is a schematic diagram of the elastic state corridor and real-time control trajectory of the present invention;

[0040] Figure 4 This is a schematic diagram of the strategy-model mismatch monitoring and triggering mechanism of the present invention. Detailed Implementation

[0041] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0042] Please see the appendix Figure 1 - Appendix Figure 4This embodiment provides a boiler thermal storage peak-shaving optimization control system based on uncertainty quantification and dual adaptation. The system is deployed on an industrial control computer or server and communicates with field equipment of the thermal system through a data interface. The system may include a data acquisition and preprocessing unit 10, a probabilistic load prediction unit 20, an upper-level robust planning unit 30, a lower-level real-time execution unit 40, and a mismatch identification and model adaptation unit 50.

[0043] The data acquisition and preprocessing unit 10 connects to the field programmable logic controller or distributed control system via a data interface to acquire real-time and historical process data of the thermal system. This data includes heat load, ambient temperature, electricity price, boiler operating power, thermal storage tank liquid level and temperature, and water pump frequency. The data acquisition and preprocessing unit 10 performs data cleaning, outlier handling, timestamp alignment, and formatting operations on the acquired raw data to form structured time-series data, and then transmits the processed data to other functional units within the system.

[0044] The data input terminal of the probabilistic load forecasting unit 20 is connected to the output terminal of the data acquisition and preprocessing unit 10. Based on historical load and meteorological data, the probabilistic load forecasting unit 20 uses forecasting models such as quantile regression to calculate and output the probabilistic forecast of the heat load for a future planning time domain T. This result is represented in the form of a set of quantile curves.

[0045]

[0046] in,

[0047] L q (t) represents the future time t;

[0048] The probability that the heat load is lower than this value is q;

[0049] Q is a predefined set of quantiles.

[0050] The probabilistic load forecasting unit 20 will use the result Transmit to the upper-level robust planning unit 30.

[0051] The upper-layer robust planning unit 30 is used for rolling optimization on long-term timescales (e.g., hours). Its function is not to generate specific device action sequences, but rather to plan a time-varying, flexible state corridor for the thermal storage tank charging state SoC for the lower-layer real-time execution unit 40. min (t),SoC max (t)]. The input terminals of the upper-layer robust planning unit 30 are connected to the probabilistic load prediction unit 20, the lower-layer real-time execution unit 40, and the mismatch identification and model adaptation unit 50, respectively. The upper-layer robust planning unit 30 is based on the received probabilistic load prediction results. The control capability reach set information R fed back by the lower-level real-time execution unit 40, and the system physical model parameters θ updated by the mismatch identification and model adaptation unit 50, are used to solve a multi-objective robust optimization problem, and the calculated elastic state corridor is transmitted to the lower-level real-time execution unit 40.

[0052] The lower-level real-time execution unit 40 is used for fine-grained real-time control on short-cycle timescales (e.g., minutes or seconds). The input of the lower-level real-time execution unit 40 receives the elastic state corridor issued by the upper-level robust planning unit 30 and the real-time system state provided by the data acquisition and preprocessing unit 10. The lower-level real-time execution unit 40 integrates a constraint reinforcement learning agent, which, under the boundary constraints of the elastic state corridor, outputs direct control commands such as boiler power and pump frequency to the field equipment based on real-time load and electricity price. Simultaneously, the lower-level real-time execution unit 40 evaluates its own local control capability margin, forms reachability set information R, feeds it back to the upper-level robust planning unit 30, and calculates the cost c incurred due to state variables touching the corridor boundary. t The value of this generation c t The signal is transmitted to the mismatch identification and model adaptation unit 50.

[0053] The mismatch identification and model adaptation unit 50 is used to monitor and correct the accuracy of the system's physical model online. The input of the mismatch identification and model adaptation unit 50 receives the cost c transmitted by the lower-level real-time execution unit 40. t The mismatch identification and model adaptation unit 50 continuously monitors the generation value c. t The moving average value is used to determine whether there is a significant mismatch between the physical model and the real system. When a mismatch is detected, the mismatch identification and model adaptation unit 50 automatically triggers an online parameter identification algorithm to re-estimate the system physical model parameters θ using recent high-frequency process data. This update process can be implemented by algorithms such as recursive least squares, and its core update formula is:

[0054]

[0055] in,

[0056] It is the parameter estimation vector at time t;

[0057] K(t) is the gain vector;

[0058] e(t) is the prediction error based on the parameters at the previous time step.

[0059] The mismatch identification and model adaptation unit 50 will update the model parameters. and its uncertainty covariance matrix P θ(t+1) is transmitted to the upper-level robust planning unit 30, thus forming a closed loop of model adaptation.

[0060] See attached document Figures 1-4 This embodiment provides a boiler thermal storage peak-shaving optimization control method based on uncertainty quantification and dual adaptive methods. This method is executed in a periodic rolling manner and may include the following steps:

[0061] S100: Acquire real-time status data of the thermal system at the current moment, including the charging status of the thermal storage tank; and based on historical data, use a probabilistic prediction model to generate probabilistic heat load prediction results for the future planning time domain.

[0062] S200, based on probabilistic heat load prediction results, control capability reachability information fed back from the lower layer, and the current system physical model, generates an elastic state corridor for the charging state of the thermal storage tank in the planning time domain by solving a multi-objective robust optimization problem. The elastic state corridor includes time-varying upper and lower boundaries.

[0063] Under the boundary constraints of the elastic state corridor, the S300 generates and executes direct control commands for boilers and thermal storage equipment by adopting real-time control strategies such as constraint reinforcement learning based on the real-time system status and real-time electricity price. At the same time, it evaluates and feeds back control capability accessibility information to the upper layer.

[0064] S400 monitors in real time the cost of system state variables caused by touching the elastic state corridor boundary after the execution of direct control commands, and calculates the policy-model mismatch signal based on the moving average of the cost value.

[0065] S500: Determine whether the strategy-model mismatch signal exceeds the preset trigger threshold; if so, lock the recent high-frequency process data and activate the online parameter identification algorithm to update the parameters of the system physical model so that the updated model parameters can be used for the flexible state corridor planning of the next planning cycle.

[0066] At the start of the execution cycle of the optimization control method, step S100 is executed. This step provides accurate initial state and uncertainty inputs for subsequent optimization planning. Specifically, this step may include:

[0067] S101, the system acquires the real-time state data of the thermal system at the current moment through the data acquisition and preprocessing unit 10, and constructs the data into a real-time state vector x(t). The real-time state vector x(t) includes at least the charging state SoC(t) of the thermal storage tank and the water supply temperature T of the heating network. sup (t), return water temperature T ret (t), Current operating power P of the boiler b (t) and the external ambient temperature Tamb (t).

[0068] x(t) = [SoC(t), T sup (t),x(t)=[SoC(t),T sup (t),T ret (t),P b (t),T amb (t),…] T .

[0069] S102, the system loads the system physical model parameter vector that was updated or maintained at the end of the previous control cycle from its internal system dynamic model library. and its uncertainty covariance matrix P θ (t-1). Parameter vector It includes key physical quantities that describe the dynamic characteristics of the system, such as:

[0070]

[0071] in,

[0072] k loss It is the heat loss coefficient per unit time of the thermal storage tank.

[0073] η b It refers to the thermal efficiency of an electric boiler.

[0074] Covariance matrix P θ This quantifies the confidence level of these parameter estimates, with the diagonal elements representing the variance of each parameter.

[0075] S103, the probabilistic load forecasting unit 20 performs probabilistic heat load forecasting to generate probabilistic distribution information of heat load within the future planning time domain T (e.g., 24 hours), rather than a single deterministic forecast value.

[0076] To ensure full transparency, the specific implementation methods of the probabilistic prediction model may include, but are not limited to, the following:

[0077] Firstly, a quantile regression neural network is used. This network takes historical load data, meteorological data (such as temperature, humidity, and wind speed), and date-type data (such as weekdays and holidays) as input. By setting different quantile targets (e.g., 0.1, 0.25, 0.5, 0.75, 0.9, 0.95), it directly outputs the predicted future heat load value corresponding to the quantile.

[0078] Secondly, a Gaussian process regression model is employed. This model not only provides the expected load forecast curve but also the variance of each forecast point, thus yielding a complete forecast probability distribution. This distribution can then be used to calculate the load value at any quantile.

[0079] Third, a Bayesian neural network based on Monte Carlo Dropout is employed. By enabling the Dropout layer in the network multiple times during the prediction phase for forward propagation, a set of different prediction results can be obtained. This set of results approximates the posterior distribution of the load. By performing statistical analysis on this set of results, the required quantile prediction can be obtained.

[0080] Regardless of the specific implementation method used, the probabilistic load prediction unit 20 ultimately outputs a set of heat load quantile curves for the future planning time domain. This set will be transmitted as uncertainty information to the upper-level robust planning unit 30.

[0081]

[0082] in,

[0083] L q (t) represents the predicted value of the heat load being lower than this value at time t in the future, with a probability of q.

[0084] T represents the total duration of the planning time domain, for example, T = 24 hours;

[0085] Q represents a preset set of quantiles, for example, Q = {0.1, 0.5, 0.9, 0.95};

[0086] L 0.95 (t) refers to the extreme high-load scenario that needs to be prioritized for protection.

[0087] After executing step S100, the system continues to execute step S200. This step is executed by the upper-level robust planning unit 30, whose core task is to plan an operating boundary for the system that can guarantee both operational safety and economy within the future planning time domain T based on uncertain inputs, namely, a flexible state corridor for the charging state of the thermal storage tank [SoC]. min (t),SoC max (t)]. This step may specifically include:

[0088]

[0089] in:

[0090] π represents the operating strategy to be optimized.

[0091] J(π) represents the total objective function value under policy π;

[0092] C elec (t) represents the unit electricity price at time t;

[0093] P b(t|π) represents the boiler's input electrical power at time t under strategy π;

[0094] This represents the load uncertainty distribution generated in step S100. Find the expected value;

[0095] I(π(t),P θ (t) represents the information gain term resulting from executing policy π(t) at time t;

[0096] P θ (t): The covariance matrix of the system physical model parameter vector θ at time t, whose trace or determinant reflects the overall uncertainty of the model parameters;

[0097] γ(t) represents the dynamic weighting coefficient used to balance economy and information gain.

[0098] To ensure full disclosure, the information gain term I(π(t),P) θ The specific implementation of π(t) aims to quantify the contribution of the operating strategy to reducing the uncertainty of model parameters. It can be specifically defined as the change in the model parameter covariance matrix P before and after implementing the strategy π(t). θ The expected reduction in the trace:

[0099]

[0100] The more complex the operation, the greater the information gain, as it generates richer data with a higher signal-to-noise ratio. The value of the dynamic weight coefficient γ(t) is positively correlated with the uncertainty of the model; for example, its value can be set to be the same as the trace (P) of the covariance matrix. θ The value of γ(t) is proportional to the objective function. When the model uncertainty is high, the value of γ(t) increases, making the objective function more inclined to choose active detection actions that can bring high information gain, even if the action is not optimal in short-term economics.

[0101] S202, when solving the optimization problem, constraints must be satisfied. This set of constraints integrates heating safety, physical feasibility, and equipment operation limitations. Specifically, the constraints include: First, heating robustness constraints. To ensure uninterrupted heating under high-load scenarios such as extreme weather, the total heating capacity of the system at any time t must be greater than or equal to the extreme high-load quantile value generated in step S100 using the probabilistic load forecast.

[0102] Secondly, control capabilities can reach set constraints. To ensure that the strategies planned at the upper layer can be physically executed by the lower-level real-time execution unit 40, the planned flexible state corridor SoC for the next time step is... min(t+1),SoC max [(t+1)] must be a subset of the control capability reachable from the lower layer under the current state x(t) and feedback.

[0103]

[0104] The reachability set R(x(t)) defines the range of states that, in the current state, regardless of any disturbances, the lower-level controller is capable of driving or maintaining the system state SoC in the next time step. This constraint quantifies the actual execution capability of the lower layer and feeds it back to the upper-level planning, preventing the planning results from deviating from physical reality. Thirdly, there are physical operation constraints on the equipment. This set of constraints includes capacity limits for thermal storage tanks, maximum and minimum power limits for boilers, power ramp-up rate limits, and charging / discharging power limits, etc.

[0105] SoC abs,min ≤SoC(t)≤SoC abs,max ;

[0106] P b,min ≤P b (t)≤P b,max ;

[0107] |P b (t)-P b (t-1)|≤ΔP b,max ;

[0108] in,

[0109] SoC abs,min and SoC abs,max These are the upper and lower limits of the absolute physical capacity of the thermal storage tank;

[0110] P b,min and P b,max This refers to the minimum and maximum operating power of the boiler.

[0111] ΔP b,max It is the maximum power change rate of the boiler.

[0112] S203 uses a numerical optimization solver to solve the aforementioned constrained multi-objective optimization problem. After completion, the system outputs two sets of core results:

[0113] 1. Flexible state corridor within the future planning time domain T [SoC] min (t),SoC max (t)]. The width of the corridor is dynamically related to the uncertainty of load forecasting (i.e., the difference between different quantile curves) and the uncertainty of model parameters.

[0114] 2. Active Probe Sequence. This sequence is usually empty and is only planned when the system determines that the long-term benefits of active probing (improved model accuracy) outweigh the short-term economic losses.

[0115] After generating the flexible state corridor in step S200, the system proceeds to step S300. This step is executed cyclically at a high frequency by the lower-level real-time execution unit 40. Its task is to perform fine-grained real-time control within the boundaries of the upper-level planning and provide feedback to the upper-level planning unit regarding its own execution capabilities. This step may specifically include:

[0116] S301, the system constructs and executes a constrained reinforcement learning agent to make decisions within a constrained Markov decision process. The elements of the decision process are defined as follows:

[0117] state space s t The state vector observed by the agent at each decision time t contains all the information needed to make the optimal economic decision and satisfy the security constraints.

[0118] s t =[SoC(t),L actual (t),C elec (t),SoC min (t),SoC max (t)] T ;

[0119] in,

[0120] SoC(t) represents the current real-time charging state of the thermal storage tank;

[0121] L actual (t) represents the current real-time monitored heat load;

[0122] C elec (t) represents the current real-time electricity price;

[0123] SoC min (t) and SoC max (t) is the lower and upper boundaries of the elastic state corridor at the current moment, issued by the upper robust planning unit 30.

[0124] Action space a t : A vector of direct control commands output by the agent and applied to the physical device, wherein the elements of the vector are continuous values.

[0125]

[0126] in,

[0127] P b (t) is the boiler's input electrical power command;

[0128] and These are the heat storage and heat release power commands for the thermal storage tank.

[0129] reward function r t The immediate reward signal from the environment after an agent performs an action, with the goal of maximizing economic benefits.

[0130] r t =-C elec (t)·P b (t)·Δt;

[0131] in,

[0132] Δt is the time step of the lower-level control.

[0133] Cost function c t Used to punish behaviors that violate core security constraints, especially behaviors that touch or cross the boundaries of a flexible state corridor.

[0134]

[0135] in,

[0136] SoC(t+1) is the execution action a t The next state of the system;

[0137] K penalty It is a preset, sufficiently large positive number to ensure that the agent prioritizes avoiding violations of corridor constraints during the learning process;

[0138] Cost c t The value will be continuously transmitted to the mismatch identification and model adaptation unit 50.

[0139] Constrained reinforcement learning agents can employ, but are not limited to, constrained policy optimization algorithms or reward-penalty shaping methods based on Lagrange multipliers. These methods can be combined with mainstream reinforcement learning algorithms such as deep deterministic policy gradient or proximal policy optimization. The agent's objective is to maximize the expected value of the cumulative reward while ensuring that the expected value of the cumulative cost remains below a certain safety threshold.

[0140] In step S302, the lower-level real-time execution unit 40, while executing control, online evaluates its own control capability reachability set R(x(t)) and quantifies this information, feeding it back to the upper-level robust planning unit 30 for dynamic constraint of the elastic state corridor in step S202. The evaluation of the control capability reachability set aims to answer the question of within which the actuator can maintain the thermal storage tank state SoC in the current system state x(t), regardless of future short-term disturbances. Specific implementations may include:

[0141] First, evaluation based on value functions. Using the state-action value function Q(s,a) already learned by the agent, by analyzing the value distribution corresponding to different actions in the current state, the range of state space that can guide the system to a high-value, low-cost region can be evaluated. This range is an approximation of the reachability set.

[0142] Secondly, it is based on short-time-domain forward simulation. Starting from the current state x(t), multiple short-time-domain (e.g., 5-10 time steps) Monte Carlo forward simulations are performed using the system dynamic model and the agent's control strategy. Random sampling of disturbances such as load can be incorporated into each simulation. The state space covered by all simulation trajectories constitutes an estimate of the reachability set.

[0143] After the evaluation is completed, the lower-level real-time execution unit 40 parameterizes the reachable set R(x(t)), for example, by extracting the maximum and minimum values ​​of the SoC dimension to form the reachable interval [SoC]. r each,min,SoC reach,max The interval information is then transmitted to the upper-level robust planning unit 30 as a dynamic constraint for its next planning cycle.

[0144] During the real-time control execution of step S300, the system executes steps S400 and S500 in parallel. These two steps constitute an adaptive correction closed loop, which is used to realize the online self-correction of the system's physical model.

[0145] S401, the mismatch identification and model adaptation unit 50 monitors in real time the cost c generated by the lower-level real-time execution unit 40 in step S301. t The stated cost c t It directly reflects the severity of the system state variables touching or crossing the boundary of the elastic state corridor in the upper-level planning.

[0146] S402, the system is based on the aforementioned cost c t Continuous observation, calculation strategy-model mismatch signal M signal (t). This signal, as an indirect performance indicator, is used to determine whether there is a significant and persistent deviation between the physical model upon which the upper-level planning relies and physical reality. The mismatch signal is calculated as follows:

[0147]

[0148] in,

[0149] 1(·) indicates an indicator function; the function value is 1 when the condition inside the parentheses is true, and 0 otherwise.

[0150] c iThis represents the cost or value generated at historical moment i.

[0151] W represents the size of the time window used to calculate the moving average, which is a preset positive constant.

[0152] C threshold This indicates the preset mismatch trigger threshold.

[0153] The calculation principle of this signal lies in the fact that a well-trained lower-level reinforcement learning agent should be able to effectively avoid penalties. If it frequently or continuously obtains high penalty values ​​over a continuous period of time (defined by window W), causing the moving average to exceed the threshold C... threshold This strongly suggests that the problem is not a failure of the lower-level control strategy, but rather an unreasonable design of the corridor boundaries in the upper-level planning. Since the corridor boundaries are calculated based on the system's physical model, this phenomenon directly indicates a mismatch between the model and reality.

[0154] S501, the system determines the policy-model mismatch signal M at each time step. signal The value of (t). If M signal If M(t) is 0, the system considers the current model accurate and continues to execute real-time control and monitoring; if M signal If (t) is 1, the system immediately triggers the online parameter identification and model adaptive correction process.

[0155] S502, when the model calibration process is triggered, the system first automatically locks and extracts high-frequency process data from the data acquisition and preprocessing unit 10 within the time window that led to the trigger, including equipment inputs (such as boiler power P). b ) and system status output (such as the charging status of the thermal storage tank SoC).

[0156] S503, the mismatch identification and model adaptation unit 50 activates its internal online parameter identifier, using the extracted data to re-estimate the parameter vector θ of the system physical model. For full disclosure, the specific implementation of the online parameter identifier can be a recursive least squares method with a forgetting factor. The execution steps of this method are as follows:

[0157] First, the nonlinear dynamic equations of the system are linearized near the current operating point and rearranged into a linear regression form y(t)=φ(t-1). T θ.

[0158] For example,

[0159] The measured value y(t) can be defined as: SoC(t) - (1-k l oss,old)·SoC(t-1);

[0160] Information vector φ(t-1) TIt can be defined as:

[0161] The parameter vector θ to be identified can be defined as: [η b ,…] T .

[0162] Secondly, iterative updates are performed using the following formula:

[0163] 1. Calculate the prediction error e(t):

[0164]

[0165] 2. Calculate the gain vector K(t):

[0166]

[0167] 3. Update the parameter estimation vector

[0168]

[0169] 4. Update the parameter covariance matrix P θ (t):

[0170]

[0171] in,

[0172] and P θ (t-1) represents the parameter estimates and covariance matrix of the previous time step;

[0173] λ is the forgetting factor, which takes values ​​between (0, 1) and is used to adjust the weight of historical data in the identification process.

[0174] S504, After identification is completed, the system will update the model parameter vector. The covariance matrix P θ (t) The model is transferred to the system dynamic model library of the upper-level robust planning unit 30 to overwrite the old model parameters. In this way, when step S200 is executed at the beginning of the next planning cycle, the system will automatically call the corrected model that is more in line with the current physical reality for optimization planning, thereby completing the closed loop of the entire adaptive control.

[0175] This invention provides an advanced control scheme that can effectively balance the economic efficiency of thermal system operation and the robustness of heating safety under complex uncertainty environments by integrating probabilistic uncertainty quantification, two-level optimization planning and real-time execution, and a dual adaptive correction mechanism of strategy and model.

Claims

1. A boiler thermal storage peak-shaving optimization control method based on load forecasting, characterized in that, Includes the following steps: Obtain the probabilistic heat load forecast results within the future planning time domain, and obtain the current real-time status of the boiler thermal storage system; Based on the probabilistic heat load prediction results and the system physical model, an elastic state corridor for the charging state of the thermal storage equipment within the planning time domain is generated through optimization. The elastic state corridor includes a time-varying upper boundary and a lower boundary. Under the boundary constraints of the elastic state corridor, a real-time control strategy is executed based on the current real-time state and the real-time electricity price, generating and issuing direct control commands for the boiler and thermal storage equipment. Monitor the system state after the execution of the direct control command, and calculate the policy-model mismatch signal to characterize the effect of the real-time control strategy. When the strategy-model mismatch signal exceeds a preset trigger threshold, online parameter identification and correction of the system physical model is triggered.

2. The boiler thermal storage peak-shaving optimization control method based on load forecasting according to claim 1, characterized in that, The probabilistic heat load prediction results are specifically a set of heat load quantile prediction curves within the future planning time domain.

3. The boiler thermal storage peak-shaving optimization control method based on load forecasting according to claim 1, characterized in that, In the step of generating the elastic state corridor, the optimization solution is performed under the condition of heating robustness constraint, which is: Ensure that the total heating capacity of the system is not lower than the extreme high load quantile predicted value in the probabilistic heat load prediction results at all times.

4. The boiler thermal storage peak-shaving optimization control method based on load forecasting according to claim 1, characterized in that, The real-time control strategy is a constraint reinforcement learning strategy, the goal of which is to maximize economic operating benefits under the boundary constraints of the elastic state corridor.

5. The boiler thermal storage peak-shaving optimization control method based on load forecasting according to claim 1, characterized in that, The strategy-model mismatch signal is calculated as follows: Obtain the cost of the constrained reinforcement learning policy when the system state touches the boundary of the elastic state corridor; The moving average of the cost value within a preset time window is calculated to obtain the policy-model mismatch signal.

6. The boiler thermal storage peak-shaving optimization control method based on load forecasting according to claim 4, characterized in that, The online parameter identification and correction steps specifically include: When the strategy-model mismatch signal exceeds the preset trigger threshold, recent high-frequency process data is locked. An online identification algorithm is used to update the parameters of the system's physical model using the high-frequency process data.

7. The boiler thermal storage peak-shaving optimization control method based on load forecasting according to claim 1, characterized in that, In the step of generating the elastic state corridor, the optimization solution is also constrained by the model confidence region; The model confidence region is a region in the system state space where the prediction variance of the system physical model is lower than a preset variance threshold. The objective function for optimization includes a penalty term, which is proportional to the distance by which the planned state trajectory deviates from the model's confidence region.

8. The boiler thermal storage peak-shaving optimization control method based on load forecasting according to claim 3, characterized in that, The heating robustness constraint also includes an adaptive risk aversion coefficient, the value of which is proportional to the distance of the state trajectory from the model confidence region, and is used to dynamically adjust the required safety margin.

9. The boiler thermal storage peak-shaving optimization control method based on load forecasting according to claim 1, characterized in that, The method further includes: Calculate the cumulative value of the martingale used to characterize the historical cumulative forecast bias; When the cumulative value of the risk martingale continuously exceeds the preset risk threshold, an information gain term is temporarily introduced into the objective function of the optimization solution to generate an active detection action for actively stimulating the system dynamics, and the data generated by the active detection action is used to target and update the physical model of the system.

10. A boiler thermal storage peak-shaving optimization control system based on load forecasting, applied to the boiler thermal storage peak-shaving optimization control method based on load forecasting as described in any one of claims 1-9, characterized in that, include: The online forecasting unit is used to obtain probabilistic heat load forecast results within the future planning time domain; The upper-level robust planning unit, connected to the online prediction unit, is used to generate an elastic state corridor of the charging state of the thermal storage equipment in the planning time domain by optimizing the solution based on the probabilistic heat load prediction results and the system physical model. The lower-level real-time execution unit is connected to the upper-level robust planning unit and is used to execute real-time control strategies and generate and issue direct control commands under the boundary constraints of the elastic state corridor. The system also includes a mismatch identification and model adaptation unit, which is connected to the lower-level real-time execution unit. This unit monitors the execution effect of the real-time control strategy and calculates the strategy-model mismatch signal. When the strategy-model mismatch signal exceeds a preset trigger threshold, it triggers online parameter identification and correction of the system's physical model.