Building heating layered safety control method and system under condition of no historical operation data

CN122504901APending Publication Date: 2026-08-04NORTHEAST DIANLI UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NORTHEAST DIANLI UNIVERSITY
Filing Date
2026-05-14
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

[0007]本发明的目的在于提供一种无历史运行数据条件下的建筑供热分层安全控制方法及系统,以解决现有技术中存在的冷启动不安全、模型失配难以实时修正、教师控制与学习控制之间切换生硬以及动态电价下经济性与舒适性难以兼顾的问题

Benefits of technology

(1)不依赖部署前的大规模历史运行数据,在系统冷启动阶段即可投入运行,适合新建建筑、重新调试建筑以及历史数据缺失场景;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122504901A_ABST
    Figure CN122504901A_ABST
Patent Text Reader

Abstract

This invention discloses a tiered safety control method and system for building heating under conditions lacking historical operating data, belonging to the field of building energy and intelligent control. The method acquires indoor temperature, comfort boundary, meteorological and electricity price predictions, and online operating samples; based on these samples, it performs online physical adaptation on the building's gray box thermal model to obtain current model parameters and disturbance terms; a target temperature is generated by a rule-based teacher and a high-level learning strategy; then, the model predictive control teacher and the low-level learning strategy output actions respectively, and the teacher dependency weights are determined based on action divergence and overheating risk to achieve weighted fusion of actions and generate the final control command. This method can achieve safe, smooth, and economical online control of building heating during the cold start phase.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of building energy system control and intelligent control technology, and in particular to a method and system for tiered safety control of building heating under conditions where historical operating data is unavailable. The method is applicable to building heating systems employing heat pumps, circulating pumps, fans, radiator terminals, or other adjustable heating actuators, and is also suitable for building energy management scenarios that require balancing indoor thermal comfort and operational economy under dynamic electricity pricing conditions. Background Technology

[0002] With the development of building electrification and demand-side response, building heating systems are no longer simply static devices that aim to meet heat loads, but are gradually becoming flexible loads capable of participating in grid interaction and peak shaving and valley filling. Especially under dynamic electricity pricing scenarios, buildings themselves have thermal inertia, and theoretically, through reasonable preheating, slow release, and temperature regulation strategies, some energy consumption can be shifted to periods with lower electricity prices, thereby reducing operating costs and improving the overall efficiency of the system.

[0003] Among existing building heating control methods, model predictive control (MMDC) can explicitly handle comfort constraints, actuator constraints, and economic objectives, and has therefore been widely studied in the field of building energy conservation control. However, MMDC typically relies on relatively accurate thermal models and parameter identification results. When the system is newly built and put into operation, undergoing re-commissioning, or lacking historical operating data, model parameters are difficult to obtain in a timely manner, and disturbance terms are difficult to fully characterize. This makes MMDC susceptible to model mismatch and time-varying disturbances during the cold start phase, thereby reducing control effectiveness.

[0004] Reinforcement learning, which can progressively improve control strategies through interaction with the environment, is attractive for complex, nonlinear, and difficult-to-model systems. However, in comfort-sensitive scenarios such as building heating, untrained or undertrained strategies often suffer from problems such as unsafe exploration, large control fluctuations, and lack of interpretability of output in the early stages of deployment, resulting in high risks in practical applications.

[0005] To mitigate these risks, existing research has attempted to combine model predictive control with reinforcement learning, or to constrain learning strategies through safety filtering, hard switching, and action masking. However, existing solutions generally suffer from the following shortcomings: First, many methods still rely on offline training data or sufficient pre-deployment modeling, making them unsuitable for scenarios without historical operational data; second, hard switching or binary replacement mechanisms are often used between the teacher controller and the learning controller, resulting in abrupt transfer of control and potentially leading to overly conservative or discontinuous control; and third, there is a lack of a complete solution that unifies and couples high-level interpretable target generation, low-level constraint-aware control, and online physical adaptation.

[0006] Therefore, there is an urgent need for a building heating control method that can be deployed directly without historical operating data, while taking into account safety, interpretability, economy and online adaptive capabilities. Summary of the Invention

[0007] The purpose of this invention is to provide a method and system for tiered safety control of building heating under conditions without historical operating data, so as to solve the problems existing in the prior art, such as unsafe cold start, difficulty in real-time correction of model mismatch, abrupt switching between teacher control and learning control, and difficulty in balancing economy and comfort under dynamic electricity prices.

[0008] To achieve the above objectives, the present invention adopts the following technical solution.

[0009] A tiered safety control method for building heating under conditions of no historical operating data is applied to a building heating system, wherein the building heating system includes at least a building thermal environment object, a heating actuator, and a controller, and the method includes the following steps: Step 1: Obtain the indoor temperature, comfort boundary, outdoor temperature prediction, solar irradiance prediction, electricity price prediction, and online accumulated operating samples at the current control moment; Step 2: Based on the accumulated online running samples, perform online physical adaptation on the building ash box thermal model to obtain the current model parameters and effective perturbation terms; Step 3: Based on the rules, the teacher generates a rule target temperature and based on the high-level learning strategy, generates a learning target temperature. After applying safety constraints to the learning target temperature, the rule target temperature and the learning target temperature are merged to obtain the execution target temperature. Step 4: Based on the target temperature and the building ash box thermal model updated online physically, the model predicts and controls the teacher to solve for the teacher's actions, and the low-level learning strategy outputs the learned actions. Step 5: Based on the action divergence between the teacher action and the learning action, and the overheating risk of the current indoor temperature relative to the future upper limit comfort boundary, determine the teacher dependency weight, and use the teacher dependency weight to perform weighted fusion of the teacher action and the learning action to obtain the final supervisory control instruction. Step 6: Map the final supervisory control command to the control quantity of the heating actuator and execute it to achieve safe online control of the building heating system under conditions of no historical operating data.

[0010] The online physical adaptation includes: using a gray box thermal model to characterize the building's dominant thermal dynamics; performing a moving time-domain parameter update on the gray box thermal model within an estimation window composed of multiple recent control samples to obtain updated thermal resistance, heat capacity, solar heat gain related parameters, baseline perturbation terms, and heat supply mapping parameters; updating the residual perturbation terms based on the single-step prediction error, and superimposing the residual perturbation terms with the baseline perturbation terms to form the effective perturbation terms, in order to compensate for the rapid time-varying perturbation between the two parameter updates.

[0011] The generation of the execution target temperature includes: generating a rule-based target temperature based on current and predicted electricity prices, meteorological information, comfort boundaries, and time period information; outputting a learned target temperature by a high-level learning strategy; performing safety trimming on the learned target temperature to ensure that it falls within a safe range defined by the comfort boundaries and the rule-based target temperature; determining the fusion weight between the rule-based target temperature and the safety-trimmed learned target temperature based on a preset preheating dependency factor and / or high-level value assessment results, and obtaining the execution target temperature accordingly.

[0012] The teacher dependency weight is jointly determined by the action divergence term and the overheating risk gating term. The action divergence term is obtained by mapping the difference between the teacher's action and the learning action. The overheating risk gating term is obtained by mapping the degree of proximity of the current indoor temperature to the upper limit comfort boundary within multiple future prediction steps. When the action divergence increases and / or the overheating risk increases, the teacher dependency weight is increased to smoothly transfer control to the model predicting and controlling the teacher.

[0013] The supervisory control command is mapped to a control quantity of at least one actuator among heat pumps, circulating pumps, fans, valves, radiators, or underfloor heating circuits to drive the building heating system to operate; after execution, the system writes new state transition data into the experience buffer and updates the high-level learning strategy and low-level learning strategy online.

[0014] The present invention also provides a tiered safety control system for building heating under conditions of no historical operating data, including a data acquisition and prediction unit, an online physical adaptive unit, a high-rise target generation unit, a teacher control unit, a low-level learning control unit, a risk gating fusion unit, and an actuator control unit, for implementing the above method.

[0015] Compared with the prior art, the present invention has at least the following beneficial effects: (1) It does not rely on large-scale historical operation data before deployment and can be put into operation during the system cold start phase, making it suitable for new buildings, re-debugged buildings, and scenarios where historical data is missing; (2) Through the overall architecture of online physical adaptation, high-level rule and learning fusion, and low-level risk gating continuous takeover, the deployability under model mismatch and time-varying perturbation conditions is improved, and the comfort risk caused by direct control of the learning strategy in the initial stage is reduced. (3) Instead of using a hard switch to replace the teacher controller and the learning controller, the teacher dependence weight is continuously adjusted according to the action divergence and overheating risk, which can achieve a smoother transfer of control and reduce abrupt changes and excessive conservatism. (4) Incorporating dynamic electricity price information into the generation of high-rise target temperature and the optimization of low-rise model prediction and control is beneficial to reduce operating costs, energy consumption and related emissions while ensuring indoor thermal comfort. (5) It can implement interpretable, safe, and online evolution control strategies on building heating, which is a constrained, time-varying, and thermally inertial object, and has good engineering promotion value. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the embodiments are briefly described below. The drawings are only used to schematically illustrate the technical concept of the present invention and do not constitute a limitation on the scope of protection of the present invention.

[0017] Figure 1 This is a schematic diagram of the overall process of the building heating stratified safety control method in an embodiment of the present invention.

[0018] Figure 2 This is a schematic diagram of the building heating tiered safety control system in an embodiment of the present invention.

[0019] in, Figure 2 The meanings of the reference numerals in the attached figures are as follows: 1-Data acquisition and prediction unit; 2-Online physical adaptive module; 3-High-rise target generation module; 4-Model prediction control teacher module; 5-Low-level learning control module; 6-Risk gating fusion module; 7-Actuator control module; 8-Building heating object; 9-Online learning and experience playback module. Detailed Implementation

[0020] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments described herein are for illustrative purposes only and are not intended to limit the invention. Unless otherwise specified, the technical features in the following embodiments can be combined with each other.

[0021] Example 1: System Overall Structure.

[0022] like Figure 2As shown, the building heating tiered safety control system of this embodiment includes a data acquisition and prediction unit 1, an online physical adaptive module 2, a high-rise target generation module 3, a model prediction control teacher module 4, a low-level learning control module 5, a risk gating fusion module 6, an actuator control module 7, a building heating object 8, and an online learning and experience playback module 9.

[0023] The data acquisition and prediction unit 1 is used to acquire the observations and external information required for building heating control, including but not limited to indoor temperature, upper and lower comfort limits, outdoor temperature prediction, solar irradiance prediction, electricity price prediction, historical control commands, time period information, and occupancy pattern information. The building heating object 8 can be a single-zone or multi-zone building heating system, specifically including heat pumps, circulating pumps, fan coil units, radiators, underfloor heating terminals, and the indoor thermal environment.

[0024] Example 2: Safety control process for tiered building heating.

[0025] like Figure 1 As shown, at each control moment, the controller sequentially performs steps such as data acquisition, online physical adaptation, target temperature generation, teacher action solving, learning action generation, risk gating fusion, and actuator mapping control. After the control is executed, the new state information is written into the experience buffer to support subsequent online updates.

[0026] Step 1 is used to obtain the indoor temperature, comfort boundary, outdoor temperature prediction, solar irradiance prediction, electricity price prediction at the current control moment, as well as the online accumulated operating samples after the start of control; Step 2 is used to update the building gray box thermal model based on online accumulated samples to obtain the current model parameters and effective perturbation terms; Step 3 is used to generate the target temperature for execution; Steps 4 and 5 are used to achieve action coordination of risk perception between the model predictive control teacher and the lower-level learning control; Step 6 is used to map the final supervisory control instruction into a specific actuator control quantity and execute it.

[0027] Example 3: Online physical adaptation.

[0028] The online physical adaptive module 2 identifies building thermal dynamics online based on the information provided by the data acquisition and prediction unit 1. Building thermal dynamics can be characterized using a gray box thermal model. In one embodiment, the gray box thermal model uses an equivalent thermal resistance-capacity model, whose continuous-time form can be expressed as: in, Indicates indoor temperature. The outdoor temperature is represented by R, the equivalent thermal resistance is represented by C, and the equivalent heat capacity is represented by C. This indicates the heat supply corresponding to the control input. This represents parameters related to solar heat gain. Indicates solar irradiance, This represents the effective disturbance term. To facilitate prediction and solution within the discrete control period, the above continuous-time thermal dynamics model can be discretized according to the control period Δt, yielding the state update relationship of the indoor temperature in the next control step with respect to the current indoor temperature, outdoor temperature, control input, solar irradiance, and effective disturbance term.

[0029] The This represents the mapping relationship between monitoring and control commands and the heating capacity. In one embodiment, the mapping relationship can be approximated by a quadratic polynomial as follows: in, , and These are the parameters for mapping heating capacity. In other embodiments, the mapping relationship can also be implemented using a piecewise linear function, a lookup table function, or other approximate models. The online physical adaptation module 2 is also used to perform mobile time-domain parameter updates. The parameter vector may include thermal resistance, heat capacity, solar heat gain-related parameters, baseline perturbation terms, and heat supply mapping parameters. In one embodiment, the parameter vector may be represented as: The controller selects the most recent N during each update. mhe An estimation window is constructed using control samples, and the current parameters are constrained and optimized based on the error between the predicted and measured indoor temperature values ​​within the estimation window. The constraint optimization process may further include: setting penalty terms for changes in thermal resistance and heat capacity relative to the previous update results; setting penalty terms for changes in baseline perturbation terms and heat supply mapping parameters; and setting upper and lower bounds for each parameter to keep the online identification results within a reasonable physical range.

[0030] Relying solely on moving-time domain parameter updates is insufficient to promptly reflect short-term disturbance changes. Therefore, this invention further incorporates a residual disturbance correction process. The controller updates the fast residual disturbance term based on the deviation between the current model's single-step temperature prediction and the actual indoor temperature collected at the next moment, and then superimposes it with the baseline disturbance term to form the current effective disturbance term. This effective disturbance term remains available to the model's predictive control teacher between two parameter updates, thereby improving short-term prediction effectiveness and enhancing adaptability to time-varying disturbances.

[0031] Example 4: Generating the temperature of a high-altitude target.

[0032] The high-level target generation module 3 simultaneously receives thermal parameter features provided by the online physical adaptation module 2, as well as external information provided by the data acquisition and prediction unit 1. The high-level target generation module 3 comprises two parts: a rule-based teacher and a high-level learning strategy.

[0033] The rule-making teacher generates the target temperature based on current and future electricity prices, outdoor temperature, solar irradiance, comfort boundaries, time period information, and occupancy patterns. The basic idea is to appropriately increase the target temperature during periods of low electricity prices, low overheating risk, or suitable preheating; and to maintain a relatively conservative target temperature setting during periods of strong solar heat gain or high overheating risk.

[0034] The high-level learning strategy outputs a target temperature based on the enhanced observation vector. Considering the initial instability of the learning strategy, this invention performs a safety trimming on the target temperature, ensuring it is neither lower than the safety lower bound determined by the lower comfort boundary and the rule-based target temperature, nor higher than the control upper bound. In one implementation, the controller also sets a preheating dependence factor α based on the occupancy mode. warm In the initial stage after a certain occupancy mode begins, α warm The alpha level is relatively high to ensure a strong reliance on rule-based teachers; as online control progresses, alpha... warm Gradually decrease.

[0035] Once a sufficient number of valid samples have accumulated in the experience buffer, a high-level value evaluation network can be activated to evaluate the long-term value of the rule-based target temperature and the learning target temperature after safe pruning, and to generate learning weights w in conjunction with the preheating dependency factor. learn The final target temperature can be expressed as: in, Indicates the target temperature of the rule. This represents the target temperature after safety-tuning. Thus, the system retains the interpretability of the prior rules while allowing the learning strategy to gradually develop its online adaptability.

[0036] Example 5: Coordination of low-risk gating actions.

[0037] Model predictive control teacher module 4 solves an optimization problem in the prediction time domain based on the current gray-box thermal model, effective disturbance terms, comfort boundaries, and target execution temperature to obtain the teacher action. The objectives of the optimization problem may include time-of-use electricity cost, target execution temperature tracking error, control smoothing term, and comfort violation relaxation. After solving, the first control variable in the obtained control sequence is preferably executed as the teacher action.

[0038] The low-level learning control module 5 outputs learning actions based on the enhanced observation sequence. To avoid simply selecting between teacher and learning actions using binary logic, this invention introduces a risk-gated fusion module 6. This risk-gated fusion module 6 first determines the action divergence based on the difference between the teacher and learning actions; then it determines the overheating risk based on the proximity of the current indoor temperature to the upper comfort boundary for several future steps; and finally, it generates teacher-dependent weights based on the action divergence and the overheating risk. The final supervisory control instruction can be expressed as: Therefore, when the divergence in actions increases or the risk of overheating rises, the system automatically increases its reliance on model-predictive control of the teacher; when the learning actions are relatively consistent with the teacher's actions and the risk of overheating is low, the system grants the learning control a higher degree of autonomy, thereby achieving a smooth, safe, and continuous transfer of control.

[0039] Example 6: Actuator Mapping and Online Learning.

[0040] Upon receiving a supervisory control command, the actuator control module 7 maps the command to a specific actuator control quantity. For example, in a hydroelectric heating system with a heat pump as the main heat source, the supervisory control command can be mapped to the heat pump compressor load, circulating pump speed, fan speed, or valve opening. When the supervisory control command is below the start-up threshold, the relevant auxiliary equipment can be shut down; when it is above the start-up threshold, the heat pump and auxiliary equipment are started synchronously according to a preset relationship.

[0041] After control execution, the data acquisition and prediction unit 1 reads the new indoor temperature, reward signal, and next state information, and writes information such as state transition, teacher action, rule target temperature, execution target temperature, teacher dependency weight, and physical parameters into the online learning and experience playback module 9. Both high-level and low-level learning strategies can use neural network structures with historical sequence encoding capabilities, such as temporal convolutional networks, recurrent neural networks, or Transformer structures, and are optimized online using off-policy update methods.

[0042] Example 7: A specific application case.

[0043] In a single-zone building heating application example, the control cycle can be set to 900 seconds, the model predictive control prediction time domain can be set to 24 control steps, the moving time domain parameter update window can be set to 96 control steps, and the update cycle can also be set to 96 control steps. From deployment onwards, the controller does not rely on historical operating data, and completes target temperature generation, teacher action solving, learning action generation, and risk gating fusion under the combined effect of dynamic electricity prices and weather forecasts. After continuous operation, the system can gradually establish more accurate thermal dynamic parameters and improve operational economy while ensuring indoor comfort boundaries. It should be noted that the above parameters are merely examples and do not constitute a limitation of the present invention.

[0044] It should be noted that this invention is not limited to single-zone building heating systems. The concepts of high-rise target generation, online physical adaptation, and low-level risk gating coordination in this invention can be extended to multi-zone buildings, multi-heat-source collaborative systems, air-source or ground-source heat pump systems, building energy systems with thermal storage devices, and flexible load scenarios interacting with the power grid. Equivalent substitutions or modifications made by those skilled in the art regarding the thermal model order, network structure, value function form, action mapping relationship, and parameter update solver, without departing from the spirit of this invention, should all fall within the protection scope of this invention.

Claims

1. A method and system for tiered safety control of building heating under conditions of no historical operating data, characterized in that, Applied to building heating systems, the building heating system includes at least a building thermal environment object, a heating actuator, and a controller, wherein the controller performs the following steps at each control moment: Step 1: Obtain the current indoor temperature, comfort boundary, outdoor temperature prediction, solar irradiance prediction, electricity price prediction, and online accumulated operating samples after control begins; Step 2: Based on the accumulated online running samples, perform online physical adaptation on the building ash box thermal model to obtain the current model parameters and effective perturbation terms; Step 3: Based on the rules, the teacher generates a rule target temperature and based on the high-level learning strategy, generates a learning target temperature. After applying safety constraints to the learning target temperature, the rule target temperature and the learning target temperature are merged to obtain the execution target temperature. Step 4: Based on the target temperature and the building ash box thermal model updated online physically, the model predicts and controls the teacher to solve for the teacher's actions, and the low-level learning strategy outputs the learned actions. Step 5: Based on the action divergence between the teacher action and the learning action and the overheating risk of the current indoor temperature relative to the future upper limit comfort boundary, determine the teacher dependency weight, and use the teacher dependency weight to perform weighted fusion of the teacher action and the learning action to obtain the final supervisory control instruction. Step 6: Map the final supervisory control command to the control quantity of the heating actuator and execute it to achieve safe online control of the building heating system under conditions of no historical operating data.

2. The method for tiered safety control of building heating according to claim 1, characterized in that, The online physical adaptation includes the following steps: Step 1: The dominant thermal dynamics of the building are characterized using a gray box thermal model; Step 2: Within the estimation window formed by the most recent multiple control samples, perform a moving time domain parameter update on the parameters of the gray box thermal model to obtain the updated thermal resistance, heat capacity, solar heat gain related parameters, baseline perturbation term and heat supply mapping parameters. Step 3: Update the residual disturbance term based on the single-step prediction error; Step 4: The residual perturbation term is superimposed with the baseline perturbation term to form the effective perturbation term, in order to compensate for the rapid time-varying perturbation between the two parameter updates.

3. The method for tiered safety control of building heating according to claim 1, characterized in that, The generation of the target temperature includes the following steps: Step 1: Generate the target temperature based on current and predicted electricity prices, meteorological information, comfort boundaries, and time period information; Step 2: The high-level learning strategy outputs the target temperature. Step 3: Perform safety trimming on the learning target temperature so that the learning target temperature is within a safe range defined by the comfort boundary and the rule target temperature; Step 4: Based on the preset preheating dependency factor and / or high-level value assessment results, determine the fusion weight of the rule target temperature and the learning target temperature after safety trimming, and obtain the execution target temperature accordingly.

4. The method for tiered safety control of building heating according to claim 1, characterized in that, The determination of the teacher dependency weights includes the following steps: Step 1: Determine the action divergence term based on the difference between the teacher's action and the learning action; Step 2: Determine the overheating risk gating item based on how close the current indoor temperature is to the upper limit comfort boundary within multiple future prediction steps; Step 3: Determine the teacher dependency weight based on the action divergence term and the overheating risk gating term; Step 4: When action divergence increases and / or the risk of overheating rises, increase the teacher dependency weight to smoothly transfer control to the model-predicted control teacher.

5. A tiered safety control system for building heating under conditions of no historical operating data, characterized in that, include: The data acquisition and prediction unit is used to acquire current indoor temperature, comfort boundary, weather forecast information, electricity price forecast information, and online accumulated operating samples after control begins; The online physical adaptive unit is used to perform moving time-domain parameter updates and residual perturbation corrections on the building ash box thermal model, so as to output the current model parameters and effective perturbation terms; The high-level goal generation unit is used to generate execution goal temperatures based on rule-based teachers and high-level learning strategies; The teacher control unit is used to solve for the teacher's actions based on the target temperature and the current model parameters; The low-level learning control unit is used to output learning actions; The risk gating fusion unit is used to determine the teacher dependency weight based on action divergence and overheating risk, and to perform weighted fusion of the teacher action and the learning action to obtain the final supervisory control instruction; An actuator control unit is used to map the final supervisory control command into a control quantity of the building heating actuator and execute it.