Building demand response lightweight rule control method based on machine learning self-adaption
Through the adaptive building demand response lightweight rule control method based on machine learning, the problem that building demand response control strategies in the prior art are difficult to take into account both effectiveness and lightweight implementation requirements, and efficient and low-cost building demand response control is achieved.
Patent Information
- Application Number
- CN202510304445.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art is difficult to balance the effectiveness of building demand response control strategies and lightweight implementation needs, especially in terms of computing resources and time, and it is difficult to meet real-time control needs.
Adaptive building demand response lightweight rule control method based on machine learning is adopted to establish building thermal models by collecting operational data, generate training data sets, train machine learning models, and automatically adjust rule-based control parameters in online control.
It effectively alleviates the limitations of traditional model prediction control in computing resources and time, improves the efficiency and effectiveness of building demand response control, reduces computing costs, and has good practical value.
Smart Images

Figure CN120161718A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of building demand response control, and in particular to a lightweight rule control method for building demand response based on machine learning adaptation. Background Art
[0002] With the development of cities, the scale of buildings is getting larger and larger, and the accompanying huge building energy demand. Currently, the direct and indirect energy consumption of buildings accounts for 30% of the global energy consumption, which poses a significant challenge to the carrying capacity of the energy supply side.
[0003] There are various flexible energy storage resources in buildings, such as equipment, walls, and furniture. These resources endow buildings with significant energy storage potential, which can effectively reduce peak loads through the charging and discharging processes of energy and reduce the system operation cost. Formulating good energy management and control strategies can give full play to the flexibility of buildings, reduce the impact of building peak loads on the energy supply side while ensuring the normal and orderly operation of buildings.
[0004] Traditional simple and easy-to-implement rule-based control (RBC) is difficult to adapt to the increasingly complex control requirements. Compared with RBC, model predictive control can adapt to more complex models to the greatest extent, and can also formulate relevant control strategies according to demand response, and can play a role in peak shaving and valley filling for the energy supply side. However, model predictive control is restricted by computing resources and computing time. The huge and complex buildings greatly increase the cost of modeling and calculation, and the computing time and cost of model predictive control completely based on physical simulation models cannot meet the requirements of real-time control.
[0005] Therefore, aiming at the difficulties faced in the construction of building demand response control strategies, exploring a solution that combines optimized control efficiency and low computational complexity is an urgent problem to be solved. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a lightweight rule control method for building demand response based on machine learning adaptation, which can solve the deficiencies of the prior art and take into account the effectiveness of building demand response control strategies and the lightweight implementation requirements.
[0007] To solve the above technical problems, the technical solutions adopted by the present invention are as follows.
[0008] A lightweight rule control method for building demand response based on machine learning adaptation includes the following steps:
[0009] A. Collect operation data and establish a building thermal model;
[0010] B. Generate a training data set containing optimal control actions;
[0011] C. Train a machine learning model for automatically adjusting the parameters of a rule-based controller;
[0012] D. The machine learning model performs online control and automatically adjusts the rule-based control parameters at fixed time intervals.
[0013] Preferably, the building thermal model consists of a building thermal dynamic model and a control optimization model.
[0014] Preferably, the linear time-invariant continuous state space form of the building thermal dynamic model is as follows:
[0015]
[0016] Q cooling = c p m flow (T r - T s ),
[0017] where [T r T m T represents the system state, which are the indoor air temperature and the thermal mass temperature respectively; [T amb I rr O cc T represents the disturbance input, which are the outdoor ambient dry bulb temperature, solar irradiance and occupancy respectively; u represents the cooling power of the control input (Q cooling ); defines the output matrix of the system; Q cooling = c p m flow (T r - T s ) defines the calculation equation of Q cooling , where m flow , T s , c p are the air volume, supply air temperature and specific heat capacity respectively; A, B, D, C are the state matrix, control input matrix, disturbance input matrix and output matrix respectively, and the C matrix is fixed as [1, 0] to physicalize the state of the model.
[0018] Preferably, the control optimization model first solves the optimization problem of the time-of-use electricity price scenario, and the optimization objective J phase1 is as follows:
[0019] J phase1 = Min(C power + Pen u )
[0020] ylb,k ≤y k ≤y ub,k
[0021] u lb,k ≤u k ≤u ub,k ;
[0022] where C power represents the electricity cost; Pen u represents an additional penalty term; k ∈ {0, 1, 2,..., N - 1} represents the set of integers; N represents the prediction horizon length; the subscript k represents the value of the variable at the k-th step; Pr TOU represents the time-of-use electricity price; E represents the system energy consumption, which is a time series matrix composed of k control steps, and each element corresponds to the state of the data set after passing through a control step, and is used to record and represent the evolution process of the data set under a series of control steps; Δu k = u k - u k-1 represents the difference between two adjacent control actions; R represents a diagonal matrix used to constrain the change rate of the control input; y lb,k ≤y k ≤y ub,k and u lb,k ≤u k ≤u ub,k are respectively the lower bound lb and the upper bound ub of the control output and input, that is, the constraints on the indoor temperature and the cooling capacity input;
[0023] The control optimization model takes E as the energy consumption baseline E baseline , and solves the optimization problem of J phase2 .
[0024] J phase2 = M in (C power - C DR + Pen u )
[0025]
[0026] u lb , k ≤ u k ≤ u ub,k ,
[0027] where C DR represents the additional reward obtained during the DR event; Pr DR represents the reward price; represents the additional allowable constraint upper bound during DR, which refers to the temperature threshold that can be tolerated during DR compared to the temperature set value in the normal occupancy mode.
[0028] Preferably, under varying external disturbances and grid signals, after clustering different working conditions, a building thermal model is used to construct a model predictive control problem, and offline closed-loop control is executed in batches to obtain a training data set containing optimal control actions.
[0029] Preferably, parameters related to demand response events adopt a grid-based scheme. For external disturbance inputs, data from a typical meteorological year is used as the input boundary, and several boundaries are selected from it using K-means clustering as the boundaries for offline closed-loop control.
[0030] Preferably, the training process of the machine learning model is to select an input feature set and extract the corresponding predefined rule control parameters as the training objective, and use a machine learning algorithm to perform a regression task for training.
[0031] Preferably, the extraction of rule-based control parameters includes the following steps.
[0032] Take the starting point of precooling in the optimized variable trajectory diagram as the first point where the cooling power of the optimized variable is greater than 0, take the lowest precooling temperature point as the point where the cooling power reaches the maximum value, and take the end point of precooling as the point where the cooling power is equal to 0 again.
[0033] Preferably, the machine learning model is deployed in the supervisory control layer to automatically adjust rule-based control parameters according to meteorological conditions, demand response signals, and the upper limit of indoor temperature during demand response events; the RBC controller and the PID controller form the local control layer. The RBC controller is used to receive the output parameter values of the machine learning model and generate a setpoint sequence using an interpolation method, and the PID controller is used to track the setpoint and adjust and control the HVAC system equipment.
[0034] Preferably, the RBC controller has four modes, namely unoccupied precooling, occupied by people, demand response period, and unoccupied by people. At the beginning of each day, it first enters the unoccupied by people mode and the system remains closed; after the system is awakened, it enters the unoccupied precooling mode, and then the temperature will drop until the lowest point and then start to climb. After the precooling ends, the temperature setpoint is adjusted to the same temperature as during the occupied period; since the temperature cannot reach the setpoint, the valve opening of the variable air volume terminal device is reduced, and at the same time, the system power is reduced; then it enters the occupied mode to maintain the thermal comfort of indoor personnel; during this period, it switches to the demand response mode according to the grid day-ahead demand response signal to reduce the system power.
[0035] The beneficial effects brought by the above technical solutions are as follows: The present invention uses a machine learning agent to learn from a dataset generated by an offline MPC closed-loop simulation and uses it to assist in adjusting the parameters of a rule-based controller (RBC) during the online phase. This effectively alleviates the problem that traditional MPC is difficult to be directly deployed in existing building systems. Since most existing building control systems are equipped with RBC, the control strategy (ML-RBC) of the present invention can be better adapted. It effectively solves the problem that it is difficult for the prior art to balance the effectiveness of the building demand response control strategy and the lightweight implementation requirements. Compared with the traditional baseline strategy and RBC strategy, it can save a large amount of costs; the online control of traditional MPC is relatively complex. The present invention approximates the control ability of MPC by learning the behavior of MPC, which can reduce the burden of online control while ensuring control performance and can save a large amount of computing costs, and has good practical value. The present invention has low computing costs and can be deployed in a building automation system with limited performance; the low computing cost is reflected in that the proposed method does not require optimization, only needs to run the machine learning model and execute the interpolation program at the beginning of each day. The RBC with adjustable parameters can realize pre-cooling the building at night to store energy and raising the temperature upper limit during DR events to cut system power. Machine learning is used for cross-validation and hyperparameter tuning to ensure that the machine learning model has accuracy and generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a schematic flow chart of the present invention.
[0037] Figure 2 is the relationship and control frequency between the control architecture implemented in an actual multi-zone building and the controllers used in the present invention.
[0038] Figure 3 is a schematic diagram of the RBC with adjustable parameters used in the present invention.
[0039] Figure 4 is a schematic flow chart of the offline model predictive control batch closed-loop simulation used in the present invention to obtain a dataset of optimal control sequences.
[0040] Figure 5 is a schematic diagram of the trajectory of the optimized variable in the optimal control sequence used in the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0041] Referring to Figure 1 , the present invention includes the following steps:
[0042] A. Collect operation data and establish a building thermal model; the building thermal model consists of a building thermal dynamic model and a control optimization model.
[0043] The linear time-invariant continuous state space form of the building thermal dynamic model is as follows:
[0044]
[0045]
[0046] Q coling = cpm flow (T r - T s ),
[0047] where [T r T m T represents the system states, namely the indoor air temperature and the thermal mass temperature; [T amb Irr Occ] T represents the disturbance inputs, namely the outdoor ambient dry-bulb temperature, solar irradiance, and occupancy; u represents the cooling power of the control input (Q cooling ); defines the output matrix of the system; Q cooling = c p m flow (T r - T s ) defines the calculation equation of Q cooling , where m flow , T s , c p are the air volume flow rate, supply air temperature, and specific heat capacity respectively; A, B, D, C are the state matrix, control input matrix, disturbance input matrix, and output matrix respectively. The C matrix is fixed as [1, 0] to physicalize the states of the model.
[0048] The control optimization model first solves the optimization problem for the time-of-use electricity price scenario, and the optimization objective J phase1 is as follows,
[0049] J phase1 = Min(C power + Pen u )
[0050]
[0051] y lb,k ≤ y k ≤ y ub,k
[0052] u lb,k ≤ u k ≤ u ub,k ;
[0053] where C power represents the electricity cost; Pen u represents an additional penalty term; k ∈ {0, 1, 2, ..., N - 1} represents the set of integers; N represents the prediction horizon length; the subscript k represents the value of the variable at the k-th step; Pr TOU represents the time-of-use electricity price; E represents the system energy consumption, which is a time series matrix composed of k control steps, and each element corresponds to the state of the dataset after a control step, used to record and represent the evolution process of the dataset under a series of control steps; Δu k = u k - u k-1 represents the difference between two adjacent control actions; R represents a diagonal matrix used to constrain the change rate of the control input; y lb,k ≤ y k ≤ y ub,k and u lb,k ≤ u k ≤ u ub,k are respectively the lower limit lb and the upper limit ub of the control output and input, that is, the constraints on the indoor temperature and the cooling capacity input;
[0054] The control optimization model takes E as the energy consumption baseline E baseline , and solves the optimization problem of J phase2 ,
[0055] J phasse2 = Min(C power - C DR + Pen u )
[0056]
[0057] u lb,k ≤ u k ≤ u ub,k ,
[0058] where C DR represents the additional reward obtained during the DR event; Pr DR represents the reward price; represents the additional allowable constraint upper limit during DR, which refers to the temperature threshold that can be tolerated during DR compared to the temperature set value in the normal occupancy mode.
[0059] B. Under changing external disturbances and grid signals, after clustering different working conditions, use the building thermal model to construct a model predictive control problem and execute batch offline closed-loop control to obtain a training dataset containing optimal control actions.
[0060] C. Select the input feature set and extract the corresponding pre-defined rule control parameters as the training target, and use machine learning algorithms to perform regression tasks to train a machine learning model for automatically adjusting the rule controller parameters.
[0061] D. The machine learning model performs online control and automatically adjusts the rule-based control parameters at fixed time intervals.
[0062] Refer to Figure 2 , the control architecture of the present invention is divided into a supervisory control layer and a local control layer;
[0063] The supervisory control layer deploys a machine learning model, which is mainly used to automatically adjust the predefined parameters of the RBC according to meteorological conditions, DR signals, and the upper limit of indoor temperature during DR events.
[0064] The local control layer includes an RBC and a PID controller. The RBC is used to receive the output parameter values of the machine learning model and generate a setpoint sequence using an interpolation method. The PID controller focuses on tracking the above set values and adjusts and controls the HVAC system equipment.
[0065] The architecture processes multiple hot zones by adopting independent temperature setpoint control for each zone and equipping with separate ML-RBCs.
[0066] In this embodiment, the temperature setpoint control is completed by adjusting the valve opening of the variable air volume terminal unit (VAVBox), which does not represent the only way.
[0067] The overall computational cost of the architecture is low, and it can be deployed in building automation systems with limited performance. The proposed method does not require optimization, and only needs to run the machine learning model and execute the interpolation program at the beginning of each day.
[0068] Refer to Figure 3 , the RBC that adjusts parameters makes up for the disadvantage of the traditional RBC that cannot make adaptive adjustments according to weather and dynamic characteristics of hot zones;
[0069] The RBC can switch between 4 different modes, namely unoccupied precooling, occupied by people, demand response period, and unoccupied by people; at the beginning of each day, it first enters the unoccupied by people mode and the system remains closed; at point A, the system is awakened and enters the unoccupied precooling mode, after which the temperature will slowly drop until the lowest point B and then start to climb, and the precooling ends at point C and the temperature set value is adjusted to the same temperature as during the occupied period; at this time, there is still some time until the occupied mode, and since the temperature cannot reach the set value, the local control will reduce the VAVBox valve opening and at the same time reduce the system power; then it enters the occupied mode to maintain the thermal comfort of indoor personnel. During this period, it will switch to the demand response mode (point D) according to the grid day-ahead DR signal to reduce the system power.
[0070] Figure 3For the continuous sequence of temperature set points shown, an interpolation method needs to be used between parameter points to approximate the optimal control trajectory. In this embodiment, the transition curve between control points (A, B, C) constructed using the Hermite interpolation method does not represent the only interpolation method.
[0071] Referring to Figure 4 , the process of obtaining the optimal control sequence dataset includes:
[0072] (1) Generate boundaries for simulation, including meteorological boundaries and input parameters related to day-ahead DR;
[0073] (2) Perform offline simulation.
[0074] The boundaries should cover all potential scenarios as much as possible to ensure that the machine learning model can obtain sufficient information from the data, so as to be able to handle various possible situations in the future. In this embodiment, a mixed input boundary construction method is used, which does not represent the only input boundary construction method.
[0075] Specifically, the parameters related to DR events adopt a grid scheme. For external disturbance inputs, the data of a typical meteorological year is used as the input boundary, and K-means clustering is used to select representative boundaries from them for closed-loop simulation.
[0076] Figure 5 It is a schematic diagram of the optimized variable trajectory in the optimal control sequence, and RBC parameters are extracted through the variable trajectory.
[0077] There are a pre-cooling start point A, a lowest pre-cooling temperature point B, and a pre-cooling end point C in the trajectory. The essence of the extraction of output features is to find the above three points from a series of generated control trajectories; the pre-cooling start point A is the first point where the cooling power of the optimized variable is greater than 0; the lowest pre-cooling temperature point B is the point where the cooling power reaches the maximum value; the pre-cooling end point C is the point where the cooling power is equal to 0 again.
[0078] In the embodiment of the present invention, the XGBoost model is used to perform the feature selection task, evaluate different feature subsets, and find a point where the performance no longer improves significantly to determine the appropriate feature input and model complexity.
[0079] In the embodiment of the present invention, the automatic machine learning toolbox scikit-optimize is used for cross-validation and hyperparameter tuning to ensure that the XGBoost model has accuracy and generalization ability; the cross-validation is implemented using the 10-fold method; the hyperparameter tuning is implemented using the Bayesian optimization algorithm.
[0080] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on the present invention.
[0081] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and all these changes and improvements fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. A lightweight rule control method for building demand response based on machine learning and self-adaptation, characterized in that The following steps are involved: A. Collect operating data and establish a building thermal model; B. Generate a training data set containing optimal control actions; C. Train a machine learning model for automatically adjusting the parameters of the rule controller; D. The machine learning model performs online control and automatically adjusts the rule-based control parameters at a fixed time.
2. The building demand response lightweight rule control method based on machine learning adaptation according to claim 1 is characterized by: The building thermal model consists of a building thermal dynamic model and a control optimization model.
3. The building demand response lightweight rule control method based on machine learning adaptation according to claim 2 is characterized by: The linear time-invariant continuous state space form of the building thermal dynamic model is as follows: Q cooling =c p m flow (T r -T s ), Where [T T T m ] T represents the system state, which are the indoor air temperature and thermal mass temperature [T amb I rr O cc ] T represents the interference input, which are the outdoor environment dry bulb temperature, solar irradiance and occupancy rate; u represents the cooling power of the control input (Q cooling ); Defines the output matrix of the system; Q cooling =c p m flow (T r -T s ) defines Q cooling The calculation equation is: flow , T s , c p are air volume, supply air temperature and specific heat capacity respectively; A, B, D, C are state matrix, control input matrix, interference input matrix and output matrix respectively, and the C matrix is fixed to [1,0] to make the state of the model physical.
4. The building demand response lightweight rule control method based on machine learning adaptation according to claim 3 is characterized by: The control optimization model first solves the optimization problem of the time-of-use electricity price scenario, and the optimization target J phase1 As shown below, J phase1 =Min(C power +Pen u ) and lb,k ≤y k ≤y ub,k in lb,k in k in ub,k ; Among them C power Pen represents the cost of electricity; u represents an additional penalty term; k∈{0, 1, 2, ..., N-1} represents an integer set; N represents the length of the prediction time domain; the subscript k represents the value of the variable at the kth step; Pr TOU Represents the time-of-use electricity price; E represents the system energy consumption, which is a time series matrix composed of k control steps. Each element corresponds to the state of the data set after one control step, which is used to record and represent the evolution process of the data set under a series of control steps; Δu k =u k -u k-1 represents the difference between two adjacent control actions; R represents a diagonal matrix, which is used to constrain the rate of change of the control input; y lb,k ≤y k ≤y ub,k and u lb,k ≤u k ≤u ub,k They refer to the lower limit lb and upper limit ub of the control output and input, i.e. the constraints of the indoor temperature and cooling capacity input respectively; The control optimization model takes E as the energy baseline E baseline , solve for J phase2 The optimization problem J phase2 =Min(C power -C DR +Pen u ) in lb,k in k in ub,k , Among them C DR Represents the additional rewards obtained during the DR event; Pr DR represents the reward price; Represents the additional upper limit of the constraint allowed during DR, which refers to the temperature threshold that can be tolerated during DR compared to the temperature setting value of the normal occupancy mode.
5. The building demand response lightweight rule control method based on machine learning adaptation according to claim 1 is characterized by: Under changing external disturbances and grid signals, after clustering analysis of different operating conditions, the building thermal model is used to construct the model predictive control problem, and offline closed-loop control is executed in batches to obtain a training data set containing the optimal control actions.
6. The building demand response lightweight rule control method based on machine learning adaptation according to claim 5 is characterized by: The parameters related to demand response events are gridded. For external disturbance input, the data of a typical meteorological year are used as the input boundary, and K-means clustering is used to select several boundaries as the boundaries for offline closed-loop control.
7. The building demand response lightweight rule control method based on machine learning adaptation according to claim 1 is characterized by: The training process of the machine learning model is to select the input feature set and extract the corresponding predefined rule control parameters as the training target, and use the machine learning algorithm to perform the regression task for training.
8. The building demand response lightweight rule control method based on machine learning adaptation according to claim 1 is characterized by: Extracting rule-based control parameters includes the following steps: The precooling starting point in the optimization variable trajectory diagram is taken as the first point where the optimization variable cooling power is greater than 0, the lowest precooling temperature point is taken as the point where the cooling power reaches the maximum value, and the precooling end point is taken as the point where the cooling power is equal to 0 again.
9. The building demand response lightweight rule control method based on machine learning adaptation according to claim 1 is characterized by: The machine learning model is deployed in the supervisory control layer to automatically adjust the rule-based control parameters according to meteorological conditions, demand response signals, and the upper limit of indoor temperature during demand response events. The RBC controller and PID controller constitute the local control layer. The RBC controller is used to accept the output parameter values of the machine learning model and generate a set point sequence using the interpolation method. The PID controller is used to track the set values and adjust the HVAC system equipment.
10. The building demand response lightweight rule control method based on machine learning adaptation according to claim 9 is characterized by: The RBC controller has four modes, namely unoccupied pre-cooling, occupied, demand response period and unoccupied. At the beginning of each day, it first enters the unoccupied mode and the system remains off; after the system is awakened, it enters the unoccupied pre-cooling mode, after which the temperature will drop to the lowest point and then begin to climb. After the pre-cooling ends, the temperature set value is adjusted to the same temperature as during the occupied period; since the temperature cannot reach the set value, the valve opening of the variable air volume terminal device is reduced, and the system power is reduced at the same time; then it enters the occupied mode to maintain the thermal comfort of indoor occupants; during this period, it switches to the demand response mode to reduce the system power according to the demand response signal of the power grid on the previous day.
Citation Information
Cited By
Realization method of demand response strategy of office building refrigeration system based on time-of-use electricity price
CN122083466A