Reinforcement learning method and device for balancing personalized thermal comfort and hvac energy consumption
By constructing a mechanism-based energy consumption model and a user-personalized comfort model for HVAC systems, and combining them with the Q-learning algorithm, the decision-making process of HVAC systems is optimized. This solves the problems of dynamic changes in user-personalized comfort and the effects of multivariate coupling, and achieves the goal of meeting personalized comfort needs while optimizing energy consumption.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TSINGHUA UNIVERSITY
- Filing Date
- 2023-11-22
- Publication Date
- 2026-05-19
AI Technical Summary
In existing technologies, HVAC systems struggle to effectively address the dynamic changes in user-specific comfort and the impact of multivariate coupling during design. This results in HVAC systems failing to meet personalized comfort requirements when optimizing energy consumption.
By employing reinforcement learning, we construct a mechanism-based energy consumption model for HVAC systems, a room heat transfer mechanism model based on the heat balance method, and a user-personalized comfort model based on the PMV index. Combined with the Q-learning algorithm, we differentiate user heating and cooling needs through metabolic rate, train the decision-making library, calculate comfort index and energy consumption weight, and optimize the decision-making process of the HVAC system.
It achieves the optimization of HVAC energy consumption while dynamically adjusting indoor temperature to meet personalized comfort needs, reduces computational load, and takes into account users' energy-saving potential, sensitivity, and consumption habits, thereby improving the personalized comfort and energy efficiency of the HVAC system.
Smart Images

Figure CN117606133B_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the field of energy-saving optimization in intelligent buildings, and in particular to a reinforcement learning method and apparatus for balancing personalized thermal comfort and HVAC energy consumption. Background Technology
[0002] With global warming, scientifically reducing equipment energy consumption and effectively implementing carbon reduction measures are imperative. Buildings, as typical demand-side energy-consuming systems, account for approximately 40% of total social energy consumption; and of this, about 40% is consumed by HVAC systems (Heating, Ventilation, and Air Conditioning systems). Therefore, researching energy-saving optimization of HVAC systems has significant practical implications. At the same time, as people's living standards improve, their demands for indoor comfort are also increasing. Clearly, higher comfort requirements mean higher energy consumption for HVAC systems, creating a trade-off between the two. Therefore, how to ensure that HVAC systems operate energy-efficiently while maximizing the personalized thermal comfort of indoor occupants has become a pressing issue.
[0003] Existing literature on HVAC energy consumption optimization typically involves first establishing a model to form an optimization problem, and then seeking the optimal solution.
[0004] In HVAC system modeling, there are three main approaches: physics-based, software simulation, and data-driven. First, physics-based methods typically use physical formulas to describe the heat transfer process within a building, calculating energy consumption by analyzing equipment mechanisms. Second, software simulation methods often utilize commercial software such as EnergyPlus and TRNSYS to simulate the dynamic thermal processes of a building, generating large amounts of data, but consuming significant time and computational resources. Finally, data-driven methods typically employ artificial intelligence methods such as machine learning to build "black box" models based on full data (e.g., artificial neural networks) and "gray box" models based on semi-data (e.g., parameter identification models).
[0005] Regarding optimization methods, existing literature employs many classic solution techniques, such as mixed-integer programming and model predictive control. Many heuristic methods, such as genetic algorithms and particle swarm optimization, are also used, as they can search for approximate optimal solutions without requiring gradient information. Furthermore, some literature utilizes reinforcement learning algorithms, which solve problems through interaction with the environment, exhibiting greater adaptability to dynamic changes.
[0006] To balance thermal comfort and HVAC system energy consumption, thermal comfort is often considered a constraint; as long as temperature and humidity are within a set range, the user is considered comfortable. For example, existing technologies assume a user comfort range of 22°C-26°C and humidity of 40%-60%. However, user comfort is influenced by numerous factors. This means that setting separate optimization ranges for different comfort variables fails to consider the coupled effects of these factors, resulting in a one-sided characterization of user comfort. Furthermore, different users perceive heat differently in the same environment due to individual differences, and their consumption habits also influence the trade-off between thermal comfort and system energy consumption. These personalized factors are not considered in the optimization processes of existing literature.
[0007] Therefore, how to incorporate a comfort measurement mechanism that can reflect both the dynamic changes in user comfort and the combined effects of multiple variables and individual differences among users into HVAC system operation decisions has become a problem that requires further research. Summary of the Invention
[0008] To address the problem that current HVAC energy consumption optimization methods do not consider user-specific comfort levels, resulting in a one-sided portrayal of user comfort and an inability to meet the dynamic changes in user comfort levels, this specification provides a reinforcement learning method and apparatus for balancing personalized thermal comfort and HVAC energy consumption. The implementation steps of this technical method are as follows: (1) Constructing a mechanism-based HVAC system energy consumption model; (2) Constructing a room heat transfer mechanism model based on the heat balance method; (3) Constructing a user-specific comfort model based on the PMV index; (4) Constructing a framework for solving the optimization problem of balancing personalized comfort and energy consumption and a reinforcement learning algorithm based on Q-learning.
[0009] To solve any of the above-mentioned technical problems, the specific technical solutions of the embodiments in this specification are as follows:
[0010] On the one hand, embodiments of this specification provide a reinforcement learning decision-making method for balancing personalized thermal comfort and HVAC energy consumption, including:
[0011] The state variables, consisting of the difference between the indoor temperature, the wall temperature and the indoor temperature, and the time variable, and the action variables, consisting of the difference between the outlet air temperature of the FCU fan in the HVAC system and the indoor temperature, and the air volume of the FCU fan, are used as input data.
[0012] The input data is trained using a reinforcement learning algorithm with predetermined parameters for multiple predetermined metabolic rates to obtain a decision base corresponding to each metabolic rate. The decision base includes decision values corresponding to each pair of multiple state variables and multiple action variables. The decision values are calculated based on the reward function value and the state variables and action variables corresponding to multiple decision stages. The state variables corresponding to the current decision stage are obtained by iteratively calculating the state variables and action variables of the previous decision stage using a pre-built room heat transfer mechanism model based on the heat balance method. The action variables are selected using an ε-greedy strategy. The reward function value is calculated using the comfort index calculated from the state variables, the energy consumption of the HVAC system calculated from the state variables and action variables, and weighting coefficients calculated based on user energy-saving potential, user sensitivity, and user consumption habits.
[0013] When making decisions regarding the airflow rate and outlet temperature of the FCU fan in the target room where the target user is located, a target decision library is determined based on the target user's metabolic rate. The target decision library is then used to make decisions based on the actual air temperature and wall temperature in the target room, resulting in target airflow rate and target outlet temperature for the FCU fan at multiple decision stages. This allows for the control of the FCU fan in the target room using the target airflow rate and target outlet temperature at the time corresponding to each decision stage, thereby ensuring that the indoor temperature in the target room meets the comfort level of the target user.
[0014] Furthermore, the formula for calculating the comfort index using the aforementioned state variables is as follows:
[0015] PMV=[0.0303exp(-0.036M)+0.028]L;
[0016] L = M - 3.05 × 10 -3 ×[5733-6.99MP a -0.42*[M-58.15]-1.7
[0017] ×10 -5 M(5867-P a -0.0014M(34-T) a )-F cl *h c *(T cl -T a )
[0018] -3.96×10 -8 F cl [(T cl +273)4 -(T r +273) 4 ];
[0019] T cl =35.7-0.028MI cl F cl 3.96×10 -8 [(T cl +273) 4 -(T r +273) 4 ]+
[0020] h c (T cl -T a )};
[0021]
[0022]
[0023] P a =6.1094h r ×exp[17.625T a / (T a +243.04)];
[0024] Where PMV represents the comfort index, which is the calculated PMV value; M represents the metabolic rate, measured in W / m³. 2 L represents the human body heat load, measured in W / m². 2 ;P a T is the vapor pressure of water, measured in Pa; a It is the indoor air temperature, T cl It refers to the surface temperature of the clothing, T. r This is the mean radiant temperature, and the unit is °C; F cl It is the surface area coefficient of clothing; I cl It is the thermal insulation coefficient of the clothing surface, and the unit is m. 2 K / W; h c It is the convective heat transfer coefficient, with units of W / m². 2 K;V a It is the air velocity, measured in m / s; h r Relative humidity is expressed as a percentage.
[0025] Furthermore, the formula for calculating the weighting coefficients based on user energy-saving potential, user sensitivity, and user consumption habits is as follows:
[0026] λ=(m*β-n*α)*γ;
[0027] Where λ represents the weighting coefficient; m and n are adjustable parameters; α represents the user's energy-saving potential, defined as the energy reduction rate when PMV = 0.5 is relaxed to 0.1 as a strict constraint; β represents the user's sensitivity, defined as the average change rate of PMV value for every 1°C change in temperature within the adjustable temperature range; γ represents the user's consumption habits, defined as a number between 0 and 1, the larger the value, the more willing the user is to pay more costs to improve their comfort experience.
[0028] Furthermore, the formula for calculating the reward function value is as follows:
[0029] reward=-(Energy_cost+λ*Comfort_cost);
[0030] Where reward represents the reward function value, Energy_cost represents the energy consumption of the HVAC system, and λ represents the weighting coefficient;
[0031]
[0032] Comfort_cost represents the comfort cost calculated based on the comfort index PMV.
[0033] Furthermore, by using a pre-constructed room heat transfer mechanism model based on the heat balance method to iteratively calculate the state variables and action variables of the previous decision-making stage, the formula for the state variables of the current decision-making stage is obtained as follows:
[0034]
[0035] Where, m a It refers to the air quality inside the room. Q represents the indoor temperature at the (k+1)th decision stage, Δt represents the time length corresponding to a single decision cycle, and Q is the temperature at the (k+1)th decision stage. g It is the heat production rate per person. and h represents the heat generated by light and indoor equipment at the k-th decision stage, respectively. gs It is the heat transfer coefficient of indoor and outdoor air through the window, A gs It is the area of the window. The outdoor temperature at the k-th decision stage. The indoor air temperature, h, is at the k-th decision stage. w It is the heat transfer coefficient between the wall and the indoor air, A w It is the wall area. The wall temperature, C, is at the k-th decision stage. p It is the specific heat of air. This represents the air output of the FCU fan in the k-th decision stage. It is the outlet air temperature of the FCU fan in the k-th decision stage. It is the energy contained in the remaining air obtained after removing the air pumped out by the FCU in the room;
[0036]
[0037] Among them, C w It is the heat capacity of the wall, m w It's about the quality of the wall. Let S represent the difference between the wall temperature at the (k+1)th decision stage and the wall temperature at the kth decision stage. w This represents the amount of heat absorbed by indoor air and walls from solar radiation.
[0038] Furthermore, the reinforcement learning algorithm is a Q-learning algorithm, and the decision value is the Q-value;
[0039] The formula for calculating the decision value based on the reward function value and the state and action variables corresponding to multiple decision stages is as follows:
[0040] Q M (x,u) new =Q M (x,u) old *[1-θ]+θ[reward+γ*
[0041] max u_next Q M (x next ,u next )];
[0042] Among them, Q M (x,u) new This represents the Q-value calculated by the Q-learning algorithm in the current iteration when the metabolic rate is M, based on the state variable x and the action variable u. M (x,u) old When the metabolic rate is M, Q represents the Q-value calculated by the state variable x and action variable u in the previous iteration of the Q-learning algorithm; θ represents the learning rate in the predetermined parameters; reward represents the reward function value; γ represents the discount coefficient in the predetermined parameters; and max represents the maximum value. u_next Q M (x next ,u next This indicates that the state variable x for the next decision stage is obtained by iteratively calculating the state variables and action variables of the current decision stage using a pre-built room heat transfer mechanism model based on the heat balance method. next And based on the state variable x next Iterate through all action variables u next For the corresponding Q value, select the maximum value among the Q values.
[0043] Furthermore, the formula for the state variable x is:
[0044]
[0045] Where, x k This represents the state variable at the k-th decision stage. The indoor air temperature at the k-th decision stage. The wall temperature is at the k-th decision stage, and t represents the current decision time variable.
[0046] The formula for the action variable u is:
[0047]
[0048] Among them, u k This represents the action variable in the k-th decision stage. This represents the outlet air temperature of the FCU fan in the k-th decision stage. Indoor air temperature at the k-th decision stage The difference between them This represents the air output of the FCU fan in the k-th decision stage.
[0049] Furthermore, the step of determining the target decision base based on the target user's metabolic rate includes:
[0050] Using formula Calculate the metabolic rate of the target user, where M represents the metabolic rate of the target user, N represents the total number of votes cast by the user, and Vote represents the actual vote value cast by the user. i This represents the environmental parameters in which the user is voting at the time of the i-th vote, including the indoor air temperature T. a Mean radiation temperature T r Relative humidity h r air velocity V a Clothing coefficient I cl PMV(M, parameter) i () indicates that the metabolic rate is M and the environmental parameter is parameter. i The PMV value calculated at that time;
[0051] Determine the decision library corresponding to the metabolic rate of the target user to obtain the target decision library;
[0052] The steps for making decisions based on the actual air temperature and actual wall temperature inside the target room using the target decision database, and obtaining the target air volume and target air outlet temperature of the FCU fan corresponding to multiple decision stages, include:
[0053] Construct the current state variables based on the actual air temperature, the actual wall temperature, and the current decision-making stage.
[0054] In the target decision base, determine the maximum decision value corresponding to the current state variable, and take the action variable corresponding to the maximum decision value as the optimal action;
[0055] Based on the current state variables and the optimal action, the next state variables corresponding to the next decision stage are calculated using the room heat transfer mechanism model based on the heat balance method.
[0056] The next state variable is used as the current state variable, and the step of determining the current optimal action corresponding to the current state variable in the target decision base is repeated until the optimal action corresponding to all decision stages is determined. The optimal action includes the target air volume of the FCU fan and the target air temperature of the FCU fan.
[0057] On the other hand, embodiments of this specification also provide a reinforcement learning device for balancing personalized thermal comfort and HVAC energy consumption, the device comprising:
[0058] The input data construction unit is used to take the state variables consisting of the difference between the indoor temperature, the wall temperature and the indoor temperature, and the time variable, as well as the action variables consisting of the difference between the outlet air temperature of the FCU fan in the HVAC system and the indoor temperature, and the air volume of the FCU fan as input data.
[0059] The decision base training unit is used to train the input data using a reinforcement learning algorithm with predetermined parameters for multiple predetermined metabolic rates to obtain a decision base corresponding to each metabolic rate. The decision base includes decision values corresponding to each pair of multiple state variables and multiple action variables. The decision values are calculated based on the reward function value and the state variables and action variables corresponding to multiple decision stages. The state variables corresponding to the current decision stage are obtained by iteratively calculating the state variables and action variables of the previous decision stage using a pre-built room heat transfer mechanism model based on the heat balance method. The action variables are selected through an ε-greedy strategy. The reward function value is calculated using the comfort index calculated from the state variables, the energy consumption of the HVAC system calculated from the state variables and action variables, and weight coefficients calculated based on user energy-saving potential, user sensitivity, and user consumption habits.
[0060] The decision-making unit is used to determine a target decision library based on the target user's metabolic rate when making decisions about the air volume and outlet temperature of the FCU fan in the target room where the target user is located. It then uses the target decision library to make decisions based on the actual air temperature and actual wall temperature in the target room, obtaining target air volume and target outlet temperature of the FCU fan for multiple decision stages. This allows for the control of the FCU fan in the target room using the target air volume and target outlet temperature at the time corresponding to each decision stage, thereby ensuring that the indoor temperature in the target room meets the comfort level of the target user.
[0061] On the other hand, embodiments of this specification also provide a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the above-described method.
[0062] This specification's embodiments abandon the traditional method of using a uniform comfort range to make decisions about the operating conditions of HVAC systems. Instead, these embodiments use metabolic rate to differentiate users with different heating and cooling needs, training a decision base corresponding to each metabolic rate. When processing training data, temperature difference is introduced into the state variables to reduce the aggregation of state and action spaces, significantly reducing computational load. Furthermore, time variables are introduced into the state variables to eliminate the time dependency of decision values. In addition, during the calculation of decision values in the decision base, comfort indices are calculated using state variables instead of the fixed comfort temperature range in traditional methods. This reflects both the dynamic changes in user comfort and the impact of multi-variable coupling. A personalized method for determining the reward function weight parameters is proposed, calculating weight coefficients for user energy-saving potential, user sensitivity, and user consumption habits. This fully considers users' energy-saving potential, sensitivity, and consumption habits, thereby constructing decision bases with different energy-saving potentials, heating and cooling preferences, and consumption habits. Finally, when making decisions about the operating conditions of the HVAC system in the target room where the target user is located, a target decision library is determined based on the user's metabolic rate. Then, the target decision library is used to make decisions based on the actual air temperature and actual wall temperature in the target room to obtain the HVAC system operating conditions at each decision stage. This allows the FCU fan in the target room to be controlled using the HVAC system operating conditions at each decision stage at the corresponding time in that decision stage, thereby ensuring that the indoor temperature in the target room meets the comfort level of the target user. Attached Figure Description
[0063] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0064] Figure 1 The diagram shown is a flowchart illustrating a reinforcement learning method for balancing personalized thermal comfort and HVAC energy consumption in an embodiment of this specification.
[0065] Figure 2 The diagram shown is a typical structural diagram of the HVAC system in the embodiments of this specification;
[0066] Figure 3 The diagram shown is a flowchart illustrating the process of making decisions using the target decision library based on the actual air temperature and actual wall temperature in the target room, as described in this specification embodiment.
[0067] Figure 4 The diagram shown is a schematic representation of the computational flow of a reinforcement learning method that balances personalized thermal comfort and HVAC energy consumption in an embodiment of this specification.
[0068] Figure 5 The diagram shown is a structural schematic of a reinforcement learning device that balances personalized thermal comfort and HVAC energy consumption in an embodiment of this specification.
[0069] Figure 6 The diagram shown is a structural schematic of the computer device in an embodiment of this specification.
[0070] [Explanation of Figure Markers]:
[0071] 501. Input data construction unit; 502. Decision base training unit; 503. Decision unit; 602. Computer equipment; 604. Processor; 606. Memory; 608. Drive mechanism; 610. Input / output module; 612. Input device; 614. Output device; 616. Presentation device; 618. Graphical user interface; 620. Network interface; 622. Communication link; 624. Communication bus. Detailed Implementation
[0072] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the embodiments of this specification, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the embodiments of this specification.
[0073] It should be noted that the terms "first," "second," etc., in the description, claims, and accompanying drawings of the embodiments herein are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, apparatus, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0074] It should be noted that the acquisition, storage, use, and processing of data in the technical solution of this application all comply with the relevant provisions of relevant laws and regulations.
[0075] This specification provides a reinforcement learning method for balancing personalized thermal comfort and HVAC energy consumption. The implementation steps of this method are as follows: (1) constructing a mechanism-based HVAC system energy consumption model; (2) constructing a room heat transfer mechanism model based on the heat balance method; (3) constructing a user-personalized comfort model based on the PMV index; (4) constructing a framework for solving the personalized comfort and energy consumption trade-off optimization problem and a reinforcement learning algorithm based on Q-learning. Figure 1 The diagram illustrates a learning-based decision-making method for balancing personalized thermal comfort and HVAC energy consumption, as described in an embodiment of this specification. The diagram depicts the process of constructing a decision base and using it to make decisions. The order of steps listed in the embodiment is merely one possible execution order among many and does not represent the only possible order. In actual system or device products, the methods shown in the embodiment or the accompanying drawings can be executed sequentially or in parallel. Specifically, as shown... Figure 1 As shown, the method can be executed by a processor in the main controller of an indoor HVAC system and may include:
[0076] Step 101: Use the state variables consisting of the difference between the indoor temperature, the wall temperature and the indoor temperature, and the time variable, and the action variables consisting of the difference between the outlet air temperature of the FCU fan in the HVAC system and the indoor temperature, and the air volume of the FCU fan as input data.
[0077] Step 102: Use a reinforcement learning algorithm with predetermined parameters to train the input data for multiple predetermined metabolic rates to obtain a decision base corresponding to each metabolic rate;
[0078] In this step, the decision base includes multiple state variables and multiple action variables corresponding to each other, and the decision values are calculated based on the reward function value and the state variables and action variables corresponding to multiple decision stages;
[0079] The state variables corresponding to the current decision-making stage are obtained by iteratively calculating the state variables and action variables of the previous decision-making stage using a pre-built room heat transfer mechanism model based on the heat balance method. The action variables are selected through the ε-greedy strategy.
[0080] The reward function value is calculated using the comfort index calculated from the state variables, the energy consumption of the HVAC system calculated from the state variables and action variables, and the weighting coefficients calculated based on the user's energy-saving potential, user sensitivity, and user consumption habits.
[0081] Step 103: When making decisions about the air volume and air temperature of the FCU fan in the target room where the target user is located, a target decision library is determined based on the target user's metabolic rate. The target decision library is used to make decisions based on the actual air temperature and actual wall temperature in the target room to obtain the target air volume and target air temperature of the FCU fan corresponding to multiple decision stages.
[0082] After obtaining the target air volume and target air temperature of the FCU fan corresponding to multiple decision stages, the FCU fan in the target room can be controlled using the target air volume and target air temperature at the time corresponding to each decision stage, so that the indoor temperature in the target room meets the comfort level of the target user.
[0083] This specification's embodiments abandon the traditional method of using a uniform comfort range to make decisions about the operating conditions of HVAC systems. Instead, these embodiments use metabolic rate to differentiate users with different heating and cooling needs, training a decision base corresponding to each metabolic rate. When processing training data, temperature difference is introduced into the state variables to reduce the aggregation of state and action spaces, significantly reducing computational load. Furthermore, time variables are introduced into the state variables to eliminate the time dependency of decision values. In addition, during the calculation of decision values in the decision base, comfort indices are calculated using state variables instead of the fixed comfort temperature range in traditional methods. This reflects both the dynamic changes in user comfort and the impact of multi-variable coupling. A personalized method for determining the reward function weight parameters is proposed, calculating weight coefficients for user energy-saving potential, user sensitivity, and user consumption habits. This fully considers users' energy-saving potential, sensitivity, and consumption habits, thereby constructing decision bases with different energy-saving potentials, heating and cooling preferences, and consumption habits. Finally, when making decisions about the operating conditions of the HVAC system in the target room where the target user is located, a target decision library is determined based on the user's metabolic rate. Then, the target decision library is used to make decisions based on the actual air temperature and actual wall temperature in the target room to obtain the HVAC system operating conditions at each decision stage. This allows the FCU fan in the target room to be controlled using the HVAC system operating conditions at each decision stage at the corresponding time in that decision stage, thereby ensuring that the indoor temperature in the target room meets the comfort level of the target user.
[0084] This specification outlines several implementations, including a mechanism-based HVAC system energy consumption model, a room heat transfer mechanism model based on the heat balance method, a user-personalized comfort model based on the PMV index, and a framework for constructing an optimization problem that balances personalized comfort with energy consumption, along with a reinforcement learning algorithm. Among these:
[0085] (1) The mechanism-based HVAC system energy consumption model is used to calculate the energy consumption of the HVAC system based on indoor temperature, air outlet temperature of FCU fan in HVAC system, air outlet volume of FCU fan, etc.
[0086] (2) The room heat transfer mechanism model based on the heat balance method is used to calculate the indoor temperature, wall temperature, FCU fan outlet temperature, FCU fan outlet volume, etc. at each decision stage in the room.
[0087] (3) A user-personalized comfort model based on PMV index is used to calculate the user's comfort index based on indoor temperature, wall temperature, etc.
[0088] (4) The problem of optimizing the trade-off between personalized comfort and energy consumption and the use of a reinforcement learning algorithm framework for training the decision library.
[0089] During the training of the decision base, the reward function value is calculated based on the energy consumption of the HVAC system, the user's comfort index, and the weighting coefficient. Then, the decision value is calculated based on the reward function value and the indoor temperature, wall temperature, FCU fan outlet temperature, and FCU fan outlet volume corresponding to multiple decision stages.
[0090] like Figure 4 The diagram shown is a schematic representation of the computational flow of a reinforcement learning method that balances personalized thermal comfort and HVAC energy consumption in an embodiment of this specification.
[0091] The process of constructing a mechanism-based energy consumption model for HVAC systems is as follows:
[0092] First, let's analyze the mechanism of HVAC systems:
[0093] HVAC systems can be used to regulate indoor air temperature and humidity and provide fresh air. A typical structure is as follows: Figure 2 As shown. The system specifically includes: fan coil units (FCU), fresh air units (FAU), chillers, circulating pumps, and cooling towers. Each room is equipped with an FCU, while the FAU is shared by all rooms. In summer, the chiller supplies cooling water to the FCU and FAU. When the heat exchange coils are cooled by the cooling water, the fans blow hot, humid air across the coils to cool and dehumidify it, thus providing cool air to the building space. The cooling water that has absorbed heat exits the heat exchange coils and flows back to the chiller, where it is cooled again, and the process is repeated. The airflow is controlled by the fan speed, and the cooling water flow rate is controlled by water valves. The difference between the FCU and FAU lies in the source of the air. Indoor air is drawn into the FCU unit, primarily for temperature and humidity regulation; outdoor air is drawn into the FAU unit, primarily for fresh air supply. In winter, the chiller produces hot water, and the rest of the process is similar to that in summer.
[0094] The embodiments in this specification make the following assumptions when analyzing the above-described HVAC system:
[0095] 1) Only the air conditioning energy consumption of a single room is considered, without considering the mutual influence and coupling between different rooms;
[0096] 2) There is only one user in the room, and the comfort needs of that user are the target for the comfort level of the room;
[0097] 3) The windows are in the default closed state, meaning the impact of natural ventilation is not considered at this time;
[0098] 4) Ignore the effect of FAU and only consider the effect of FCU on temperature regulation;
[0099] 5) Since only a single room is considered, it is assumed that the cooling demand of a single room will not exceed the maximum load of the chiller;
[0100] 6) Consider the air conditioning operation strategy for the room in the next 24 hours, and make decisions in 30-minute intervals, numbered 1-48.
[0101] Then calculate the energy consumption of the HVAC system:
[0102] The energy consumption of an HVAC system mainly consists of two parts: fan energy consumption and cooling energy consumption.
[0103] FCU fan power P FCU Depending on the fan speed, this variable can be expressed as the air volume G. FCU Characterization. Existing literature indicates that the energy consumption of the fan is non-linearly related to the air volume, which can be characterized by formula (1):
[0104]
[0105] in, This represents the power of the FCU fan in the k-th decision stage; It is the rated power of the FCU fan in the k-th decision stage; This represents the air output of the FCU fan in the k-th decision stage; It is the rated output air volume of the FCU fan in the k-th decision stage.
[0106] Refrigeration energy consumption mainly depends on the refrigeration capacity. The refrigeration capacity can be characterized by the energy difference before and after air enters the FCU, and the physical quantity "enthalpy" precisely represents the energy contained in air and water vapor. Therefore, the refrigeration capacity C... FCU It can be measured by the air enthalpy at the air inlet (EN) FCU,inlet and air enthalpy at the air outlet (EN) FCU,outlet The difference is used for calculation. For the FCU, it draws in indoor air, so the cooling capacity of the FCU can be expressed by the following formula (2):
[0107]
[0108] in, It is the cooling capacity of the FCU fan in the k-th decision stage, EN FCU,inlet It is the air enthalpy at the air inlet of the FCU fan in the kth decision stage; C is the air enthalpy at the FCU fan outlet during the k-th decision stage; p It is the specific heat of air; These are the indoor air temperature and the FCU outlet air temperature, respectively, during the k-th decision-making stage. These are the absolute humidity of the indoor air and the air at the FCU outlet during the k-th decision-making stage, respectively.
[0109] After calculating the cooling capacity, the energy consumption of other equipment such as chillers, circulating pumps, and cooling towers can be calculated using the Coefficient of Performance (COP), which is defined as the ratio of the cooling capacity of the FCU to the electrical power consumed by the chiller, circulating pumps, and cooling towers.
[0110] After calculating the hourly energy consumption of the air conditioner, this instruction manual uses the time-of-use electricity price (c). k Calculate energy consumption, where c k This represents the electricity value corresponding to the k-th decision stage. For example, see Table 1 for specific time-of-use electricity prices:
[0111] Table 1 Time-of-use Electricity Price Table
[0112]
[0113] Therefore, the energy consumption of the HVAC system can be calculated as follows (3):
[0114]
[0115] Where Energy_cost represents the energy consumption of the HVAC system; c k Δt represents the time-of-use electricity price at the k-th decision stage; Δt represents a single decision cycle of 30 minutes (1800 seconds); COP represents the performance coefficient. Equation (3) represents the total energy consumption of the HVAC system from the start of the decision to the current decision stage.
[0116] The process of constructing a room heat transfer mechanism model based on the heat balance method is as follows:
[0117] To assess the impact of HVAC systems on indoor users, it is essential to characterize the dynamic heat transfer process within the room. The indoor heat transfer process involves numerous components and complex calculations, primarily including: heat exchange between air at different indoor temperatures, heat exchange between walls and air, heat conduction within the walls themselves, heat exchange between indoor and outdoor air at different temperatures, heat absorption by walls and air from solar radiation, and heat generation by indoor heat sources (users and electrical equipment).
[0118] Indoor air temperature, humidity, wall temperature, and CO2 concentration are often used as key parameters to describe the aforementioned heat transfer kinetics, because air, water vapor, and CO2 can act as energy carriers and reflect dynamic characteristics such as energy conversion and transfer. Therefore, it is necessary to characterize the dynamic changes of these variables under the influence of an HVAC system.
[0119] In order to reasonably model the room heat transfer dynamics based on the above variables, and at the same time simplify the calculations as much as possible, this paper makes the following assumptions:
[0120] 1) The room air temperature is uniformly distributed, meaning that the current indoor air is a uniform mixture of the previous indoor air and the air after cooling and dehumidification by the FCU, and has undergone sufficient heat exchange with the walls and outdoor air, so there is no spatial temperature difference.
[0121] 2) The total mass of air in the room is constant, that is, the pressure in the room remains constant;
[0122] 3) The temperature of the two surfaces of the building's interior walls is the same, meaning that the heat exchange process between adjacent rooms through the walls is not considered.
[0123] Based on the above assumptions and the laws of conservation of energy and mass, the state variables change and update according to the following rules:
[0124] First, regarding indoor air temperature, the indoor temperature in the (k+1)th decision stage is affected by the following factors:
[0125] 1) Heat generated by users and equipment;
[0126] 2) Heat transferred from the walls;
[0127] 3) Heat transferred from the outside air through windows;
[0128] 4) Heat provided by the FCU;
[0129] 5) The heat already contained in the air.
[0130] According to one embodiment of this specification, the quantitative change relationship of indoor temperature can be characterized by formula (4):
[0131]
[0132]
[0133] Where, m a It refers to the air quality inside the room. Q represents the indoor temperature at the (k+1)th decision stage, Δt represents the time length corresponding to a single decision cycle, and Q is the temperature at the (k+1)th decision stage. g It is the heat production rate per person. and h represents the heat generated by light and indoor equipment at the k-th decision stage, respectively. gs It is the heat transfer coefficient of indoor and outdoor air through the window, A gs It is the area of the window. The outdoor temperature at the k-th decision stage. The indoor air temperature, h, is at the k-th decision stage. w It is the heat transfer coefficient between the wall and the indoor air, A w It is the wall area. The wall temperature, C, is at the k-th decision stage. p It is the specific heat of air. This represents the air output of the FCU fan in the k-th decision stage. It is the outlet air temperature of the FCU fan in the k-th decision stage. It is the energy contained in the remaining air after removing the air pumped out by the FCU in the room.
[0134] Secondly, regarding wall temperature, the wall temperature in the (k+1)th decision stage is mainly affected by the following factors:
[0135] 1) Convective heat transfer between the wall and the indoor air;
[0136] 2) Absorption of solar heat by walls and indoor air.
[0137] Specifically, according to one embodiment of this specification, applying the mass and energy conservation formula to a wall yields the following formula (5):
[0138]
[0139] Among them, C w It is the heat capacity of the wall, m w It's about the quality of the wall. Let S represent the difference between the wall temperature at the (k+1)th decision stage and the wall temperature at the kth decision stage. w This represents the amount of heat absorbed by indoor air and walls from solar radiation.
[0140] Based on formulas (4) and (5), the variables characterizing room heat transfer are dynamically updated.
[0141] The process of constructing a user-personalized comfort model based on the PMV index is as follows:
[0142] Thermal comfort, as an important indicator of building performance evaluation, plays a crucial role in the energy-efficient operation of buildings. Thermal comfort reflects users' satisfaction with their thermal environment and depends on the heat interaction between users and their surroundings. Users feel relatively comfortable when the heat they generate is equal to the heat they release to the environment.
[0143] Most existing literature sets variables such as temperature and humidity that affect comfort as a static range and uses them as constraints in optimization problems. However, comfort is determined by both environmental variables and personal factors, specifically including six factors: indoor air temperature, mean radiant temperature, relative humidity, air velocity, clothing coefficient, and the user's metabolic rate. To characterize the coupled effects of these variables and reflect the dynamic changes in comfort with the environment, this specification introduces the "Predicted Average Value" (PMV) index to characterize comfort.
[0144] The above six indicators can be mapped to PMV values according to certain rules, with a value range of [-3, +3]. The correspondence between the values and the sensation of hot and cold is shown in Table 2.
[0145] Table 2. Correspondence between thermal sensation and PMV value
[0146]
[0147] Specifically, according to one embodiment of this specification, the PMV value is calculated according to the following formulas (6)-(11):
[0148] PMV=[0.0303 exp(-0.036M)+0.028]L (6)
[0149] L = M - 3.05 × 10 -3 ×[5733-6.99MP a -0.42*[M-58.15]-1.7
[0150] ×10 -5 M(5867-P a -0.0014M(34-T) a )-F cl *h c *(T cl -T a )
[0151] -3.96×10 -8 F cl [(T cl +273) 4 -(T r +273) 4 (7)
[0152] T cl =35.7-0.028M
[0153] -I cl F cl 3.96×10 -8 [(T cl +273) 4 -(T r +273) 4 ]+h c (T cl -T a )} (8)
[0154]
[0155]
[0156] P a =6.1094h r ×exp[17.625T a / (T a +243.04)] (11)
[0157] Where PMV represents the comfort index, which is the calculated PMV value; M represents the metabolic rate, measured in W / m³. 2 L represents the human body heat load, measured in W / m². 2 ;P a T is the vapor pressure of water, measured in Pa; a It is the indoor air temperature, T cl It refers to the surface temperature of the clothing, T. r This is the mean radiant temperature, and the unit is °C; F cl It is the surface area coefficient of clothing; I cl It is the thermal insulation coefficient of the clothing surface, and the unit is m. 2 K / W; h c It is the convective heat transfer coefficient, with units of W / m². 2 K;V a It is the air velocity, measured in m / s; h r Relative humidity is expressed as a percentage.
[0158] Indoor temperature and relative humidity can be directly measured by sensors. Clothing surface temperature, convective heat transfer coefficient, clothing surface area coefficient, and water vapor pressure can be calculated using the formulas mentioned above. The insulation coefficient of the clothing surface depends on the user's attire and can be flexibly adjusted according to the season. According to relevant research, air conditioning has replaced traditional fans in most offices, therefore indoor air velocity is typically below 0.2 m / s. To simplify the calculation, the average radiant temperature T... r Approximately the air temperature T a equal.
[0159] During the training phase, the metabolic rate M is a set value. Optionally, the value of M is in the range of [0.5, 1.5], with an interval of 0.1, to build multiple decision bases.
[0160] Therefore, given any environmental variable parameters, the real-time PMV value can be calculated based on the above principle.
[0161] The recommended comfort range is [-0.5, +0.5]. To characterize the cost of violating comfort in the current environment, the comfort cost is defined in the following (12) form:
[0162]
[0163] Comfort_cost represents the comfort cost calculated based on the comfort index PMV.
[0164] Under this definition of comfort, any deviation from the normal range will result in a certain degree of penalty. This additional penalty value can be negotiated with energy consumption costs, thereby achieving a balance between comfort and energy consumption.
[0165] In some other embodiments of this specification, the comfort cost can also be calculated using formula (13):
[0166]
[0167] Where β is the weighting coefficient.
[0168] The process of constructing the personalized comfort and energy consumption trade-off optimization problem and the reinforcement learning algorithm framework based on Q-learning is as follows:
[0169] First, we need to address the issue of balancing personalized comfort with energy consumption:
[0170] The problem studied in this specification is a typical optimization problem with constraints, defined as achieving the optimal trade-off between HVAC system energy consumption and user-personalized comfort under the premise of satisfying equipment constraints and system dynamic laws, i.e., formula (14):
[0171]
[0172] Where J represents the trade-off index, k represents the total number of decision-making stages, and λ represents the weighting coefficient, used to balance energy costs and comfort costs.
[0173] Then, a reinforcement learning algorithm framework based on Q-learning is proposed:
[0174] To solve the above problems, this document adopts a reinforcement learning algorithm framework that is more adaptable to dynamic changes, and the steps are as follows:
[0175] Step 1: Determine the state variables and decision variables
[0176] This manual focuses on temperature control and does not consider changes in humidity and CO2 concentration. Therefore, the indoor air temperature at decision stage k is selected. Wall temperature As a state variable, i.e., formula (15):
[0177]
[0178] Where, x k This represents the state variable at the k-th decision stage. The indoor air temperature at the k-th decision stage. It is the wall temperature at the k-th decision stage.
[0179] HVAC modeling reveals that the most significant factors influencing energy consumption are outlet air temperature and outlet air volume. Since this specification does not consider the role of the FAU (Fan Activated Unit), the outlet air temperature of the FCU (Fan Controller Unit) will be selected at the current decision-making stage. With air volume As an action variable, i.e., formula (16):
[0180]
[0181] Among them, u k This represents the action variable for the k-th decision stage.
[0182] ①State and Action Space Aggregation
[0183] Because this manual takes into account energy-saving scheduling strategies for air conditioning in summer, the indoor air temperature T is set. a The value range is 20℃~35℃, with a discrete interval of 1℃; wall temperature T w The value range is 20℃~35℃, with a discrete interval of 0.1℃; the air conditioner outlet temperature T FCU The value range is 18-30℃, with a discrete interval of 1℃; air conditioner output air volume G FCU Four values are available: zero, 30% of rated air volume, 60% of rated air volume, and 100% of rated air volume.
[0184] As can be seen, the state space of this problem is extremely large, with thousands of possible combinations of indoor air temperature and wall temperature alone. However, in reality, due to the continuous heat exchange process, the air temperature and wall temperature cannot differ significantly. Furthermore, during the actual operation of an air conditioner in summer, the outlet air temperature cannot be higher than the indoor air temperature. Therefore, many state-action pairs obtained using the aforementioned state-action space design method are unlikely to occur in actual operation. The Q-learning method requires a large-scale search of the state-action space, which to some extent leads to a significant waste of computational resources.
[0185] To eliminate a large amount of useless state space, we consider introducing a "temperature difference" to modify the state and action variables. Assuming the temperature difference between the wall and the air does not exceed 2°C, we introduce the variable... The difference between the wall temperature and the air temperature is represented by a value in the range of [-2, +2], with a distance of 0.1℃. The corrected state variable can then be expressed as (17):
[0186]
[0187] Similarly, introducing variables The difference between the air conditioner outlet temperature and the air temperature is represented by a value in the range of [-15, 0] and a distance of 1℃. The corrected action variable can then be expressed as (18):
[0188]
[0189] Through the above correction process, the size of the state-action space is greatly compressed, which greatly improves the performance of the Q-learning algorithm.
[0190] ② Time dependence of Q-value processing
[0191] As can be seen from formula (3), total energy consumption is affected by real-time electricity prices. Assuming the same action is taken under the same conditions, different decision-making periods will result in different time-of-use electricity prices, which in turn will lead to different energy costs, ultimately affecting the Q value. This makes the Q value time-dependent, meaning that the Q value calculated at this time is only valid for states with specific time-of-use electricity prices. However, this time information cannot be directly reflected in the existing state variables, making it difficult to train and apply the Q value.
[0192] To solve the above problems, consider directly adding the current decision time variable t to the state variables, i.e., formula (19):
[0193]
[0194] Using the above method, the corresponding time-of-use electricity price can be determined based on the current time t, then state transitions can be performed and the Q value can be calculated. In this way, the time information of the Q value is reflected in the time variable in the state variables.
[0195] Step 2: Design the reward function
[0196] The reward function is the core part of reinforcement learning for control. Many existing papers use a fixed comfort range as a constraint, so the reward function only includes the energy consumption component. Since this specification introduces a comprehensive and dynamic PMV index to characterize comfort and calculates the comfort cost accordingly, both energy consumption cost and comfort cost should be included in the reward function to balance their relationship. Therefore, the objective function from problem (14) is directly used as the definition of the reward function. To transform the problem of minimizing energy consumption into the problem of maximizing reward, a negative sign is added to the reward, as shown in formula (20).
[0197] reward=-(Energy_cost+λ*Comfort_cost) (20)
[0198] Where reward represents the reward function value, Energy_cost represents the energy consumption of the HVAC system, and λ represents the weighting coefficient.
[0199] It should be noted that, in addition to using the linear weighting of Energy_cost and Comfort_cost, other non-linear weighting methods can also be used to define the reward function. This specification does not impose any restrictions on the embodiments.
[0200] It is worth noting that the weighting coefficients of energy consumption and comfort also reflect the user's personalized preferences. Existing literature does not have a specific study on how to determine the weighting coefficients. Most of them are obtained through parameter tuning. This method of parameter selection often relies too much on the experience of the tuner and does not take into account the individual differences of different users.
[0201] Therefore, this specification proposes a personalized weight parameter determination rule, which believes that the weight parameter λ is affected by the following three factors: user energy-saving potential, user sensitivity, and user consumption habits.
[0202] ① User energy-saving potential
[0203] The user's energy-saving potential α is defined as the rate of energy consumption reduction when the PMV = 0.5 is relaxed by 0.1 as a strict constraint. That is, if the total energy consumption after one decision cycle is E1 under the strict constraint of PMV = 0.6, and the total energy consumption after operating under PMV = 0.5 is E2, then ΔE = (E2 - E1) / E2 * 100% is the energy-saving potential. Clearly, the greater the energy-saving potential, the more "appropriate" it is to sacrifice comfort for energy reduction, and the smaller the value of λ should be.
[0204] ② User sensitivity
[0205] User sensitivity β is defined as the average rate of change in PMV value for every 1°C change in temperature within an adjustable temperature range. The more sensitive the user, the greater the cost of violating comfort should be; therefore, the value of λ should be larger.
[0206] ③ User consumption habits
[0207] User consumption habits (γ) refer to whether a user is willing to pay more to improve their comfort experience. Time-of-use electricity pricing affects a user's comfort range, which depends on their consumption habits—whether they are frugal, wasteful, or neutral. Obviously, the more "frugal" a user is, the lower the cost of compromising comfort should be, so the value of λ should be smaller. This indicator can be defined as a number between 0 and 1; the larger the number, the more "luxurious" the user is.
[0208] Since this manual uses metabolic rate as a personalized characteristic for users, it can assess users' energy-saving potential and sensitivity under different metabolic rate conditions. To simplify calculations, other indicators in the PMV index besides metabolic rate can be treated as statistically significant averages. Users' consumption habits can be discretized into three criteria: "thrifty," "neutral," and "wasteful," which can be assessed through questionnaires and other methods regarding consumption habits.
[0209] Therefore, the weighting coefficient λ can be seen as the weighted sum of the above three user personalization indicators (21):
[0210] λ=(m*β-n*α)*γ (21)
[0211] Where λ represents the weighting coefficient; m and n are adjustable parameters; α represents the user's energy-saving potential, defined as the energy reduction rate when PMV = 0.5 is relaxed to 0.1 as a strict constraint; β represents the user's sensitivity, defined as the average change rate of PMV value for every 1°C change in temperature within the adjustable temperature range; γ represents the user's consumption habits, defined as a number between 0 and 1, the larger the value, the more willing the user is to pay more costs to improve their comfort experience.
[0212] Step 3: Algorithm Design
[0213] Considering that the state variables and action variables in this problem are both finite and discrete, Q-learning, a representative algorithm of temporal difference learning in reinforcement learning, is chosen for algorithm design.
[0214] The training process involves several scenarios, each with a termination condition. Since the state variables in this specification include time, each scenario is designed with 48 decisions from 0:00 to 24:00. The scenario is considered to have met the termination condition when 24 hours have elapsed.
[0215] The overall algorithm for Q-learning applied to HVAC system optimization can be divided into two parts: Q-value training and Q-value decision-making. The Q-value training algorithm is used before decision-making, and through continuous training and learning, it fully evaluates the "future reward" of different state-action pairs. The Q-value decision-making algorithm is used for actual HVAC system control, and uses the learned experience and actual measurement data to output the optimal decision.
[0216] It is worth noting that the metabolic rate characterizes the individual differences of users. Therefore, different metabolic rate values can be used for training during the algorithm training stage, such as between 1.0 and 2.1. Based on this, the user's "energy-saving potential" and "sensitivity" can be evaluated and added to the reward function for training. At the same time, based on the user's questionnaire feedback, the user's consumption habits can be discretized into three standards: "thrifty", "neutral", and "wasteful" and added to the reward function for training. In this way, a strategy library for people with "different hot and cold preferences and different consumption habits" can be learned. In the algorithm decision-making stage, based on the real-time voting data of the users on site, the metabolic rate value of the user can be fitted according to the formula (23). At the same time, by asking the user about their consumption habits, the type of the user on site can be found. The corresponding strategy can be found and executed according to the strategy library.
[0217] The specific algorithm flow is as follows:
[0218]
[0219]
[0220] Through iterative training using the above algorithm, a decision library is finally obtained when the metabolic rate is M.
[0221] It should be noted that, in addition to Q-learning, the framework proposed in this specification is also applicable to other types of reinforcement learning algorithms.
[0222] It is worth noting that different users will have different physical reactions and different thermal sensations in the same environment. For example, under the conditions of 22℃-26℃ commonly used in existing literature, some people may feel "hot" and 20℃ may be more suitable for them, while others may feel "cold" and 28℃ may be sufficient to meet their thermal needs. This personalized thermal comfort preference can be reflected by the metabolic rate in the PMV model. Therefore, in order to enable the statistically significant PMV index model to adapt to personalized comfort needs, the user's metabolic rate can be obtained by training a machine learning algorithm from the user's actual votes in the real environment (also a number between -3 and +3), as follows (23):
[0223]
[0224] Where N represents the total number of votes cast by the user, Vote represents the actual number of votes cast by the user, and parameter i This represents the environmental parameters in which the user is voting at the time of the i-th vote, including the indoor air temperature T. a Mean radiation temperature T r Relative humidity h r air velocity V aClothing coefficient I cl PMV(M, parameter) i () indicates that the metabolic rate is M and the environmental parameter is parameter. i The PMV value calculated at that time.
[0225] According to one embodiment of this specification, the steps for determining the target decision base based on the target user's metabolic rate include:
[0226] The metabolic rate of the target user was calculated using formula (23);
[0227] Determine the decision library corresponding to the metabolic rate of the target user to obtain the target decision library;
[0228] like Figure 3 As shown, the steps of using the target decision library to make decisions based on the actual air temperature and actual wall temperature in the target room, and obtaining the target air volume and target air outlet temperature of the FCU fan corresponding to multiple decision stages, include:
[0229] Step 301: Construct the current state variables based on the actual air temperature, the actual wall temperature, and the current decision-making stage.
[0230] Step 302: Determine the maximum decision value corresponding to the current state variable in the target decision base, and take the action variable corresponding to the maximum decision value as the optimal action;
[0231] Step 303: Based on the current state variables and the optimal action, use the room heat transfer mechanism model based on the heat balance method to calculate the next state variables corresponding to the next decision stage;
[0232] Step 304: Take the next state variable as the current state variable, and repeatedly execute the step of determining the current optimal action corresponding to the current state variable in the target decision library until the optimal action corresponding to all decision stages is determined. The optimal action includes the target air volume of the FCU fan and the target air temperature of the FCU fan.
[0233] Specifically, the algorithm is as follows:
[0234]
[0235]
[0236] Thus, a reinforcement learning decision-making method has been developed that balances user-specific thermal comfort with building HVAC system energy consumption, maximizing energy savings while meeting user comfort needs.
[0237] The methods described in this specification can fully meet the thermal comfort needs of users with different heating and cooling preferences and different consumption habits, avoiding unnecessary energy waste; can make full use of the low electricity price advantage to achieve energy saving through the pre-cooling mechanism; and can adjust the operation strategy in a timely manner based on real-time user feedback, exhibiting strong robustness.
[0238] Based on the same inventive concept, embodiments of this specification also provide a reinforcement learning device that balances personalized thermal comfort with HVAC energy consumption, such as... Figure 5 As shown, it includes:
[0239] The input data construction unit 501 is used to take the state variables consisting of the difference between the indoor temperature, the wall temperature and the indoor temperature, and the time variable, as well as the action variables consisting of the difference between the outlet air temperature of the FCU fan in the HVAC system and the indoor temperature, and the air volume of the FCU fan as input data.
[0240] The decision base training unit 502 is used to train the input data using a reinforcement learning algorithm with predetermined parameters for multiple predetermined metabolic rates to obtain a decision base corresponding to each metabolic rate. The decision base includes decision values corresponding to each pair of multiple state variables and multiple action variables. The decision values are calculated based on the reward function value and the state variables and action variables corresponding to multiple decision stages. The state variables corresponding to the current decision stage are obtained by iteratively calculating the state variables and action variables of the previous decision stage using a pre-built room heat transfer mechanism model based on the heat balance method. The action variables are selected through an ε-greedy strategy. The reward function value is calculated using the comfort index calculated by the state variables, the energy consumption of the HVAC system calculated by the state variables and action variables, and weight coefficients calculated based on user energy-saving potential, user sensitivity, and user consumption habits.
[0241] The decision unit 503 is used to determine a target decision library based on the target user's metabolic rate when making decisions about the air volume and air outlet temperature of the FCU fan in the target room where the target user is located. The decision unit uses the target decision library to make decisions based on the actual air temperature and actual wall temperature in the target room, and obtains the target air volume and target air outlet temperature of the FCU fan corresponding to multiple decision stages. This allows the FCU fan in the target room to be controlled using the target air volume and target air outlet temperature at the time corresponding to each decision stage, so that the indoor temperature in the target room meets the comfort level of the target user.
[0242] Since the principle of the above-mentioned device in solving the problem is similar to that of the above-mentioned method, the implementation of the above-mentioned device can refer to the implementation of the above-mentioned method, and the repeated parts will not be described again.
[0243] like Figure 6 The illustration shows a computer device provided in an embodiment of this specification. The apparatus described herein can be the computer device described in this embodiment, performing the methods described above. The computer device 602 may include one or more processors 604, such as one or more central processing units (CPUs), each of which can implement one or more hardware threads. The computer device 602 may also include any memory 606 for storing information of any kind, such as code, settings, data, etc. Without limitation, for example, memory 606 may include any type of RAM, any type of ROM, flash memory, hard disk, optical disk, etc. More generally, any memory can use any technology to store information. Further, any memory can provide volatile or non-volatile retention of information. Further, any memory may represent a fixed or removable component of the computer device 602. In one case, when processor 604 executes associated instructions stored in any memory or combination of memories, the computer device 602 can perform any operation of the associated instructions. The computer device 602 also includes one or more drive mechanisms 608 for interacting with any memory, such as hard disk drive mechanisms, optical disk drive mechanisms, etc.
[0244] Computer device 602 may also include an input / output module 610 (I / O) for receiving various inputs (via input device 612) and providing various outputs (via output device 614). A specific output mechanism may include a presentation device 616 and an associated graphical user interface (GUI) 618. In other embodiments, the input / output module 610 (I / O), input device 612, and output device 614 may be omitted, and the device may function solely as a computer device within a network. Computer device 602 may also include one or more network interfaces 620 for exchanging data with other devices via one or more communication links 622. One or more communication buses 624 couple the components described above together.
[0245] Communication link 622 can be implemented in any way, such as via a local area network, a wide area network (e.g., the Internet), a point-to-point connection, or any combination thereof. Communication link 622 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc., governed by any protocol or combination of protocols.
[0246] This specification also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0247] This specification also provides computer-readable instructions, wherein when a processor executes the instructions, the program therein causes the processor to perform the above-described method.
[0248] It should be understood that in the various embodiments of this specification, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this specification.
[0249] It should also be understood that, in the embodiments of this specification, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the embodiments of this specification, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0250] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this specification can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of the embodiments in this specification.
[0251] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0252] In the embodiments provided in this specification, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections through some interfaces, devices, or units, or they may be electrical, mechanical, or other forms of connection.
[0253] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments described in this specification, depending on actual needs.
Claims
1. A reinforcement learning method for balancing personalized thermal comfort and HVAC energy consumption, characterized in that, The method includes: The state variables, consisting of the difference between the indoor temperature, the wall temperature and the indoor temperature, and the time variable, and the action variables, consisting of the difference between the outlet air temperature of the FCU fan in the HVAC system and the indoor temperature, and the air volume of the FCU fan, are used as input data. The input data is trained using a reinforcement learning algorithm with predetermined parameters for multiple predetermined metabolic rates to obtain a decision base corresponding to each metabolic rate. The decision base includes decision values corresponding to each pair of multiple state variables and multiple action variables. The decision values are calculated based on the reward function value and the state variables and action variables corresponding to multiple decision stages. The state variables corresponding to the current decision stage are obtained by iteratively calculating the state variables and action variables of the previous decision stage using a pre-built room heat transfer mechanism model based on the heat balance method. The action variables are selected using an ε-greedy strategy. The reward function value is calculated using the comfort index calculated from the state variables, the energy consumption of the HVAC system calculated from the state variables and action variables, and weighting coefficients calculated based on user energy-saving potential, user sensitivity, and user consumption habits. When making decisions regarding the airflow rate and outlet temperature of the FCU fan in the target room where the target user is located, a target decision library is determined based on the target user's metabolic rate. The target decision library is then used to make decisions based on the actual air temperature and wall temperature in the target room, resulting in target airflow rate and target outlet temperature for the FCU fan at multiple decision stages. This allows for the control of the FCU fan in the target room using the target airflow rate and target outlet temperature at the time corresponding to each decision stage, thereby ensuring that the indoor temperature in the target room meets the comfort level of the target user.
2. The method according to claim 1, characterized in that, The formula for calculating the comfort index using the aforementioned state variables is as follows: PMV=[0.0303exp(-0.036M)+0.028]L; L=M-3.05×10 -3 ×[5733-6.99M-P a ]-0.42*[M-58.15]-1.7×10 -5 M(5867-P a )-0.0014M(34-T a )-F cl *h c *(T cl -T a )-3.96×10 -8 F cl [(T cl +273) 4 -(T r +273) 4 ]; T cl =35.7-0.028M-I cl F cl {3.96×10 -8 [(T cl +273) 4 -(T r +273) 4 ]+h c (T cl -T a )}; P a =6.1094h r ×exp[17.625T a / (T a +243.04)]; Where PMV represents the comfort index, which is the calculated PMV value; M represents the metabolic rate, measured in W / m³. 2 L represents the human body heat load, measured in W / m². 2 ;P a T is the vapor pressure of water, measured in Pa; a It is the indoor air temperature, T cl It refers to the surface temperature of the clothing, T. r This is the mean radiant temperature, and the unit is °C; F cl It is the surface area coefficient of clothing; I cl It is the thermal insulation coefficient of the clothing surface, and the unit is m. 2 K / W; h c It is the convective heat transfer coefficient, with units of W / m². 2 K;V a It is the air velocity, measured in m / s; h r Relative humidity is expressed as a percentage.
3. The method according to claim 2, characterized in that, The formula for calculating the weighting coefficients based on user energy-saving potential, user sensitivity, and user consumption habits is as follows: λ=(m*β-n*α)*γ; Where λ represents the weighting coefficient; m and n are adjustable parameters; α represents the user's energy-saving potential, defined as the energy reduction rate when PMV = 0.5 is relaxed to 0.1 as a strict constraint; β represents the user's sensitivity, defined as the average change rate of PMV value for every 1°C change in temperature within the adjustable temperature range; γ represents the user's consumption habits, defined as a number between 0 and 1, the larger the value, the more willing the user is to pay more costs to improve their comfort experience.
4. The method according to claim 3, characterized in that, The formula for calculating the reward function value is: reward=-(Energy_cost+λ*Comfort_cost); Where reward represents the reward function value, Energy_cost represents the energy consumption of the HVAC system, and λ represents the weighting coefficient; Comfort_cost represents the comfort cost calculated based on the comfort index PMV.
5. The method according to claim 1, characterized in that, Using a pre-built room heat transfer mechanism model based on the heat balance method, the state variables and action variables of the previous decision-making stage are iteratively calculated to obtain the formula for the state variables of the current decision-making stage: Where, m a It refers to the air quality inside the room. Q represents the indoor temperature at the (k+1)th decision stage, Δt represents the time length corresponding to a single decision cycle, and Q is the temperature at the (k+1)th decision stage. g It is the heat production rate per person. and h represents the heat generated by light and indoor equipment at the k-th decision stage, respectively. gs It is the heat transfer coefficient of indoor and outdoor air through the window, A gs It is the area of the window. The outdoor temperature at the k-th decision stage. The indoor air temperature, h, is at the k-th decision stage. w It is the heat transfer coefficient between the wall and the indoor air, A w It is the wall area. The wall temperature, C, is at the k-th decision stage. p It is the specific heat of air. This represents the air output of the FCU fan in the k-th decision stage. It is the outlet air temperature of the FCU fan in the k-th decision stage. It is the energy contained in the remaining air obtained after removing the air pumped out by the FCU in the room; Among them, C w It is the heat capacity of the wall, m w It's about the quality of the wall. Let S represent the difference between the wall temperature at the (k+1)th decision stage and the wall temperature at the kth decision stage. w This represents the amount of heat absorbed by indoor air and walls from solar radiation.
6. The method according to claim 5, characterized in that, The reinforcement learning algorithm is a Q-learning algorithm, and the decision value is the Q-value; The formula for calculating the decision value based on the reward function value and the state and action variables corresponding to multiple decision stages is as follows: Q M (x,u) new =Q M (x,u) old *[1-θ]+θ[reward+γ*max u_next Q M (x next ,u next )]; Among them, Q M (x,u) new This represents the Q-value calculated by the Q-learning algorithm in the current iteration when the metabolic rate is M, based on the state variable x and the action variable u. M (x,u) old When the metabolic rate is M, Q represents the Q-value calculated by the state variable x and action variable u in the previous iteration of the Q-learning algorithm; θ represents the learning rate in the predetermined parameters; reward represents the reward function value; γ represents the discount coefficient in the predetermined parameters; and max represents the maximum value. u_next Q M (x next ,u next This indicates that the state variable x for the next decision stage is obtained by iteratively calculating the state variables and action variables of the current decision stage using a pre-built room heat transfer mechanism model based on the heat balance method. next And based on the state variable x next Iterate through all action variables u next For the corresponding Q value, select the maximum value among the Q values.
7. The method according to claim 6, characterized in that, The formula for the state variable x is: Where, x k This represents the state variable at the k-th decision stage. The indoor air temperature at the k-th decision stage. The wall temperature is at the k-th decision stage, and t represents the current decision time variable. The formula for the action variable u is: Among them, u k This represents the action variable in the k-th decision stage. This represents the outlet air temperature of the FCU fan in the k-th decision stage. Indoor air temperature at the k-th decision stage The difference between them This represents the air output of the FCU fan in the k-th decision stage.
8. The method according to claim 7, characterized in that, The steps for determining the target decision base based on the target user's metabolic rate include: Using formula Calculate the metabolic rate of the target user, where M represents the metabolic rate of the target user, N represents the total number of votes cast by the user, and Vote represents the actual vote value cast by the user. i This represents the environmental parameters in which the user is voting at the time of the i-th vote, including the indoor air temperature T. a Mean radiation temperature T r Relative humidity h r air velocity V a Clothing coefficient I cl PMV(M, parameter) i () indicates that the metabolic rate is M and the environmental parameter is parameter. i The PMV value calculated at that time; Determine the decision library corresponding to the metabolic rate of the target user to obtain the target decision library; The steps for making decisions based on the actual air temperature and actual wall temperature inside the target room using the target decision database, and obtaining the target air volume and target air outlet temperature of the FCU fan corresponding to multiple decision stages, include: Construct the current state variables based on the actual air temperature, the actual wall temperature, and the current decision-making stage. In the target decision base, determine the maximum decision value corresponding to the current state variable, and take the action variable corresponding to the maximum decision value as the optimal action; Based on the current state variables and the optimal action, the next state variables corresponding to the next decision stage are calculated using the room heat transfer mechanism model based on the heat balance method. The next state variable is used as the current state variable, and the step of determining the current optimal action corresponding to the current state variable in the target decision base is repeated until the optimal action corresponding to all decision stages is determined. The optimal action includes the target air volume of the FCU fan and the target air temperature of the FCU fan.
9. A reinforcement learning device that balances personalized thermal comfort with HVAC energy consumption, characterized in that, The device includes: The input data construction unit is used to take the state variables consisting of the difference between the indoor temperature, the wall temperature and the indoor temperature, and the time variable, as well as the action variables consisting of the difference between the outlet air temperature of the FCU fan in the HVAC system and the indoor temperature, and the air volume of the FCU fan as input data. The decision base training unit is used to train the input data using a reinforcement learning algorithm with predetermined parameters for multiple predetermined metabolic rates to obtain a decision base corresponding to each metabolic rate. The decision base includes decision values corresponding to each pair of multiple state variables and multiple action variables. The decision values are calculated based on the reward function value and the state variables and action variables corresponding to multiple decision stages. The state variables corresponding to the current decision stage are obtained by iteratively calculating the state variables and action variables of the previous decision stage using a pre-built room heat transfer mechanism model based on the heat balance method. The action variables are selected through an ε-greedy strategy. The reward function value is calculated using the comfort index calculated from the state variables, the energy consumption of the HVAC system calculated from the state variables and action variables, and weight coefficients calculated based on user energy-saving potential, user sensitivity, and user consumption habits. The decision-making unit is used to determine a target decision library based on the target user's metabolic rate when making decisions about the air volume and outlet temperature of the FCU fan in the target room where the target user is located. It then uses the target decision library to make decisions based on the actual air temperature and actual wall temperature in the target room, obtaining target air volume and target outlet temperature of the FCU fan for multiple decision stages. This allows for the control of the FCU fan in the target room using the target air volume and target outlet temperature at the time corresponding to each decision stage, thereby ensuring that the indoor temperature in the target room meets the comfort level of the target user.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 8.