Time-varying reinforcement learning for robust adaptive estimator design with application to HVAC flow control
Patent Information
- Application Number
- JP2022196383
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-04-21
- Filing Date
- 2022-12-08
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing control methods for dynamic systems, such as HVAC units, face challenges in accurately modeling nonlinear systems and designing optimal controllers due to the complexity of partial differential equations and the need for large amounts of data, leading to inefficient and inaccurate control strategies that do not consider the physical properties of the system.
A data-driven approach using reinforcement learning (RL) to develop a reduced-order model combined with a closure model, which captures the patterns of system dynamics, allowing for efficient and constrained control strategies that mimic the physical behavior of the system.
This method enables robust and efficient control of HVAC systems by accurately estimating system behavior and ensuring stability, even under uncertain conditions, while reducing computational complexity and incorporating constraints.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to system modeling, prediction, and control. More particularly, it relates to methods and apparatus for robust data-driven model adaptation with dynamic mode decomposition for controlling HVAC units.
Background Art
[0002] Control theory in control system engineering is a subfield of mathematics that deals with the control of dynamic systems that operate continuously in engineered processes and machines. The object of the present invention is to develop control strategies for controlling such systems in an optimal manner without delay or overshoot and to ensure control stability.
[0003] For example, optimization-based control and estimation techniques such as model predictive control (MPC) enable a model-based design framework that can directly take into account system dynamics and constraints. MPC is used in many applications for controlling dynamic systems of various complexities. Examples of such systems include production lines, automotive engines, robots, numerical control machining, motors, satellites, and generators. As used herein, a model of the dynamics of a system or a model of a system describes the dynamics of the system using differential equations. For example, the most general model of a linear system having p inputs u, q outputs y, and n state variables x is written in the following form:
Number
[0004] However, in some situations, the model of a controlled system is nonlinear, which can make it difficult to design, difficult to use in real time, or inaccurate. Examples of such cases are prevalent in robotics, building control (HVAC), smart grids, factory automation, transportation, self-regulating machinery, and transit networks. In addition, even when a nonlinear model is available accurately, designing an optimal controller is inherently a difficult task because it requires solving a partial differential equation known as the Hamilton-Jacobi-Bellman (HJB) equation.
[0005] In the absence of precise models of dynamic systems, some control methods utilize behavioral data generated by the dynamic system to construct feedback control policies that stabilize system dynamics or to embed quantifiable control-related performance. The use of behavioral data to design control policies is called data-driven control. There are two types of data-driven control methods: (i) an indirect method in which a model of the system is first constructed and then the model is used to design the controller, and (ii) a direct method in which the control policy is constructed directly from the data without an intermediate model construction step.
[0006] A drawback of indirect methods is the potential need for large amounts of data during the model-building phase. Furthermore, in indirect control methods, the controller is calculated from the estimated model, for example, according to the certainty equivalence principle; however, in practice, the model estimated from the data does not capture the physical characteristics of the system's dynamics. Therefore, some model-based control techniques cannot be used in conjunction with such data-driven models.
[0007] To overcome this problem, some methods use direct control methods, directly mapping experimental data onto the controller without identifying any model in between. However, direct control methods result in a black-box design of the control policy, which directly maps the system state to control commands. However, such control policies are not designed considering the physical characteristics of the system. Furthermore, the control designer cannot influence the data-driven decisions of the control policy.
[0008] Therefore, methods and apparatus for optimally controlling the system are still needed. [Overview of the project]
[0009] The objective of some embodiments is to provide apparatus and methods for data-driven design of system dynamics models to generate system dynamics models that capture the dynamics of system behavior. Thus, embodiments simplify the model design process while retaining the advantages of having a system model when designing control applications. However, current data-driven methods are not suitable for estimating system models that capture the physical dynamics of a system.
[0010] For example, reinforcement learning (RL) is a field of machine learning that deals with how to take actions in an environment to maximize some concept of cumulative reward (or, conversely, minimize cumulative loss / cost). Reinforcement learning relates to optimal control in a continuous state input space, and it primarily concerns the existence and characterization of optimal control policies, as well as algorithms for their computation in the absence of mathematical models of the controlled system and / or environment.
[0011] Considering the advantages offered by the RL method, some embodiments aim to develop RL techniques that yield optimal control policies for dynamic systems, which can be described using differential equations. However, control policies map the system state to control commands, and this mapping is not based on, or at least does not need to be based on, the physical dynamics of the system. Therefore, RL-based data-driven estimation of models that have physical meaning and one or more differential equations to describe the system's dynamics remains unexplored in the field of control.
[0012] Some embodiments of RL data-driven learning of system dynamics models with physical significance are based on the recognition that the reward function can be viewed as a virtual control problem, which is minimizing the difference between the system's behavior as determined by the learned model and the system's actual behavior. In particular, the system's behavior is a higher-level characterization of the system, such as system stability and boundedness of states. In fact, systems also exhibit behavior in uncontrolled situations. Unfortunately, estimating such models via RL is computationally difficult.
[0013] To this end, some embodiments are based on the understanding that a model of a system can be represented by a reduced-order model, which we call a closure model, combined with a virtual control term. For example, if a model of a system based on complete physical laws is typically captured by partial differential equations (PDEs), the reduced-order model may be represented by ordinary differential equations (ODEs). ODEs express the dynamics of a system as a function of time, but are less precise than the representation of dynamics using PDEs. Therefore, the purpose of the closure model is to reduce this gap.
[0014] As used herein, the closure model is a nonlinear function of the system's state that captures the difference in the system's behavior estimated by the ODE and PDE. Thus, the closure model is also a function of time representing the difference in dynamics between the dynamics captured by the ODE and the dynamics captured by the PDE. Some embodiments are based on the understanding that, because solving the PDE equations is computationally expensive, representing the system's dynamics as a combination of the ODE and the closure model can simplify subsequent control of the system. Therefore, some embodiments attempt to simplify data-driven estimation of the system's dynamics by representing the dynamics using the ODE and the closure model and updating only the closure model. However, this problem, while computationally simpler, is difficult when formulated within the framework of RL. This is because RL is typically used to learn control policies to precisely control a system. Here, in this analogy, RL should attempt to precisely estimate the closure model, which is difficult.
[0015] However, some embodiments are based on the recognition that, in certain modeling situations, it is sufficient to represent a pattern of behavior rather than the exact behavior of the system's dynamics itself. For example, if the exact behavior captures the system's energy at each point in time, the pattern of behavior captures the rate of change of energy. By analogy, when a system is excited, its energy increases. Knowing the exact behavior of the system's dynamics makes it possible to estimate such an energy increase. Knowing the pattern of behavior of the system's dynamics makes it possible to estimate the rate of increase in order to estimate a new energy value that is proportional to the actual energy value.
[0016] Therefore, while the pattern of the system's dynamics is not the exact behavior itself, in some model-based control applications, the pattern of the system's dynamics is sufficient to design Lyapunov stabilization control. Examples of such control applications include stabilization control, which aims to stabilize the state of a system.
[0017] To this end, some embodiments use RL to update the closure model so that the dynamics of the ODE and the updated CL mimic the pattern of the system's dynamics. Some embodiments are based on the understanding that the pattern of dynamics can be represented by the shape of the state trajectory, which is determined as a function of time, in contrast to the system's state values. The state trajectory can be measured during the system's online functioning. In addition, or alternatively, the state trajectory can be simulated using PDE.
[0018] To this end, some embodiments control the system using a system model that includes a combination of ODE and a closure model, and update the closure model with RL having a value function that reduces the difference between the actual shape of the state trajectory and the shape of the state trajectory estimated using the ODE together with the updated closure model.
[0019] However, after convergence, the ODE with the updated CL represents a pattern of the system's behavioral dynamics, but not the actual values of the behavior. In other words, the ODE with the updated CL is a function proportional to the system's actual physical dynamics. For this reason, some embodiments involve gains from closure models, which are learned later during online control of the system, in a way that is more suitable for model-based optimization than RL. Examples of these methods include extremum search and optimization based on Gaussian processes.
[0020] In addition, or alternatively, some embodiments use a model of the system determined by data-driven adaptation in various model-based predictive control systems, such as MPC. These embodiments allow us to leverage MPC's ability to consider constraints in the control of the system. For example, conventional RL methods are not suitable for data-driven control of constrained systems. This is because conventional RL methods do not consider the satisfaction of state and input constraints in a continuous state-operation space; that is, conventional RL cannot guarantee that the state of the controlled system acted upon by the control inputs satisfies state and input constraints throughout the operation.
[0021] However, some embodiments allow for learning the physical properties of a system using RL and combining the data-driven advantages of RL with model-based constrained optimization.
[0022] Accordingly, one embodiment discloses a device for controlling the operation of a system. The device comprises: an input interface configured to receive a state trajectory of the system; a memory configured to store a model of the system's dynamics, including a combination of at least one differential equation and a closure model; a processor configured to update the closure model using reinforcement learning (RL) having a value function that reduces the difference between the shape of the received state trajectory and the shape of the state trajectory estimated using the model together with the updated closure model, and to determine control commands based on the model and the updated closure model; and an output interface configured to transmit control commands to the system's actuators to control the operation of the system.
[0023] Another embodiment discloses a method for controlling the operation of a system. The method uses a processor coupled to memory storing a model of the system's dynamics, which includes a combination of at least one differential equation and a closure model, the processor coupled to stored instructions that, when executed by the processor, perform the steps of the method, the method including receiving a state trajectory of the system; updating the closure model using reinforcement learning (RL) which has a value function that reduces the difference between the shape of the received state trajectory and the shape of the state trajectory estimated using the model together with the updated closure model; determining control commands based on the model and the updated closure model; and transmitting the control commands to the system's actuators to control the operation of the system.
[0024] According to some embodiments of the present invention, a computer-implemented method is provided for controlling a heating, ventilation, and air conditioning (HVAC) system including actuators, using a reinforcement learning-trained order reduction estimator (RL-trained ROE) and a robust closure model. The method uses a processor coupled with a memory storing instructions for performing the method, and when executed by the processor, the steps of the method include: obtaining a setpoint for the HVAC system from user input via an input interface; obtaining measurement data from sensors located within the HVAC system; calculating a higher-dimensional state estimate using the measurement data and an order reduction state estimate from the RL-trained ROE; determining a controller with respect to the setpoint using the RL-trained ROE; generating a control command based on the controller; and transmitting the control command to the actuators of the HVAC system via an output interface.
[0025] Furthermore, some embodiments of the present invention provide an apparatus for controlling a heating, ventilation, and air conditioning (HVAC) system including an actuator. The apparatus may include an input interface configured to obtain set values of the HVAC system from user input and measurement data from sensors disposed in the HVAC system, at least one memory configured to store instructions for implementing a method implemented by a computer, and at least one processor coupled to the at least one memory. When the instructions are executed by the at least one processor, the steps of the method implemented by the computer are executed to calculate a high-dimensional state estimate using the measurement data and an estimated value of the reduced-order state from the RL-trained ROE, determine a controller with respect to the set value by using the RL-trained ROE, and generate a control command based on the controller. The apparatus may further include an output interface configured to transmit a control command including a control command for controlling an actuator that operates the HVAC system.
Brief Description of the Drawings
[0026] [Figure 1] FIG. 8 is a two-stage block diagram for generating a robust reduced-order model for online control in an offline manner according to an embodiment of the present invention. [Figure 2] FIG. 11 is a schematic diagram of the principle used by some embodiments to control the operation of the system. [Figure 3] FIG. 14 is a block diagram of an apparatus for controlling the operation of the system according to some embodiments of the present invention. [Figure 4] FIG. 17 is a flowchart diagram of the principle for controlling the system according to some embodiments of the present invention. <This is a schematic diagram of a reinforcement learning (RL)-based degree reduction model according to several embodiments of the present invention. [Figure 6B] This is a flowchart illustrating the operation of updating a closure model using RL according to one embodiment of the present invention. [Figure 7] This figure shows the difference between the actual behavior and the estimated behavior of the system according to several embodiments of the present invention. [Figure 8A] This is a schematic diagram of a training algorithm for learning the optimal policy to be used in a closure model, according to one embodiment of the present invention. [Figure 8B] This is a schematic diagram of a training algorithm for learning the optimal policy to be used in a closure model, according to one embodiment of the present invention. [Figure 8C] This is a schematic diagram of a training algorithm for learning the optimal policy to be used in a closure model, according to one embodiment of the present invention. [Figure 9A] This is a schematic diagram of a control algorithm based on a robust order reduction model according to several embodiments of the present invention. [Figure 9B] This is a schematic diagram of a control algorithm based on a robust order reduction model according to several embodiments of the present invention. [Figure 9C] This is a schematic diagram of a control algorithm based on a robust order reduction model according to several embodiments of the present invention. [Figure 10] This figure shows an exemplary real-time implementation example of a device for controlling an air conditioning system, according to an embodiment of the present invention. [Modes for carrying out the invention]
[0027] The accompanying drawings are included for a further understanding of the present invention, illustrating embodiments of the invention and, together with this description, illustrating the principles of the invention. The drawings shown are not necessarily to scale and are generally focused on illustrating the principles of the embodiments of this disclosure.
[0028] The drawings identified above illustrate embodiments disclosed herein, but other embodiments are contemplated, as will be discussed. This disclosure presents exemplary embodiments, not limiting ones. Those skilled in the art can devise numerous other modifications and embodiments that fall within the scope and spirit of the principles of the embodiments of this disclosure.
[0029] In the following description, for illustrative purposes and to facilitate a full understanding of the disclosure, numerous specific details are provided. However, it will be apparent to those skilled in the art that the disclosure may be implemented without these specific details. In other examples, the apparatus and methods are shown only in block diagram form to avoid obscuring the disclosure.
[0030] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the disclosure. Rather, the following description of exemplary embodiments provides a practicable description for realizing one or more exemplary embodiments. The intent is to describe various modifications that may be made to the function and configuration of the elements without departing from the spirit and scope of the subject matter disclosed as described in the claims.
[0031] The following description provides specific details for a complete understanding of the embodiments. However, it will be understood by those skilled in the art that embodiments may be carried out without these specific details. For example, systems, processes, and other elements in the disclosed subject matter may be shown as components in the form of block diagrams so as not to obscure the embodiments with unnecessary details. In other examples, well-known processes, structures, and techniques may be shown without unnecessary details to avoid obscuring the embodiments. Furthermore, similar reference numbers and names in different drawings indicate similar elements.
[0032] In the following description, for illustrative purposes and to facilitate a full understanding of the disclosure, numerous specific details are provided. However, it will be apparent to those skilled in the art that the disclosure may be implemented without these specific details. In other examples, the apparatus and methods are shown only in block diagram form to avoid obscuring the disclosure.
[0033] Where used herein and in the claims, the words “for example,” “as an example,” and “etc.,” and the verbs “equip,” “have,” and “include,” and their other verb forms, when used with a list of one or more components or other items, should each be interpreted as open-ended, meaning that the list should not be considered to exclude any other additional components or items. The phrase “based on” means based at least in part. Furthermore, it should be understood that the expressions and terms used herein are for illustrative purposes only and should not be considered limiting. Any headings used within this description are for convenience only and have no legal or limiting effect.
[0034] In describing embodiments of the present invention, the following definitions are applicable throughout this disclosure.
[0035] A “control system” or “controller” may refer to a device or set of devices for managing, commanding, instructing, or coordinating the behavior of other devices or systems. A control system can be implemented by either software or hardware and may include one or more modules. A control system including a feedback loop may be implemented using a microprocessor. A control system may be an embedded system.
[0036] A “heating, ventilation, and air conditioning (HVAC) system” can refer to a system that uses a vapor compression cycle to move a refrigerant through its components based on the principles of thermodynamics, fluid dynamics, and / or heat transfer. HVAC systems encompass a very wide range of systems, from those that supply only outdoor air to building occupants to those that control only the building's temperature, and those that control both temperature and humidity.
[0037] The term "central processing unit (CPU)" or "processor" can refer to a computer or a component of a computer that reads and executes software instructions. Furthermore, a processor can be "at least one processor" or "one or more processors."
[0038] Figure 1 shows a schematic block diagram illustrating how large-scale systems, such as those resulting after the discretization of partial differential equations (PDEs), can be controlled and estimated using a two-stage apparatus.
[0039] In step 1, shown in 106, an offline task is performed to derive a robust reduced-order model (ROM). Data for the development of such a model may be generated by high-fidelity computational fluid dynamics (CFD) simulations or by conducting experiments.
[0040] CFD is a branch of fluid dynamics that uses numerical analysis and data structures to analyze and solve problems involving fluid flow. Computers are used to perform the calculations necessary to simulate the free flow of fluids and their interaction with surfaces defined by boundary conditions (liquids and gases). Ongoing research has led to the development of software to improve the accuracy and speed of complex simulation scenarios, such as transonic or turbulent flow, or to describe airflow in HVAC applications. Initial validation of such software is typically performed using experimental equipment such as wind tunnels. In addition, previously performed analytical or empirical analyses of a particular problem can be used for comparison.
[0041] Next, a ROM (Range of Motion) is developed using a dataset generated by either a CFD simulation or experiment, which may only be valid for trajectories obtained by the CFD. For example, the CFD in step 101 can be performed on a room with a closed window, and the ROM 102 is valid only for this condition. If the window is opened, the accuracy of the ROM 102 may degrade, becoming unstable or highly inaccurate. In this case, 1033 is trained to be used for estimation and control, using several trajectories generated by the CFD simulation or experiment of 101. All such tasks are performed offline. The model 1033 (102+103), generated by the offline step 106 and trained on the difference between the prediction of the corrected ROM 102 and the training data 105 by RL, is robust to parameter variations and can also handle unknown initial conditions.
[0042] The uncertainty in experimentation or CFD simulation 101, as described in 102, can be addressed by developing a robust ROM in 103.
[0043] A major challenge is that the ROM provides a simplified, incomplete description of the dynamics, which negatively impacts the performance of the state estimator used in online control. One potential solution is to improve the accuracy of the ROM itself by including additional closure terms with further detail, as shown in Figure 5.
[0044] Some embodiments attempt to develop a more robust ROM by various methods, for example, by using various trajectories and averaging, by using sensitivity analysis, by using problem-specific, prior known basis functions, etc.
[0045] Some embodiments develop the ROM based solely on a given trajectory and, instead of further developing the ROM, propose additional terms called closure models to improve the accuracy of the estimation. For example, a Lyapunov-based closure model, a physical law-inspired closure model (e.g., using artificial diffusion), or a reinforcement learning method can be used to develop a model of the closure term.
[0046] Some embodiments use conventional methods such as Kalman filtering to add an estimation layer to the ROM. In terms of statistics and control theory, Kalman filtering, also known as linear quadratic estimation (LQE), is an algorithm that uses a series of measurements observed over time, including statistical measurements and noise modeling, to generate estimates of the unmeasured states of a system. These estimates are more accurate than estimates based on single measurements alone by estimating the joint probability distribution across the states in each time frame.
[0047] Some embodiments utilize a reinforcement learning reduced-order estimator (RL-ROE), which can then be used for online control. The RL-ROE is constructed from ROM in a manner similar to a Kalman filter, but with a key difference: the linear filter gain function is replaced by a nonlinear stochastic policy trained through reinforcement learning (RL). The flexibility of the nonlinear policy allows the RL-ROE to compensate for ROM errors due to, for example, incomplete knowledge of dynamics.
[0048] Some embodiments describe the estimation problem as a quiescent Markov decision process (MDP) to enable RL training using the RL method for quiescent Markov decision processes (MDPs). A Markov process is a stochastic process in which, given the current situation, the future is independent of the past. Thus, Markov processes are a natural stochastic analogue of deterministic processes described by differential and difference equations. They form one of the most important classes of stochastic processes.
[0049] Several embodiments demonstrate that the trained RL-ROE outperforms a Kalman filter designed using the same ROM and exhibits robust estimation performance with respect to different reference trajectories and initial state estimates. The proposed RL-ROE is the first application of reinforcement learning to state estimation for high-dimensional systems. Further details relating thereto are given with reference to Figures 6 and 8.
[0050] Once the ROM and closure models are constructed, the resulting models can be used, first for estimation and finally for online control. For example, a robust model 108 generated using some CFD or experimental trajectory 101 may be developed for a particular room layout (e.g., rectangular, L-shaped) using some conditions for windows (e.g., open, closed, half-open) or a given number of people in the room. However, in reality, the number of people in a room may vary, and the windows may be quarter-open for layouts that are neither rectangular nor L-shaped, but a combination of the two. The closure model learned in the offline stage 106 is configured to estimate room conditions, e.g., temperature or velocity in the room, even in such unknown cases that fall within a similar trajectory generated by 101. This can be done if sensor data 109 representing partially accurate knowledge of the physical properties of the room and the HVAC installed therein is supplied to 108. Such a process is also known as data assimilation, i.e., assimilating the information itself from the sensing with possibly inaccurate model information.
[0051] Data assimilation is a mathematical field that attempts to best combine predictions (usually in the form of numerical models) with observations. For example, there may be several different objectives sought: determining the optimal state estimate of a system, determining the initial conditions of a numerical prediction model, interpolating sparse observational data using knowledge of the observed system, or identifying numerical parameters of a model from observed experimental data. Depending on the objective, different solution methods may be used. Data assimilation is distinguished from other forms of machine learning and statistical methods in that it utilizes a dynamic model of the system being analyzed. The process (process step) 110 of reconstructing room temperature and velocity is the result of such data assimilation of a robust model 108 and sensor data 109.
[0052] The offline stages 106 and online stages 107 are examples of developing a simplified, robust model 108, which can then be used for estimation and control.
[0053] Estimation theory is a branch of statistics that deals with estimating the values of parameters based on measured empirical data that have random components. Parameters describe the underlying physical setting in such a way that their values affect the distribution of the measured data. Estimators attempt to approximate unknown parameters using measured values. In estimation theory, two approaches are generally considered: the probabilistic approach (as described in this invention) assumes that the measured data are random and that the probability distribution depends on the parameter in question, and the set membership approach assumes that the measured data vector belongs to a set on which the parameter vector depends.
[0054] Examples of sensory data installed in rooms for HVAC applications include thermocouple readings, thermal camera measurements, velocity sensors, and humidity sensors.
[0055] Once the room temperature or velocity is re-established at 110, an online control step 107 may be performed for room airflow control 111. Further details are shown in Figure 9.
[0056] Figure 2 shows a schematic diagram of the principles used by several embodiments to control the operation of the system. Some embodiments provide a control device 200 configured to control a system 202. For example, the device 200 can be configured to control a continuously operating dynamic system 202 in engineering processes and machines. Hereinafter, “control device” and “device” may be used interchangeably and have the same meaning. Hereinafter, “continuously operating dynamic system” and “system” may be used interchangeably and have the same meaning. Examples of system 102 include HVAC systems, LIDAR systems, condensing units, production lines, self-regulating machines, smart grids, automobile engines, robots, numerically controlled machining, motors, satellites, generators, and transportation networks. Some embodiments are based on the understanding that the device 200 develops a control policy 206 configured to provide estimations and controls (commands) in order to control the system 202 using control actions in an optimal manner without delay or overshoot, and to ensure control stability.
[0057] In some embodiments, the device 200 develops control commands 206 for system 202 using model-based and / or optimization-based control and estimation techniques, such as model predictive control (MPC). Model-based techniques can be advantageous for controlling dynamic systems. For example, MPC enables a model-based design framework in which the dynamics and constraints of system 202 can be directly considered. MPC develops control commands 206 based on a model of system 204. Model 204 of system 202 refers to the dynamics of system 202 described using differential equations. In some embodiments, model 204 may be nonlinear, difficult to design, and / or difficult to use in real time. For example, even if a nonlinear model is precisely available, estimating the optimal control commands 206 is an inherently difficult task because it requires solving a partial differential equation (PDE) describing the dynamics of system 202, known as the Hamilton-Jacobi-Bellman (HJB) equation, which is computationally difficult.
[0058] Some embodiments use data-driven control techniques to design Model 204. Data-driven techniques utilize behavioral data generated by System 202 to construct a feedback control policy that stabilizes System 202. For example, each state of System 202 measured during its operation may be given as feedback for controlling System 202. Generally, the use of behavioral data to design a control policy and / or command 206 is called data-driven control. The objective of data-driven control is to design a control policy from data and to control the system using the data-driven control policy. In contrast to such a data-driven control approach, some embodiments use behavioral data to design a model of the control system, e.g., Model 204, and then use the data-driven model to control the system using various model-based control methods. It should be noted that the objective of some embodiments is to determine a real model of the system from the data, i.e., such a model that can be used to estimate the behavior of the system. For example, the objective of some embodiments is to determine a model of the system from data that captures the system's dynamics using differential equations. In addition, or alternatively, an objective of some embodiments is to learn a model from data that has PDE model accuracy based on physical laws.
[0059] To simplify calculations, some embodiments formulate an ordinary differential equation (ODE) 208a to describe the dynamics of system 202. In some embodiments, ODE 208a may be formulated using model degeneracy techniques. For example, ODE 208a may be a reduced-dimensional PDE. To that end, ODE 208a can be part of the PDE. However, in some embodiments, ODE 108a cannot reproduce the actual dynamics of system 202 (i.e., the dynamics described by the PDE) under uncertainty conditions. Examples of uncertainty conditions may be that the boundary conditions of the PDE change over time, or that one of the coefficients involved in the PDE changes.
[0060] To this end, several embodiments provide a reduced-order estimator (ROE) 208 that includes a ROM(DMD) 208a and a robust RL-based closure model 208b that reduces the PDE, while covering the case of uncertainty conditions. In some embodiments, the closure model 208b may be a nonlinear function of the state of system 202 that captures differences in the behavior (e.g., dynamics) of system 202 according to the ODE and PDE. The closure model 208b may be formulated using reinforcement learning (RL). In other words, the PDE model of system 202 is approximated by a combination of the ODE(ROM) 208a and the closure model 208b, and the closure model 208b is learned from the data using RL. In this way, a model that approaches the accuracy of the PDE is learned from the data.
[0061] In some embodiments, RL learns a state trajectory of system 202 that defines the behavior of system 202, rather than learning the individual states of system 202. The state trajectory may be a sequence of states of system 202. In some embodiments, model 208, comprising ODE 208a and closure model 208b, is based on the understanding that it reproduces a pattern of system 202 behavior, rather than the actual behavioral values (e.g., states) of system 202. The pattern of system 202 behavior can be expressed as a function of time, such as the shape of the state trajectory, e.g., a set of states of the system. The pattern of system 202 behavior may also represent higher-level characteristics of the model, e.g., the boundedness of its solution over time or the decay of its solution over time, but does not optimally reproduce the dynamics of the system.
[0062] To this end, some embodiments include the gains in the closure model 208b to determine the gains and best reproduce the dynamics of system 202. In some embodiments, the gains may be updated using an optimization algorithm. Model 208, including the closure model 108b with the updated gains, reproduces the dynamics of system 202. Thus, model 208 best reproduces the dynamics of system 202. Some embodiments are based on the understanding that model 208 contains fewer parameters than PDE. For this reason, model 208 is not as computationally complex as PDE, which describes a physical model of system 202. In some embodiments, the control policy 206 is determined using model 208. The control policy 206 controls the operation of system 202 by directly mapping the state of system 202 to control commands. Thus, the degenerate model 108 is used to design control for system 202 in an efficient manner.
[0063] Figure 3 shows a block diagram of a device 1200 for controlling the operation of system 1202, according to several embodiments. The device 1200 includes an input interface 1202 and an output interface 1218 for connecting the device 1200 to other systems and devices. In some embodiments, the device 1200 may include multiple input interfaces and multiple output interfaces. The input interface 1202 is configured to receive the status trajectory 1216 of system 202. The input interface 1202 includes a network interface controller (NIC) 1212 adapted to connect the device 1200 to a network 1214 via a bus 1210. Through the network 1214, either wirelessly or wired, the device 1200 receives the status trajectory 1216 of system 1202.
[0064] The state trajectory 1216 may be a set of states of system 202 that define the actual behavior of the system's dynamics. For example, the state trajectory 1216 acts as a reference continuous state space for controlling system 202. In some embodiments, the state trajectory 1216 may be received from real-time measurements of a portion of the system's states. In some other embodiments, the state trajectory 1216 may be simulated using a PDE that describes the dynamics of system 202. In some embodiments, the shape of the received state trajectory may be determined as a function of time. The shape of the state trajectory may represent the actual pattern of the system's behavior.
[0065] The device 1200 further includes a processor 1204 and a memory 1206 that stores instructions executable by the processor 1204. The processor 1204 may be a single-core processor, a multi-core processor, a computing cluster, or any number of other configurations. The memory 1206 may include random-access memory (RAM), read-only memory (ROM), flash memory, or any other suitable memory system. The processor 1204 is connected to one or more input and output devices via a bus 1210. The stored instructions implement a method for controlling the operation of the system 202.
[0066] Memory 1206 may be further extended to include storage 1208. Storage 1208 may be configured to store model 1208a, controller 1208b, update module 1208c, and control command module 1208d. In some embodiments, model 1208a may be a model describing the dynamics of system 202, including a combination of at least one differential equation and a closure model. The differential equation of model 1208 may be an ordinary differential equation (ODE) 208a. The closure model of model 208a may be a linear or nonlinear function of the state of system 202. The closure model may be learned using RL to mimic the behavior of system 202. As understood, once the closure model is learned, the closure model may become closure 208b as illustrated in Figure 1.
[0067] The controller 1208b may be configured to store instructions that, when executed by the processor 1204, will execute one or more modules in the storage 1208. Some embodiments are based on the understanding that the controller 1208b controls the system 202 by managing each module of the storage 1208.
[0068] The update module 1208c may be configured to update the closure model of model 1208a using reinforcement learning (RL) which has a value function that reduces the difference between the shape of the received state trajectory and the shape of the state trajectory estimated using model 1208a together with the updated closure model. In some embodiments, the update module 1208c may be configured to iteratively update the closure module using RL until a termination condition is met. The updated closure model is a nonlinear function of the system's states that captures the differences in the ODE and PDE and therefore the system's behavior.
[0069] Furthermore, in some embodiments, the update module 1208c may be configured to update the gains of the updated closure model. To this end, some embodiments determine a gain that reduces the error between the state of system 202 estimated using model 1208a having the updated closure model with the updated gains and the actual state of the system. In some embodiments, the actual state of the system may be a measured state. In some other embodiments, the actual state of the system may be a state estimated using a PDE that describes the dynamics of system 202. In some embodiments, the update module 1208c may update the gains using extreme value search. In some other embodiments, the update module 1208c may update the gains using optimization based on a Gaussian process.
[0070] The control command module 1208c may be configured to determine control commands based on model 1208a and the updated closure model. The control commands may control the operation of the system. In some embodiments, the operation of the system may be constrained. To this end, the control command module 1208c uses predictive model-based control to determine control commands while enforcing constraints. The constraints include state constraints in the continuous state space of system 202 and control input constraints in the continuous control input space of system 202.
[0071] The output interface 1218 is configured to send control commands to the actuator 1220 of the system 202 to control the operation of the system. Some examples of the output interface 1218 may include a control interface that submits control commands to control the system 202.
[0072] Figure 4 shows flowcharts of the principles for controlling system 202 according to several embodiments. Some embodiments are based on the understanding that system 202 can be modeled from physical laws. For example, the dynamics of system 202 can be expressed by mathematical equations using physical laws. In step 402, system 202 may be represented by a higher-dimensional model based on physical laws. The higher-dimensional model based on physical laws may be a partial differential equation (PDE) that describes the dynamics of system 402. For illustrative purposes, system 202 is considered to be an HVAC system, and its model is represented by the Boussinescu equations. The Boussinescu equations are derived from physical laws and describe the coupling between airflow and temperature in a room. Thus, the HVAC system model can be mathematically expressed as follows:
number
number
[0073]
number
[0074] In some embodiments, such abstract dynamics are obtained from the numerical discretization of a nonlinear partial differential equation (PDE), which typically requires a large number of n state dimensions.
[0075] Some embodiments are based on the recognition that a high-dimensional model of system 202 based on the physical laws needs to be solved in order to control the operation of system 202 in real time. For example, in the case of an HVAC system, the Boussinescu equations need to be solved in order to control the airflow dynamics and temperature in the room. Some embodiments are based on the recognition that a high-dimensional model of system 202 based on the physical laws involves a large number of equations and variables that are complex to solve. For example, solving a high-dimensional model based on physical laws in real time requires greater computational power. For this reason, the objective of some embodiments is to simplify the high-dimensional model based on physical laws.
[0076] In step 404, the apparatus 1200 is provided to generate a reduced-order model to reproduce the dynamics of system 202 so that the apparatus 1200 controls system 202 in an efficient manner. In some embodiments, the apparatus 1200 may use model degeneracy techniques to simplify a high-dimensional model based on physical laws to generate a reduced-order model. Some embodiments are based on the understanding that model degeneracy techniques reduce the dimensionality (e.g., variables of the PDE) of a high-dimensional model based on physical laws, and that the reduced-order model may be used in real time for predicting and controlling system 202. Furthermore, the generation of a reduced-order model for controlling system 202 will be described in detail with reference to Figure 5. In step 406, the apparatus 1200 uses the reduced-order model in real time to predict and control system 202.
[0077] Figure 5 shows schematic architectures for generating a reduced-order model according to several embodiments. Some embodiments are based on the understanding that the device 1200 generates a reduced-order model (ROM) 506 using model degeneracy techniques. The ROM 506 generated using model degeneracy techniques may be a part 502 of a higher-dimensional model based on physical laws. The part 502 of the higher-dimensional model based on physical laws may be one or more differential equations describing the dynamics of system 202. The part 502 of the higher-dimensional model based on physical laws may be an ordinary differential equation (ODE). In some embodiments, the ODE cannot reproduce the actual dynamics (i.e., the dynamics described by the PDE) under uncertainty conditions. Examples of uncertainty conditions may be that the boundary conditions of the PDE change over time, or that one of the coefficients included in the PDE changes. These mathematical changes actually reflect some actual changes in the actual dynamics. For example, in an HVAC system, opening and closing the windows and / or doors of a room changes the boundary conditions of the Boussinescu equation (i.e., PDE). Similarly, weather changes, such as daily and seasonal variations, affect the difference between indoor and outdoor temperatures, which in turn can affect some of the PDE coefficients, and thus, for example, the Reynolds number.
[0078] In all of these scenarios, model degeneracy techniques cannot have an integrated approach to obtain a reduced-order (or reduced-dimensionality) model of the system dynamics 506 that covers all of the above scenarios, namely parameter uncertainty and boundary condition uncertainty.
[0079] The objective of some embodiments is to generate a ROM506 that solves the PDE in the case of changes in boundary conditions and / or parameters. To this end, some embodiments use adaptive model degeneracy methods, regime detection methods, etc.
[0080]
number
[0081]
number
[0082]
number
[0083] As another example, in one embodiment of the present invention, the order reduction 506 takes the following quadratic form:
number
number
[0084] However, solutions to the ROM equations can lead to unstable solutions (diverging beyond the finite-time support), and these unstable solutions do not reproduce the physics of the original PDE model with its viscous term that always stabilizes the solution (i.e., being bounded to the bounded-time support). For example, ODE may lose the intrinsic properties of the actual solutions of the physical laws-based higher-dimensional model during model degeneracy. As a result, ODE may lose the boundedness of the actual solutions of the physical laws-based higher-dimensional model in space and time.
[0085] Therefore, some embodiments modify the ROM 506 by adding a closure model 504 that represents the difference between the ODE and the PDE. For example, the closure model 504 captures the lost intrinsic properties of the actual solution of the PDE and acts like a stabilizing factor. Some embodiments allow updating only the closure model 506 to reduce the difference between the ODE and the PDE.
[0086] For example, in some embodiments, ROM406 can be mathematically represented as follows:
number
[0087] Function F is the closure model 504, which is added to stabilize the solution of the ROM model 506.
number
[0088] Figure 6A shows schematic diagrams of several embodiments of a reinforcement learning (RL)-based reduction-of-order model 506. In some embodiments, an RL-based data-driven method may be used to compute an RL-based closure model 602. Some embodiments are based on the understanding that the closure model 502 is iteratively updated with RL to compute the RL-based closure model 602. The RL-based closure model 602 may be the optimal closure model. Furthermore, an iterative process for updating the closure model 504 is described in detail with reference to Figure 6B. Some embodiments are based on the understanding that the optimal closure model combined with ODE may form the optimal ROM 506. In some embodiments, the ROM 506 may estimate actual patterns of the system 202's behavior. For example, the ROM 506 mimics the shape of the received state trajectory.
[0089]
number
[0090]
number
[0091] Figure 6B shows a flowchart of the operation for updating the closure model 602 using RL according to an embodiment of the present invention. In step 604, the device 1200 may be configured to initialize an initial closure model policy and a learning cumulative reward function associated with the initial closure model policy. The initial closure model policy may be a simple linear closure model policy. The cumulative reward function may be a value function. In step 606, the device 1200 is configured to run ROM 606 containing a portion of a higher-dimensional model 502 based on physical laws and the current closure model (e.g., the initial closure model policy) to collect data along a finite time interval. To this end, the device 1200 collects data representing a pattern of behavior of the system 202's dynamics. For example, the behavior pattern captures the rate of change of the system 202's energy over a finite time interval. Some embodiments are based on the understanding that the pattern of behavior of the system 202's dynamics can be represented by the shape of the state trajectory over a finite time interval.
[0092] In step 608, the device 1200 is configured to update the cumulative reward function using the collected data. In some embodiments, the device 1200 updates the cumulative reward function (i.e., the value function) to show the difference between the shape of the received state trajectory and the shape of the state trajectory estimated using the ROM 506 together with the current closure model (e.g., the initialized closure model).
[0093] Some embodiments are based on the understanding that RL uses a neural network trained to minimize a value function. To that end, in step 610, the device 1200 is configured to update the current closure model policy using the collected data and / or updated cumulative reward function so that the value function is minimized.
[0094] In some embodiments, the device 1200 is configured to repeat steps 606, 608, and 610 until a termination condition is met. To this end, in step 612, the device 1200 is configured to determine whether the learning has converged. For example, the device 1200 determines whether the learning cumulative reward function is below a threshold limit, or whether two consecutive learning cumulative reward functions are within a small threshold limit. If the learning has converged, the device 1200 proceeds to step 616; otherwise, the device 1200 proceeds to step 614. In step 614, the device 1200 is configured to replace the closure model with an updated closure model, and the update procedure is repeated until a termination condition is met. In some embodiments, the device 1200 repeats the update procedure until the learning converges. In step 614, the device 1200 is configured to stop closure model learning and use the last updated closure model policy as the optimal closure model for the ROM 506.
[0095] Figure 7 shows the difference between the actual behavior and the estimated behavior of system 202 in several embodiments. In some embodiments, the behavior pattern of system 202 may be represented by two-dimensional axes, where the x-axis corresponds to time and the y-axis corresponds to the energy magnitude of system 202. Wave 702 may represent the actual behavior of system 202. Wave 704 may represent the estimated behavior of system 202. Some embodiments are based on the recognition that there may be a quantitative gap 706 between the actual behavior 702 and the estimated behavior 704. For example, the actual behavior 702 and the estimated behavior 704 may have similar frequencies but different amplitudes.
[0096] To this end, the objective of some embodiments is to include a policy parameter θ in the optimal closure model so that the gap 706 between the actual behavior 702 and the estimated behavior 704 is reduced. Furthermore, how the apparatus 1200 determines the policy parameter θ to reduce the gap 706 will be described in detail with reference to Figures 8A, 8B, and 8C.
[0097] Figures 8A to 8C show schematic diagrams of training algorithms for tuning the optimal closure model according to one embodiment of the present invention. Some embodiments are based on the understanding that the ROM 506 with the ODE 502 and the optimal closure model (i.e., the optimal ROM 506) may be useful for short time intervals. In other words, the optimal ROM 506 forces the behavior of the system 202 to be bounded only for small time intervals. To this end, the objective of some embodiments is to tune the policy parameter θ (also called a coefficient) of the optimal ROM 506 over time.
[0098]
number
[0099]
number
[0100]
number
[0101]
number
[0102]
number
[0103]
number
[0104]
number
[0105] Figures 9A to 9C show schematic diagrams of a control algorithm for using a robust ROM 616 trained in an offline stage 106 for use in online control of system 202, according to one embodiment of the present invention. Sensor data 109 is used for data assimilation and incorporated with an RL-based closure model to update the ROM. Once the model is available, it can be used for online control. Examples of control u are operations related to HVAC performance, such as compressor speed, fan speed, blade yaw angle, and temperature and speed at the HVAC outlet.
[0106]
number
[0107] Figure 9A shows a Lyapunov-based control used in combination with the robust ROM 616. Such a model is computationally far less demanding than the full-order model 101, making online control 107 feasible. In control theory, the control-Lyapunov function is an extension of the concept of the Lyapunov function V(x) to a system with a control input. The ordinary Lyapunov function is used to test whether a dynamic system is stable; that is, whether a system starting in state x≠0 in a region D will remain in D or eventually return to x=0 due to asymptotic stability. The control-Lyapunov function is used to test whether a system is stabilizable, that is, whether for any state x there exists a control u(x,t) such that the system can be brought to a zero state by applying the control u.
[0108] Figure 9B illustrates robust control used in conjunction with the robust ROM616. In control theory, robust control is an approach to controller design that explicitly addresses uncertainty. A robust control method is designed to function properly under the condition that uncertain parameters or disturbances are found within some (typically compact) set. A robust method aims to achieve robust performance and / or stability in the presence of bounded modeling errors. In contrast to adaptive control policies, robust control policies are static and, rather than adapting to measured fluctuations, the controller is designed to operate under the assumption that some variables are unknown but bounded. The controller may perform some calculations based on one or more sensor measurements to calculate values for one or more actuators in the vapor compression cycle so that a desired performance objective is met. In some cases, the vapor compression cycle (system) of an HVAC system is connected to a controller or optimizer that adjusts actuators such as compressor speed, valve settings, or fan speed to achieve desired operating performance. The controller may obtain information about the vapor compression cycle via sensors that may be installed on or near the vapor compression cycle to measure the state of the vapor compression cycle or its environment, including several thermal-fluid characteristic variables. Examples of such sensors are temperature sensors or pressure sensors. When an actuator in the HVAC system receives a control command, including instructions, via an output interface, the control command controls the operation of the actuator in the HVAC system vapor compression cycle, which has variable-position actuators such as a variable-speed compressor or fan.
[0109] Robust controllers 902 can account for uncertainties not addressed by RL-ROE 616 using various solution methods. In some embodiments, high-gain feedback control is used so that the effects of any parameter fluctuations can be ignored. In terms of closed-loop transfer function, high open-loop gain leads to substantial disturbance rejection in the face of system parameter uncertainties. In some other embodiments, sliding-mode control is used for robust control 902. Sliding-mode control (SMC) modifies the dynamics provided by 616 by applying discontinuous control signals (or more precisely, setpoint control signals) that cause the system to "slide" along a cross-section of the system's normal behavior. SMC is a special class of nonlinear control systems that are less sensitive to plant parameter fluctuations and disturbances of 616.
[0110] Figure 9C shows MPC control used in combination with a robust ROM616. Model predictive control (MPC) is an advanced method of process control used to control a process while satisfying a set of constraints. A model predictive controller relies on a dynamic model of the process, which in this case can be provided by a robust ROM616. The main advantage of MPC is that it can optimize the current time slot while taking future time slots into account. This is different from a linear-secondary regulator (LQR), which optimizes a finite time horizon but achieves this only by realizing the current time slot and then iteratively optimizing it again. MPC also has the ability to anticipate future events and take control actions accordingly. PID controllers do not have this predictive capability. While MPC is implemented almost universally as digital control, there is research into achieving faster response times with specially designed analog circuits.
[0111] The MPC902 is configured to determine the optimal temperature and speed setpoints designed for thermal comfort by using a predictive model (provided by RL-ROE616) to predict the temperature, speed, and humidity of the building zones during each of the multiple time steps in the optimization period, generating a cost function that takes into account the cost of operating HVAC equipment during each of the multiple time steps, optimizing the cost function under constraints on the predicted temperature, speed, and humidity of the building zones to determine the optimal temperature and speed setpoints for each of the multiple time steps.
[0112] Figure 10 shows an exemplary real-time implementation of a control device 1200 for controlling a system 202, which is an air conditioning system. In this example, room 1300 has a door 1302 and at least one window 1304. The temperature and airflow in room 1300 are controlled by the device 1200 through the air conditioning system 202 and through a ventilation unit 1306. A set of sensors 1308, such as at least one airflow sensor 1308a for measuring the velocity of the airflow at a given point in room 1300 and at least one temperature sensor 1308b for measuring the room temperature, is placed in room 1300. Other types of settings can be considered, for example, a room with multiple HVAC units or a house with multiple rooms.
[0113] Some embodiments are based on the understanding that the air conditioning system 202 can be described by a model based on physical laws, known as the Boussinesque equations, as illustrated in Figure 4. However, the Boussinesque equations involve infinite dimensions in order to solve them in order to control the air conditioning system 202. For this reason, a model including the ODE502 and an updated closure model with updated gains is formulated as described in the detailed description in Figures 1 to 9. This model reproduces the dynamics of the air conditioning system 202 (e.g., airflow dynamics) in an optimal manner. Furthermore, in some embodiments, the model of airflow dynamics correlates the value of the airflow during operation of the air conditioning system 202 (e.g., airflow velocity) with the temperature of the air-conditioned room 1300. To this end, the device 11200 optimally controls the air conditioning system 202 to generate the airflow in a conditioned manner.
[0114] The above description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the following description of exemplary embodiments provides a practical description for realizing one or more exemplary embodiments. The intent is on various modifications that may be made to the function and configuration of the elements without departing from the spirit and scope of the subject matter disclosed as described in the claims.
[0115] The following description provides specific details for a complete understanding of the embodiments. However, it will be understood by those skilled in the art that embodiments may be carried out without these specific details. For example, systems, processes, and other elements in the disclosed subject matter may be shown as components in the form of block diagrams so as not to obscure the embodiments with unnecessary details. In other examples, well-known processes, structures, and techniques may be shown without unnecessary details to avoid obscuring the embodiments. Furthermore, similar reference numbers and names in different drawings indicate similar elements.
[0116] Furthermore, individual embodiments may be described as processes shown as flowcharts, flow diagrams, data flow diagrams, structural diagrams, or block diagrams. While flowcharts can describe operations as sequential processes, many operations can be performed in parallel or simultaneously. In addition, the order of operations may be rearranged. A process may terminate when its operations are complete, but it may have additional steps that are not discussed or included in the diagrams. Moreover, not all operations in any particular process described may occur in all embodiments. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, the termination of the function may correspond to the function's return to the calling function or the main function.
[0117] Furthermore, embodiments of the disclosed subject matter may be implemented at least partially manually or automatically. Manual or automatic implementations may be performed, or at least assisted, through the use of a machine, hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof. When implemented in software, firmware, middleware, or microcode, program code or code segments for performing the required tasks may be stored in a machine-readable medium. The required tasks may be performed by a processor.
[0118] The various methods or processes outlined herein may be coded as software executable on one or more processors using any one of various operating systems or platforms. In addition, such software may be written using any of several preferred programming languages and / or programming or scripting tools, and may be compiled as executable machine language code or intermediate code that runs on a framework or virtual machine. Typically, the functionality of program modules may be combined or distributed as desired in various embodiments.
[0119] Each of the above embodiments is described as a process shown as a flowchart, flow diagram, data flow diagram, structure diagram, or block diagram. While flowcharts show operations as sequential processes, many operations can be performed in parallel or simultaneously. In addition, the order of operations may be rearranged. A process may terminate when its operations are complete, but it may have additional steps that are not discussed or included in the diagrams. Furthermore, not all operations in any particular process described may occur in all embodiments. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, the termination of the function may correspond to the function's return to the calling function or the main function.
[0120] Furthermore, embodiments of the disclosed subject matter may be implemented at least partially manually or automatically. Manual or automatic implementations may be performed, or at least assisted, through the use of a machine, hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof. When implemented in software, firmware, middleware, or microcode, program code or code segments for performing the required tasks may be stored in a machine-readable medium. The required tasks may be performed by a processor.
Claims
1. A method implemented by a computer for controlling a heating, ventilation, and air conditioning (HVAC) system including an actuator, using a reduced-order estimator trained with reinforcement learning (RL-trained ROE) and a closure model, the method using a processor coupled to a memory storing instructions for implementing the method, the instructions, when executed by the processor, performing steps of the method, the steps being obtaining setpoint values of the HVAC system from user input and measurement data from sensors disposed in the HVAC system via an input interface; calculating a high-dimensional state estimate using the measurement data and an estimate of the reduced-order state from the RL-trained ROE; determining a controller with respect to the setpoint values by using the RL-trained ROE; generating a control command corresponding to the calculated high-dimensional state estimate based on the controller; transmitting, via an output interface, the control command including instructions for controlling the operation of the actuator of the HVAC system.
2. The method of claim 1, wherein the controller is designed using model predictive control.
3. The method of claim 1, wherein the controller is designed using Lyapunov design.
4. The method of claim 1, wherein the controller is designed using robust control that takes into account any model uncertainty in the RL-trained ROE.
5. The method of claim 1, wherein the RL-trained ROE is trained using a proximal policy optimization (PPO) algorithm.
6. The method of claim 1, wherein the RL-trained ROE is trained using a trust region policy optimization (TRPO) algorithm.
7. The method of claim 1, wherein the RL-trained ROE is trained using a robust constrained Markov decision process (RCMP) algorithm.
8. The method of claim 1, wherein the RL-trained ROE is trained using a time-varying non-stationary MDP.
9. An apparatus for controlling a heating, ventilation, and air conditioning (HVAC) system including an actuator, an input interface configured to obtain set values of the HVAC system from user input and obtain measurement data from sensors disposed in the HVAC system, at least one memory configured to store instructions for implementing a method implemented by a computer, at least one processor coupled to the at least one memory, wherein the instructions, when executed by the at least one processor, in steps of the method implemented by the computer, calculate a high-dimensional state estimate using the measurement data and an estimate of the reduced-order state from the RL-trained ROE, determine a controller with respect to the set value by using the RL-trained ROE, generate a control command corresponding to the calculated high-dimensional state estimate based on the controller, and the apparatus further includes an output interface configured to transmit the control command including control instructions for controlling the actuator that operates the HVAC system.
10. The apparatus according to claim 9, wherein the controller is designed using optimal control.
11. The apparatus according to claim 9, wherein the controller is designed using Lyapunov design.
12. The apparatus according to claim 9, wherein the controller is designed using robust control.
13. The apparatus according to claim 9, wherein the RL-trained ROE is trained using a proximal policy optimization (PPO) algorithm.
14. The apparatus according to claim 9, wherein the RL-trained ROE is trained using a trust region policy optimization (TRPO) algorithm.
15. The apparatus according to claim 9, wherein the RL-trained ROE is trained using a robust constrained Markov decision process (RCMP) algorithm.
16. The apparatus according to claim 9, wherein the RL-trained ROE is trained using a time-varying non-stationary MDP.