Device and method for adaptively performing control by using a data-driven model

By combining reinforcement learning and closure models, using the combination of ODE and closure models to characterize system dynamics, and updating the closure model to reduce state trajectory differences, the problem of designing optimal control strategies in the existing technology in the absence of accurate models of dynamic systems is solved, and the physical characteristic capture of system behavior and the optimization of control strategies is achieved.

CN115298622BActive Publication Date: 2025-06-17MITSUBISHI ELECTRIC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180021437.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-20
Filing Date
2021-01-08
Publication Date
2025-06-17
Estimated Expiration
2041-01-08

AI Technical Summary

Technical Problem

The prior art is difficult to design optimal control strategies without an accurate model of the dynamic system, especially when the system dynamics are complex and the model is inaccurate.

Method used

Reinforcement learning (RL) combined with closure model is used to characterize system dynamics through a combination of ODE and closure model, and the closure model is updated using RL to reduce the differences in state trajectories.

Benefits of technology

It realizes capturing the physical characteristics of system behavior without relying on the system physical dynamics model, simplifying data-driven estimation of system dynamics, and improving the optimization effect of control strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115298622B_ABST
    Figure CN115298622B_ABST
Patent Text Reader

Abstract

A device for controlling the operation of a system is provided. The device includes: an input interface configured to receive a state trajectory of the system; and a memory configured to store a dynamics model of the system including a combination of at least one differential equation and a closure model. The device further includes a processor configured to: update the closure model using reinforcement learning (RL) having a value function that reduces the difference between the shape of the received state trajectory and the shape of the state trajectory estimated using the model with the updated closure model; and determine a control command based on the model with the updated closure model. Additionally, the device includes an output interface configured to send the control command to an actuator of the system to control the operation of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to system modeling and control, and more specifically, to a method and apparatus for data-driven model adaptation for modeling, simulating, and controlling a machine using reinforcement learning. Background Art

[0002] Control theory in control systems engineering is a subfield of mathematics that deals with the control of dynamic systems that operate continuously in engineering processes and machines. The goal is to develop control strategies that optimally use control actions to control such systems without delay or overshoot and ensure control stability.

[0003] For example, optimization-based control and estimation techniques (such as model predictive control (MPC)) allow a model-based design framework in which system dynamics and constraints can be directly considered. MPC is used in many applications to control dynamic systems of various complexities. Examples of such systems include production lines, automotive engines, robots, numerical control machining, electric motors, satellites, and generators. As used herein, a system dynamics model or system model describes system dynamics using differential equations. For example, the most general model of a linear system with p inputs u, q outputs y, and n state variables x is written in the following form:

[0004]

[0005] y(t) = C(t)x(t) + D(t)u(t).

[0006] However, in many cases, the model of the controlled system is non-linear and may be difficult to design, use in real time, or may be inaccurate. Examples of such cases are prevalent in robotics, building control (HVAC), smart grids, factory automation, transportation, self-tuning machines, and traffic networks. Additionally, even if the non-linear model is fully available, designing an optimal controller is inherently a challenging task because it requires solving a partial differential equation known as the Hamilton-Jacobi-Bellman (HJB) equation.

[0007] In the absence of an accurate model of the dynamic system, some control methods utilize the operational data generated by the dynamic system to construct a feedback control strategy that stabilizes the system dynamics or embeds quantifiable control-related performance. Using operational data to design a control strategy is called data-driven control. There are two data-driven control methods: (i) an indirect method that first constructs a system model and then uses the model to design a controller; and (ii) a direct method that directly constructs a control strategy from the data without an intermediate model construction step.

[0008] The disadvantage of the indirect method is that potentially a large amount of data is required during the model construction phase. Additionally, in the indirect control method, the controller is calculated, for example, from the estimated model according to the certainty equivalence principle, but in practice, the model estimated from data does not capture the physical characteristics of the dynamics of the system. Therefore, many model-based control techniques cannot be used with such data-driven models.

[0009] To overcome this problem, some methods use a direct control method that directly maps experimental data to the controller without the need to identify any model in between. However, the direct control method results in a black-box design of the control strategy that directly maps the system state to the control command. However, this control strategy is designed without considering the physical characteristics of the system. Additionally, the control designer cannot influence the data-driven determination of the control strategy.

[0010] Therefore, there is still a need for a method and device for controlling a system in an optimal manner. Summary of the Invention

[0011] An object of some embodiments is to provide an apparatus and method for data-driven design of a dynamic model of a system to generate a dynamic model of the system that captures the physical characteristics of the system behavior. In this way, the embodiments simplify the model design process while retaining the advantages of having a system model when designing control applications. However, current data-driven methods are not suitable for estimating a system model that captures the physical dynamics of the system.

[0012] For example, reinforcement learning (RL) is a field of machine learning that involves certain concepts of how to act in an environment to maximize cumulative reward (or equivalently, minimize cumulative loss / cost). Reinforcement learning is related to optimal control in a continuous state input space that mainly focuses on the existence and characterization of optimal control policies, and algorithms for computing them without a mathematical model of the controlled system and / or environment.

[0013] Given the advantages provided by RL methods, some embodiments aim to develop RL techniques for obtaining an optimal control policy for a dynamic system that can be described by differential equations. However, the control policy maps the state of the system to the control command, and it is not necessary or at least not mandatory to perform this mapping based on the physical dynamics of the system. Therefore, the control community has not explored RL data-driven estimation of models with one or more differential equations that describe the dynamics of the system and have physical meaning.

[0014] Some embodiments are based on the recognition that RL data-driven learning of system dynamics models with physical meaning can be regarded as a virtual control problem, where the reward function minimizes the difference between the behavior of the system according to the learned model and the actual behavior of the system. It is worth noting that the behavior of the system is a high-level representation of the system, for example, the stability of the system, the boundedness of the state. In fact, even in an uncontrolled situation, the system has behavior. Unfortunately, estimating such a model through RL is computationally challenging.

[0015] To this end, some embodiments are based on the understanding that the model of the system can be characterized by a reduced-order model combined with a virtual control term, which we call a closure model. For example, if the model of the system based on all physical properties is usually captured by a partial differential equation (PDE), the reduced-order model can be characterized by an ordinary differential equation (ODE). The ODE characterizes the dynamics of the system as a function of time, but is not as accurate as the characterization of the dynamics using the PDE. Therefore, the goal of the closure model is to reduce this gap.

[0016] As used herein, the closure model is a non-linear function of the system state that captures the difference between the system behavior estimated by the ODE and the PDE. Therefore, the closure model is also a function of time, characterizing the dynamic difference between the dynamics captured by the ODE and the PDE. Some embodiments are based on the understanding that characterizing the dynamics of the system as a combination of the ODE and the closure model can simplify the subsequent control of the system, because solving the PDE equation is computationally expensive. Therefore, some embodiments attempt to simplify the data-driven estimation of the dynamics of the system by characterizing the dynamics with the ODE and the closure model and only updating the closure model. However, although this problem is computationally simpler, it also has challenges in formulating the RL framework. This is because RL is usually used to learn control strategies to precisely control the system. Here, in this analogy, RL should attempt to accurately estimate the closure model, which is challenging.

[0017] However, some embodiments are based on the recognition that in many modeling situations, it is sufficient to characterize the behavior pattern of the system dynamics rather than the exact behavior itself. For example, when the exact behavior captures the energy of the system at each time point, the behavior pattern captures the rate of change of the energy. By analogy, when the system is excited, the energy of the system increases. Knowing the exact behavior of the system dynamics allows the evaluation of this energy increase. Knowing the behavior pattern of the system dynamics allows the evaluation of the increase rate to estimate a new energy value proportional to its actual value.

[0018] Therefore, the behavioral pattern of system dynamics itself is not an exact behavior. However, in many model-based control applications, the behavioral pattern of system dynamics is sufficient to design Lyapunov stable control. Examples of such control applications include stabilization control aimed at stabilizing the state of a system.

[0019] To this end, some embodiments use RL to update the closure model such that the dynamics of the ODE and the updated CL mimic the dynamic pattern of the system. Some embodiments are based on the recognition that the dynamic pattern can be characterized by the shape of a state trajectory that is determined as a function of time and compared with the state values of the system. The state trajectory can be measured during the online operation of the system. Additionally or alternatively, a PDE can be used to simulate the state trajectory.

[0020] To this end, some embodiments use a system model that includes a combination of an ODE and a closure model to control the system, and update the closure model with RL that has a value function that reduces the difference between the actual shape of the state trajectory and the shape of the state trajectory estimated using the ODE with the updated closure model.

[0021] However, after convergence, the ODE with the updated CL characterizes the dynamic pattern of the system behavior, rather than the actual value of the behavior. In other words, the ODE with the updated CL is a function proportional to the actual physical dynamics of the system. To this end, some embodiments include a gain in the closure model and then use a method more suitable for model-based optimization than RL to learn the closure model during the online control of the system. Examples of these methods are extremum seeking, Gaussian process-based optimization, etc.

[0022] Additionally or alternatively, some embodiments use a system model determined by data-driven adaptation in various model-based predictive controls (e.g., MPC). These embodiments allow leveraging the capabilities of MPC to account for constraints in system control. For example, classical RL methods are not applicable to data-driven control of constrained systems. This is because classical RL methods do not consider state and input constraint satisfaction in a continuous state-action space; that is, classical RL cannot guarantee that the state of the controlled system operated with control inputs satisfies state and input constraints throughout the operation.

[0023] However, some embodiments use RL to learn the physical characteristics of the system, thus allowing the combination of the data-driven advantages of RL with model-based constraint optimization.

[0024] Accordingly, one embodiment discloses an apparatus for controlling the operation of a system. The apparatus includes: an input interface configured to receive a state trajectory of the system; a memory configured to store a dynamics model of the system, the dynamics model including a combination of at least one differential equation and a closure model; a processor configured to: update the closure model using reinforcement learning (RL) having a value function that reduces the difference between the shape of the received state trajectory and the shape of the state trajectory estimated using the model with the updated closure model; and determine a control command based on the model with the updated closure model; and an output interface configured to send the control command to an actuator of the system to control the operation of the system.

[0025] Another embodiment discloses a method for controlling the operation of a system. The method uses a processor coupled to a memory that stores a dynamics model of the system including a combination of at least one differential equation and a closure model, the processor being coupled to stored instructions that, when executed by the processor, implement the steps of the method, the method including: receiving a state trajectory of the system; using reinforcement learning (RL) to update the closure model having a value function that reduces the difference between the shape of the received state trajectory and the shape of the state trajectory estimated using the model with the updated closure model; determining a control command based on the model with the updated closure model; and sending the control command to an actuator of the system to control the operation of the system.

[0026] The presently disclosed embodiments will be further described with reference to the accompanying drawings. The drawings shown are not necessarily to scale, but the focus is generally on illustrating the principles of the presently disclosed embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1

[0028] Figure 1 A schematic overview of the principles used in some embodiments for controlling the operation of a system is shown.

[0029] Figure 2

[0030] Figure 2 A block diagram of an apparatus for controlling the operation of a system according to some embodiments is shown.

[0031] Figure 3

[0032] Figure 3 A flowchart of the principles for controlling a system according to some embodiments is shown.

[0033] Figure 4

[0034] ​​​​​​​​Figure 4 Shows a schematic architecture for generating a reduced-order model according to some embodiments.

[0035] Figure 5A

[0036] Figure 5A Shows a schematic diagram of a reduced-order model based on reinforcement learning (RL) according to some embodiments.

[0037] Figure 5B

[0038] Figure 5B Shows a flowchart of operations for updating a closure model using RL according to an embodiment of the present invention.

[0039] Figure 6

[0040] Figure 6 Shows the difference between the actual behavior and the estimated behavior of a system according to some embodiments.

[0041] Figure 7A

[0042] Figure 7A Shows a schematic diagram of an optimization algorithm for tuning an optimal closure model according to an embodiment of the present invention.

[0043] Figure 7B

[0044] Figure 7B Shows a schematic diagram of an optimization algorithm for tuning an optimal closure model according to an embodiment of the present invention.

[0045] Figure 7C

[0046] Figure 7C Shows a schematic diagram of an optimization algorithm for tuning an optimal closure model according to an embodiment of the present invention.

[0047] Figure 8A

[0048] Figure 8A Shows a schematic diagram of an optimization algorithm for tuning an optimal closure model according to another embodiment of the present invention.

[0049] Figure 8B

[0050] Figure 8B Shows a schematic diagram of an optimization algorithm for tuning an optimal closure model according to another embodiment of the present invention.

[0051] Figure 8C ​​​​​​​​​​​​​​​​​​

[0052] Figure 8C Shows a schematic diagram of an optimization algorithm for tuning an optimal closure model according to another embodiment of the present invention.

[0053] Figure 9A

[0054] Figure 9A Shows a flowchart of an extremum seeking (ES) algorithm for updating gain according to some embodiments.

[0055] Figure 9B

[0056] Figure 9B Shows a flowchart of an extremum seeking (ES) algorithm for updating gain using a performance cost function according to some embodiments.

[0057] Figure 10

[0058] Figure 10 Shows a schematic diagram of an extremum seeking (ES) controller for single-parameter tuning according to some embodiments.

[0059] Figure 11

[0060] Figure 11 Shows a schematic diagram of an extremum seeking (ES) controller for multi-parameter tuning according to some embodiments.

[0061] Figure 12

[0062] Figure 12 Shows a predictive model-based algorithm for considering constraints of a control system according to some embodiments.

[0063] Figure 13

[0064] Figure 13 Shows an exemplary real-time implementation of a device for a control system, where the system is an air conditioning system.

[0065] Figure 14A

[0066] Figure 14A Shows an exemplary real-time implementation of a device for a control system, where the system is a vehicle.

[0067] Figure 14B

[0068] Figure 14B Shows a schematic diagram of the interaction between a controller and a vehicle controller according to some embodiments.

[0069] ​​​​​​​​​​​​​​​​​Figure 15

[0070] Figure 15 An exemplary real - time implementation of a device for controlling a system is shown, where the system is an induction motor. Detailed Description

[0071] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to one of ordinary skill in the art that the present disclosure may be practiced without these specific details. In other instances, devices and methods are shown only in block diagram form in order to avoid obscuring the present disclosure.

[0072] As used in this specification and the claims, when the terms "such as", "like", and "for example" and the verbs "comprising", "having", "including", and their other verb forms are used in conjunction with a list of one or more components or other items, each shall be construed as open - ended, meaning that the list should not be considered as excluding other additional components or items. The term "based on" means at least in part based on. Additionally, it should be understood that the language and terminology used herein are for the purpose of description and should not be regarded as limiting. Any headings used in this description are for convenience only and have no legal or limiting effect.

[0073] Figure 1 A schematic overview of the principles used in some implementations for the operation of a control system is shown. Some implementations provide a control device 100 configured to control a system 102. For example, the device 100 can be configured to control a dynamic system 102 that operates continuously in engineering processes and machines. Hereinafter, "control device" and "device" may be used interchangeably and will have the same meaning. Hereinafter, "power system operating continuously" and "system" may be used interchangeably and will have the same meaning. Examples of the system 102 are HVAC systems, LIDAR systems, condensing units, production lines, self - tuning machines, smart grids, automotive engines, robots, numerical control machining, electric motors, satellites, generators, transportation networks, etc. Some implementations are based on the recognition that the device 100 develops a control strategy 106 for optimally using control actions to control the system 102 without delay or overshoot and ensuring control stability.

[0074] ​In some embodiments, the device 100 uses model-based and / or optimization-based control and estimation techniques such as model predictive control (MPC) to develop control commands 106 for the system 102. Model-based techniques can be advantageous for the control of dynamic systems. For example, MPC allows for a model-based design framework in which the dynamics and constraints of the system 102 can be directly considered. MPC develops control commands 106 based on a model 104 of the system. The model 104 of the system 102 refers to the dynamics of the system 102 described using differential equations. In some embodiments, the model 104 is non-linear and may be difficult to design and / or difficult to use in real time. For example, even if the non-linear model is fully available, estimating the optimal control commands 106 is essentially a challenging task because it requires solving partial differential equations (PDEs) (referred to as the Hamilton-Jacobi-Bellman (HJB) equations) that describe the dynamics of the system 102, which is computationally challenging.

[0075] Some embodiments use data-driven control techniques to design the model 104. Data-driven techniques utilize the operational data generated by the system 102 in order to construct a feedback control strategy that stabilizes the system 102. For example, each state of the system 102 measured during operation of the system 102 can be given as feedback for controlling the system 102. Generally, using operational data to design control strategies and / or commands 106 is referred to as data-driven control. The goal of data-driven control is to design a control strategy based on data and to use the data-driven control strategy to control the system. Compared with such data-driven control methods, some embodiments use operational data to design a model of the control system (e.g., model 104), and then use the data-driven model to control the system using various model-based control methods. It should be noted that the purpose of some embodiments is to determine the actual model of the system from data, i.e., a model that can be used to estimate the behavior of the system. For example, the purpose of some embodiments is to determine a system model that captures the system dynamics using differential equations from data. Additionally or alternatively, the purpose of some embodiments is to learn a model from data that has the accuracy of a PDE model based on physical characteristics.

[0076] To simplify calculations, some embodiments formulate an ordinary differential equation (ODE) 108a to describe the dynamics of system 102. In some embodiments, model reduction techniques can be used to formulate the ODE 108a. For example, the ODE 108a can be a reduced-order PDE. To this end, the ODE 108a can be part of the PDE. However, in some embodiments, under conditions of uncertainty, the ODE 108a cannot reproduce the actual dynamics of system 102 (i.e., the dynamics described by the PDE). Examples of uncertainty conditions can be cases where the boundary conditions of the PDE are changing over time or where one of the coefficients involved in the PDE is changing.

[0077] To this end, some embodiments formulate a closure model 108b that reduces the PDE while covering cases of uncertainty conditions. In some embodiments, the closure model 108b can be a non-linear function of the state of system 102 that captures the difference in the behavior (e.g., dynamics) of system 102 based on the ODE and the PDE. Reinforcement learning (RL) can be used to formulate the closure model 108b. In other words, the PDE model of system 102 is approximated by the combination of the ODE 108a and the closure model 108b, and RL is used to learn the closure model 108b from the data. In this way, a model close to the PDE accuracy is learned from the data.

[0078] In some embodiments, RL learning defines the state trajectory of system 102 that defines the behavior of system 102, rather than learning the individual states of system 102. The state trajectory can be a sequence of states of system 102. Some embodiments are based on the recognition that model 108, which includes ODE 108a and closure model 108b, reproduces the behavior pattern of system 102, rather than the actual behavior values (e.g., states) of system 102. The behavior pattern of system 102 can represent the shape of the state trajectory, e.g., the sequence of states of the system as a function of time. The behavior pattern of system 102 can also represent the high-level features of the model, such as the boundedness of its solution over time, or the decay of its solution over time. However, it cannot optimally reproduce the dynamics of the system.

[0079] To this end, some embodiments determine a gain and include the gain in the closure model 108b to optimally reproduce the dynamics of the system 102. In some embodiments, an optimization algorithm can be used to update the gain. The model 108, including the ODE 108a and the closure model 108b with the updated gain, reproduces the dynamics of the system 102. Thus, the model 108 optimally reproduces the dynamics of the system 102. Some embodiments are based on the recognition that the model 108 includes a smaller number of parameters than the PDE. To this end, the model 108 is computationally less complex than the PDE of the physical model describing the system 102. In some embodiments, the model 108 is used to determine the control strategy 106. The control strategy 106 directly maps the state of the system 102 to a control command to control the operation of the system 102. Thus, the simplified model 108 is used to design the control of the system 102 in an efficient manner.

[0080] Figure 2 FIG. shows a block diagram of a device 200 for controlling the operation of a system 102 according to some embodiments. The device 200 includes an input interface 202 and an output interface 218 for connecting the device 200 to other systems and devices. In some embodiments, the device 200 may include multiple input interfaces and multiple output interfaces. The input interface 202 is configured to receive the state trajectory 216 of the system 102. The input interface 202 includes a network interface controller (NIC) 212, which is adapted to connect the device 200 to the network 214 via a bus 210. Through the network 214, wirelessly or via wire, the device 200 receives the state trajectory 216 of the system 102.

[0081] The state trajectory 216 may be multiple states of the system 102 that define the actual behavior of the dynamics of the system 102. For example, the state trajectory 216 serves as a reference continuous state space for controlling the system 102. In some embodiments, the state trajectory 216 can be received from real-time measurements of some parts of the system 102 state. In some other embodiments, the PDE describing the dynamics of the system 102 can be used to simulate the state trajectory 216. In some embodiments, the shape as a function of time can be determined for the received state trajectory. The shape of the state trajectory can represent the actual behavior pattern of the system 102.

[0082] The device 200 further includes a processor 204 and a memory 206 that stores instructions executable by the processor 204. The processor 204 can be a single-core processor, a multi-core processor, a computing cluster, or any number of other configurations. The memory 206 may include random access memory (RAM), read-only memory (ROM), flash memory, or any other suitable memory system. The processor 204 is connected to one or more input / output devices via a bus 210. The stored instructions implement a method for controlling the operation of the system 102.

[0083] The memory 206 can be further extended to include a storage 208. The storage 208 can be configured to store a model 208a, a controller 208b, an update module 208c, and a control command module 208d. In some embodiments, the model 208a can be a model that describes the dynamics of the system 102, which includes a combination of at least one differential equation and a closure model. The differential equation of the model 208a can be an ordinary differential equation (ODE) 108a. The closure model of the model 208a can be a linear or non-linear function of the state of the system 102. RL can be used to learn the closure model to mimic the behavior of the system 102. It should be understood that once the closure model is learned, the closure model can be the closure model 108b as Figure 1 illustrated.

[0084] The controller 208b can be configured to store instructions that, once executed by the processor 204, execute one or more modules in the storage 208. Some embodiments are based on the recognition that the controller 208b manages each module of the storage 208 to control the system 102.

[0085] The update module 208c can be configured to update the closure model of the model 208a using reinforcement learning (RL), which has a value function that reduces the difference between the shape of the received state trajectory and the shape of the state trajectory estimated using the model 208a with the updated closure model. In some embodiments, the update module 208c can be configured to iteratively update the closure module with RL until a termination condition is met. The updated closure model is a non-linear function of the system state that captures the difference in system behavior according to ODE and PDE.

[0086] In addition, in some embodiments, the update module 208c can be configured to update the gain for the updated closure model. To this end, some embodiments determine a gain that reduces the error between the state of the system 102 estimated using the model 208a with the updated closure model including the updated gain and the actual state of the system. In some embodiments, the actual state of the system can be a measured state. In some other embodiments, the actual state of the system can be a state estimated using a PDE that describes the dynamics of the system 102. In some embodiments, the update module 208c can use extremum search to update the gain. In some other embodiments, the update module 208c can use Gaussian process-based optimization to update the gain.

[0087] The update module 208c may be configured to determine a control command based on the model 208a with an updated closure model. The control command may control the operation of the system. In some embodiments, the operation of the system may be subject to constraints. To this end, while implementing the constraints, the update module 208c uses control based on a predictive model to determine the control command. The constraints include state constraints in the continuous state space of the system 102 and control input constraints in the continuous control input space of the system 102.

[0088] The output interface 218 is configured to send a control command to the actuator 220 of the system 102 to control the operation of the system. Some examples of the output interface 218 may include submitting a control command to a control interface that controls the system 102.

[0089] Figure 3 A flowchart showing the principle for controlling the system 102 according to some embodiments is shown. Some embodiments are based on the recognition that the system 102 can be modeled according to physical laws. For example, the dynamics of the system 102 can be characterized by using mathematical equations of physical laws. In step 302, the system 102 can be characterized by a high-dimensional model based on physical properties. The high-dimensional model based on physical properties can be a partial differential equation (PDE) that describes the dynamics of the system 102. For ease of explanation, the system 102 is considered to be an HVAC system, and its model is characterized by the Boussinesq equation. The Boussinesq equation is obtained from physical properties and describes the connection between air flow and indoor temperature. Thus, the HAVC system model can be mathematically characterized as:

[0090]

[0091]

[0092]

[0093] where T is the temperature scalar variable, is the three-dimensional velocity vector, μ is the reciprocal of the viscosity and the Reynolds number, k is the heat dissipation coefficient, and p is the pressure scalar variable.

[0094] The operators Δ and are defined as:

[0095]

[0096]

[0097] Some embodiments are based on the recognition that a high-dimensional model based on physical properties of system 102 needs to be solved to control the operation of system 102 in real time. For example, in the case of an HVAC system, the Boussinesq equations need to be solved to control air flow dynamics and indoor temperature. Some embodiments are based on the recognition that the high-dimensional model based on physical properties of system 102 involves solving a large number of complex equations and variables. For example, greater computational power is required to solve the high-dimensional model based on physical properties in real time. To this end, some embodiments aim to simplify the high-dimensional model based on physical properties.

[0098] In step 304, device 200 is configured to generate a reduced-order model to reproduce the dynamics of system 102 such that device 200 can control system 102 in an efficient manner. In some embodiments, device 200 may use model reduction techniques to simplify the high-dimensional model based on physical properties to generate a reduced-order model. Some embodiments are based on the recognition that model reduction techniques reduce the dimension of the high-dimensional model based on physical properties (e.g., variables of a PDE) such that the reduced-order model can be used to predict and control system 102 in real time. Additionally, reference Figure 4 is made to the detailed description of the generation of the reduced-order model for controlling system 102. In step 306, device 200 uses the reduced-order model to predict and control system 102 in real time.

[0099] Figure 4 A schematic architecture for generating a reduced-order model according to some embodiments is shown. Some embodiments are based on the recognition that device 200 uses model reduction techniques to generate a reduced-order model (ROM) 406. The ROM 406 generated using model reduction techniques can be a part 402 of the high-dimensional model based on physical properties. The part 402 of the high-dimensional model based on physical properties can be one or more differential equations that describe the dynamics of system 102. The part 402 of the high-dimensional model based on physical properties can be ordinary differential equations (ODEs). In some embodiments, in the case of uncertainty conditions, the ODEs cannot reproduce the actual dynamics (i.e., the dynamics described by the PDE). Examples of uncertainty conditions can be the case where the boundary conditions of the PDE are changing over time or the case where one of the coefficients involved in the PDE is changing. These mathematical changes actually reflect some real changes in the actual dynamics. For example, in the case of an HVAC system, the opening or closing of windows and / or doors in a room changes the boundary conditions of the Boussinesq equations (i.e., the PDE). Similarly, weather changes (such as daily and seasonal changes) can affect the difference between the indoor temperature and the outdoor temperature of the room, which in turn can affect some PDE coefficients (e.g., the Reynolds number).

[0100] In all of these scenarios, model reduction techniques fail to have a unified method to obtain a reduced-order (or reduced-dimension) model 406 of the dynamics of system 102 that covers all of the above scenarios (i.e., parametric uncertainty and boundary condition uncertainty).

[0101] Some embodiments are aimed at generating a ROM 406 of a PDE that reduces under changing boundary conditions and / or changing parameters. To this end, some embodiments use adaptive model reduction methods, physical regime detection methods, etc.

[0102] For example, in one embodiment of the present invention, the reduced-order model 406 has a quadratic form:

[0103]

[0104] where b, A, B are constants related to the constants of the PDE equation and the type of model reduction algorithm used, and x r is a vector with reduced dimension r and represents the reduced-order state. The original state x of the system can be recovered from x r using the following simple algebraic equation:

[0105] x(t)≈Φx r (t)

[0106] where x is typically a vector of high dimension n>>r, containing n states obtained from the spatial discretization of the PDE equation, and Φ is a matrix formed by cascading given vectors called modes or basis vectors of the ROM 406. These modes vary depending on which model reduction method is used. Examples of model reduction methods include proper orthogonal decomposition (POD), dynamic mode decomposition (DMD) methods, etc.

[0107] However, the solution of the equation of ROM406 may result in an unstable solution (diverging on a finite time support), which cannot reproduce the physical characteristics of the original PDE model with viscous terms that keep the solution stable all the time (i.e., bounded on a bounded time support). For example, during the model reduction process, the ODE may lose the inherent characteristics of the actual solution of the high-dimensional model based on physical characteristics. To this end, the ODE may lose the boundedness of the actual solution of the high-dimensional model based on physical characteristics in space and time.

[0108] Therefore, some embodiments modify the ROM 406 by adding a closure model 404 that characterizes the difference between the ODE and the PDE. For example, the closure model 404 captures the lost inherent characteristics of the actual solution of the PDE and acts as a stabilizing factor. Some embodiments allow only the closure model 404 to be updated to reduce the difference between the ODE and the PDE.

[0109] For example, in some embodiments, ROM 406 can be mathematically characterized as:

[0110]

[0111] The function F is the closure model 404, which is added to stabilize the solution of ROM 406. The term characterizes the ODE. The term K characterizes the coefficient vector that should be tuned to ensure stability and the fact that ROM 406 needs to reproduce the dynamics or solution of the original PDE model. In some embodiments, the closure model 404 is a linear function of the state of system 102. In some other embodiments, the closure model 404 can be a non - linear function of the state of system 102. In some embodiments, a data - driven method based on reinforcement learning (RL) can be used to calculate the closure model 404. Additionally, reference is made to Figure 5A and Figure 5B for a detailed description of calculating the closure model 404 using reinforcement learning (RL).

[0112] Figure 5A FIG. shows a schematic diagram of a reduced - order model 406 based on reinforcement learning (RL) according to some embodiments. In some embodiments, a data - driven method based on RL can be used to calculate the RL - based closure model 502. Some embodiments are based on the recognition that the closure model 402 is iteratively updated with RL to calculate the RL - based closure model 502. The RL - based closure model 502 can be an optimal closure model. Additionally, reference is made to Figure 5B for a detailed description of the iterative process for updating the closure model 404. Some embodiments are based on the understanding that the optimal closure model combined with the ODE can form the optimal ROM 406. In some embodiments, ROM 406 can estimate the actual behavior pattern of system 102. For example, ROM 406 simulates the received state trajectory.

[0113] Figure 5BFIG. 0 shows an operational flowchart of using RL to update the closure model 502 according to an embodiment of the present invention. At step 504, the device 200 may be configured to initialize an initial closure model policy and a learning cumulative reward function associated with the initial closure model policy. The initial closure model policy may be a simple linear closure model policy. The cumulative reward function may be a value function. At step 506, the device 200 is configured to run the ROM 406 including a part 402 of a high-dimensional model based on physical characteristics and the current closure model (e.g., the initial closure model policy) to collect data along a finite time interval. To this end, the device 200 collects data characterizing the dynamic behavior pattern of the system 102. For example, the behavior pattern captures the rate of change of energy of the system 102 over a finite time interval. Some embodiments are based on the recognition that the dynamic behavior pattern of the system 102 can be characterized by the shape of the state trajectory over a finite time interval.

[0114] At step 508, the device 200 is configured to update the cumulative reward function using the collected data. In some embodiments, the device 200 updates the cumulative reward function (i.e., the value function) to indicate the difference between the shape of the received state trajectory and the shape of the state trajectory estimated using the ROM 406 with the current closure model (e.g., the initialized closure model).

[0115] Some embodiments are based on the recognition that RL uses a trained neural network to minimize the value function. To this end, at step 510, the device 200 is configured to update the current closure model policy using the collected data and / or the updated cumulative reward function to minimize the value function.

[0116] In some embodiments, the device 200 is configured to repeat steps 506, 508, and 510 until a termination condition is met. To this end, at step 512, the device 200 is configured to determine whether the learning converges. For example, the device 200 determines whether the learning cumulative reward function is lower than a threshold limit, or whether two consecutive learning cumulative reward functions are within a small threshold limit. If the learning converges, the device 200 proceeds to step 516, otherwise the device 200 proceeds to step 514. At step 514, the device 200 is configured to replace the closure model with the updated closure model and iterate the update process until the termination condition is met. In some embodiments, the device 200 iterates the update process until the learning converges. At step 516, the device 200 is configured to stop the closure model learning and use the last updated closure model policy as the optimal closure model of the ROM 406.

[0117] For example, given a closure model policy u(x), some embodiments define the infinite-horizon cumulative reward function for a given initial state as

[0118]

[0119] wherein, is a positive definite function, where and {x k} refers to the state sequence generated by the closed-loop system:

[0120] x t+1 = Ax t + Bu(x t ) + Gφ(C q x t )

[0121] In some embodiments, the scalar γ ∈ (0, 1] is a forgetting / discount factor, which aims to emphasize the cost more through the current state and control actions and reduce the trust in the past.

[0122] The continuous control strategy is an admissible control strategy on if it stabilizes the closed-loop system on X and is finite for any initial state x0 in X. An optimal control strategy that achieves the optimal cumulative reward for any initial state x0 in X can be designed

[0123]

[0124] Here, refers to the set of all admissible control strategies. In other words, the optimal control strategy can be calculated as:

[0125]

[0126] Directly constructing such an optimal controller is very challenging for general nonlinear systems; this situation is further exacerbated due to the uncertain dynamics in the system. Therefore, some embodiments use adaptive / approximate dynamic programming (ADP): a class of iterative, data-driven algorithms that generate a convergent sequence of control strategies u ∞ (x).

[0127] According to the Bellman optimality principle, the discrete-time Hamilton-Jacobi-Bellman equation is given by:

[0128]

[0129]

[0130] The ADP method typically involves performing iterations on the cumulative reward function and the closed-loop model policy to ultimately converge to the optimal value function and the optimal closed-loop model policy. The key operations in the ADP method involve setting an admissible closed-loop model policy u0(x), and then iterating the policy evaluation step until convergence.

[0131]

[0132]

[0133] For example, according to some embodiments, F = K0x is an admissible closed-loop model policy, and the learning cumulative reward function approximator is:

[0134]

[0135] where ψ(x) is a set of differentiable basis functions (equivalent to, hidden layer neuron activations), and ω k is the corresponding column vector of basis coefficients (equivalent to, neural network weights). Thus, the initial weight vector is ω0.

[0136] In one embodiment, when the goal of ROM 406 is to generate a solution that minimizes the quadratic value function:

[0137]

[0138] where R and Q are two user-defined positive weight matrices.

[0139] Then, the closed-loop model policy improvement step is given by:

[0140]

[0141] Some embodiments are based on the understanding that the generated ROM406 (e.g., optimal ROM) including ODE 402 and the optimal closed-loop model simulates the actual behavior pattern of system 102, rather than the actual behavior values. In other words, ODE 402 with the optimal closed-loop model is a function proportional to the actual physical dynamics of system 102. For example, the behavior of the optimal ROM 406 (i.e., the estimated behavior) may be quantitatively similar to the actual behavior of system 102, but there may be a quantitative gap between the actual behavior of system 102 and the estimated behavior. In addition, refer to Figure 6 for a detailed description of the differences between the actual behavior and the estimated behavior.

[0142] Figure 6Shows the difference between the actual behavior and the estimated behavior of system 102 according to some embodiments. In some embodiments, the behavior pattern of system 102 can be characterized by two-dimensional axes, where the x-axis corresponds to time and the y-axis corresponds to the energy magnitude of system 102. Wave 602 can characterize the behavior of the actual system 102. Wave 604 can characterize the estimated behavior of system 102. Some embodiments are based on the recognition that there can be a quantitative gap 606 between the actual behavior 602 and the estimated behavior 604. For example, the actual behavior 602 and the estimated behavior 604 can have similar frequencies but different amplitudes.

[0143] To this end, the aim of some embodiments is to include a gain in the optimal closure model such that the gap 606 between the actual behavior 602 and the estimated behavior 604 is reduced. For example, in some embodiments, the closure model can be characterized as:

[0144]

[0145] where θ is a positive gain that needs to be optimally tuned to minimize the learning cost function Q such that the gap 606 between the actual behavior 602 and the estimated behavior 604 is reduced. Additionally, refer to Figures 7A to 7C for a detailed description of how device 200 determines the gain for reducing the gap 606.

[0146] Figures 7A to 7C Shows a schematic diagram of an optimization algorithm for tuning an optimal closure model according to an embodiment of the present invention. Some embodiments are based on the recognition that for small time intervals, the ROM 406 including the ODE 402 and the optimal closure model (i.e., the optimal ROM 406) can be useful. In other words, only for small time intervals, the optimal ROM 406 forces the behavior of system 102 to be bounded. To this end, the aim of some embodiments is to tune the gain (also referred to as a coefficient) of the optimal ROM 406 over time.

[0147] In an embodiment, device 200 uses the high-dimensional model behavior 702 based on physical characteristics (i.e., the actual behavior 602) to tune the gain of the optimal closure model. In some example embodiments, device 200 calculates the error 706 between the estimated behavior 704 corresponding to the optimal ROM 406 and the behavior 702. Additionally, device 200 determines the gain that reduces the error 706. Some embodiments are based on the recognition that device 200 determines the gain that reduces the error 706 between the state of system 102 estimated using the optimal ROM 406 (i.e., the estimated behavior 704) and the actual state of system 102 estimated using the PDE (i.e., the behavior 702). In some embodiments, device 200 updates the determined gain into the optimal closure model to include the determined gain.

[0148] Some embodiments are based on the recognition that device 200 uses an optimization algorithm to update the gain. In one embodiment, the optimization algorithm can be extremum seeking (ES) 710, as Figure 7B exemplarily shown therein. In another embodiment, the optimization algorithm can be Gaussian process-based optimization 712, as Figure 7C exemplarily shown therein.

[0149] Figures 8A to 8C A schematic diagram of an optimization algorithm for tuning an optimal closure model according to another embodiment of the present invention is shown. Some embodiments are based on the recognition that for small time intervals, the optimal ROM 406 can be useful. In other words, only for small time intervals, the optimal ROM 406 forces the behavior of system 102 to be bounded. To this end, the goal of some embodiments is to tune the gain of the optimal ROM 406 over time.

[0150] In an embodiment, device 200 uses the real-time measured state 802 (i.e., the actual behavior 602) of a part of system 102 to tune the gain of the optimal closure model. In some example embodiments, device 200 calculates an error 806 corresponding to the estimated behavior 804 of the optimal ROM 406 and the actual behavior 602 (e.g., the real-time measured state 802 of system 102). In addition, device 200 determines a gain that reduces the error 806. Some embodiments are based on the recognition that device 200 determines a gain that reduces the error 806 between the state of system 102 estimated using the optimal ROM 406 (i.e., the estimated behavior 704) and the actual state of system 102 (i.e., the real-time measured state 802). In some embodiments, device 200 updates the determined gain in the optimal closure model to include the determined gain.

[0151] Some embodiments are based on the recognition that device 200 uses an optimization algorithm to update the gain. In one embodiment, the optimization algorithm can be extremum seeking (ES) 810, as Figure 8B exemplarily shown therein. In another embodiment, the optimization algorithm can be Gaussian process-based optimization 812, as Figure 8C exemplarily shown therein.

[0152] Figure 9AFIG. 0 shows a flowchart of an extremum seeking (ES) algorithm 900 for updating a gain according to some embodiments. Some embodiments are based on the recognition that the ES algorithm 900 is a model-free learning algorithm that allows the device 200 to tune the gain of an optimal closure model. Some embodiments are based on the understanding that the ES algorithm 900 iteratively perturbs the gain of the optimal closure module with a perturbation signal until a termination condition is met. In some embodiments, the perturbation signal may be a periodic signal having a predetermined frequency. In some embodiments, the termination condition may be a condition that the gap 606 may be within a threshold limit. The gain of the optimal closure model may be a control parameter.

[0153] At step 902a, the ES algorithm 900 may perturb the control parameter of the optimal closure model. For example, the ES algorithm 900 may use a perturbation signal to perturb the control parameter. In some embodiments, the perturbation signal may be a previously updated perturbation signal. At step 904a, the ES algorithm 900 may determine a cost function Q regarding the closure model performance in response to perturbing the control parameter. At step 906a, the ES may determine the gradient of the cost function by modifying the cost function with the perturbation signal. For example, the gradient of the cost function is determined as the product of the cost function, the perturbation signal, and the gain of the ES algorithm 900. At step 908a, the ES algorithm 900 may integrate the perturbation signal with the determined gradient to update the perturbation signal for the next iteration. The iterations of the ES 900 may be repeated until the termination condition is met.

[0154] Figure 9B FIG. 7 shows a flowchart of an extremum seeking (ES) algorithm 900 for updating a gain using a performance cost function according to some embodiments. At step 904b, the ES 900 may determine a cost function of the closure model performance. In some embodiments, the ES algorithm 900 determines the cost function at step 904b, as exemplarily shown in step 904a of Figure 9A FIG. 9. In some embodiments, the determined cost function may be a performance cost function 904b-0. According to some example embodiments, the performance cost function 904b-0 may be a quadratic equation characterizing the behavior of the gap 606.

[0155] In step 906b, the ES algorithm 900 may multiply the determined cost function by a first periodic signal 906b-0 of time to generate a perturbed cost function 906b-1. In step 908b, the ES algorithm 900 may subtract a second periodic signal 908b-0 from the perturbed cost function 906b-1 to generate a derivative 908b-1 of the cost function, where the second periodic signal 908b-0 has a ninety-degree quadrature phase shift with respect to the phase of the first periodic signal 906b-0. In step 910b, the ES algorithm 900 may integrate the derivative 908b-1 of the cost function over time to generate a control parameter value 910b-0 as a function of time.

[0156] Figure 10 FIG. shows a schematic diagram of an extremum seeking (ES) controller 1000 for single-parameter tuning according to some embodiments. The ES controller 1000 injects a sinusoidal perturbation signal asin(ωt) 1002 to perturb a control parameter θ 1004. In response to perturbing the control parameter θ 1004, the ES controller 1000 determines a cost function Q(θ) 1006 regarding the performance of a closed-loop model. The ES controller 1000 multiplies the determined cost function Q(θ) 1006 by the sinusoidal perturbation signal asin(ωt) 1002 using a multiplier 1008. Additionally, the ES controller multiplies the resultant signal obtained from the multiplier 1008 by a gain l 1010 of the ES to form an estimate 1012 of the gradient of the cost function Q(θ) 1006. 1012 of the ES controller 1000 passes the estimated gradient 1012 through an integrator 1 / s 1014 to generate a parameter ξ 1016. The parameter ξ 1016 is added to the sinusoidal perturbation signal asin(ωt) 1002 using an adder 1018 to modulate the sinusoidal perturbation signal asin(ωt) 1002 for the next iteration. The iteration of the ES controller 1000 may be repeated until a termination condition is satisfied.

[0157] Figure 11 FIG. shows a schematic diagram of an extremum seeking (ES) controller 1100 for multi-parameter tuning according to some embodiments. Some embodiments are based on the recognition that the multi-parameter ES controller 1100 is derived from the single-parameter ES 1000. For example, the single-parameter ES controller 1000 may be replicated n times to obtain an n-parameter ES controller 1100. Some embodiments are based on the recognition that the n-parameter ES controller 1100 perturbs a set θ of n control parameters with corresponding n perturbation signals 1104-1 to 1104-n having n different frequencies. i1102 to update the optimal closed-loop model. In some embodiments, each of the n different frequencies is greater than the frequency response of system 102. Additionally, the n different frequencies of the n perturbation signals 1104-1 to 1104-n satisfy a convergence condition such that the sum of the first frequency of the first perturbation signal 1104-1 and the second frequency of the second perturbation signal 1104-2 in the set is not equal to the third frequency of the third perturbation signal 1104-3.

[0158] In addition, the n control parameters θ Figure 10 can be updated as described in the detailed description of i each of 1102. To this end, the n-parameter ES controller 1100 includes n control parameters θ i 1102, n perturbation signals 1104-1 to 1104-n, n estimated gradients 1108 n parameters ξ i 1110 and a common cost function Q(θ) 1106, where the common cost function Q(θ) 1106 is a function of all the estimated control parameters θ = (θ1, …, θ n ) T 1102. In some embodiments, the multi-parameter ES 1100 can be mathematically defined as:

[0159]

[0160] θ i = ξ i + a i sin(ω i t)

[0161] where the perturbation frequency ω i is such that ω i ≠ ω j , ω i + ω j ≠ ω k , i, j, k, ∈ {1, 2, …, n}, and ω i > ω * , where ω * is large enough to ensure convergence. In some embodiments, when the parameters a i , ω i and l are appropriately selected, the cost function Q(θ) 1106 converges to the neighborhood of the optimal cost function Q(θ * ).

[0162] To implement the multi-parameter ES controller 1100 in the real-time embedded system 102, a discrete version of the multi-parameter ES controller 1100 is advantageous. For example, the discrete version of the multi-parameter ES controller 1100 can be mathematically characterized as:

[0163] ξ i (k + 1) = ξ i (k) + a i lΔT sin(ω i k)Q(θ(k))

[0164] θ i (k + 1) = ξ i (k + 1) + a i sin(ω i (k))

[0165] Where k is the time step, and ΔT is the sampling time.

[0166] It should be understood that once the control parameter θ (i.e., the positive gain) is updated using the ES algorithm or Gaussian process-based optimization in the optimal closure model, the optimal closure model combined with the ODE simulates the actual behavior 602 of the system 102. For example, the estimated behavior 604 can be qualitatively and quantitatively similar to the actual behavior 602 without a gap 606.

[0167] To this end, the optimal reduced-order model 406 including the ODE and the optimal closure model with updated gains can be used to determine the control command. In some embodiments, the optimal reduced-order model 406 including the ODE and the optimal closure model with updated gains can develop a control strategy 106 for the system 102. The control strategy 106 can directly map the state of the system 102 among multiple states of the system 102 to the control command to control the operation of the coefficient 102. In the case where the system 102 is an HAVC system, examples of the control command include the position valve, the speed of the compressor, the parameters of the evaporator, etc. In the case where the system 102 is a rotor, examples of the control command include the speed of the rotor, the temperature of the motor, etc. In addition, the control command can be sent to the actuator of the system 102 via the output interface 218 to control the system 102. Some embodiments are based on the understanding that the operation of the system 102 is constrained. The constraints can include state constraints in the continuous state space of the system 102 and control input constraints in the continuous control input space of the system 102. In addition, the device 200 for controlling the constrained operation will be described in the Figure 12 detailed description.

[0168] Figure 12Illustrated is a predictive model-based algorithm 1200 for considering constraints on a control system 102 according to some embodiments. Some embodiments are based on the recognition that classical RL methods are not suitable for data-driven control of a constrained system 102. For example, classical RL methods do not consider state and input constraint satisfaction in a continuous state-action space; that is, classical RL methods cannot guarantee that the state of the controlled system 102 operated with control inputs satisfies state and input constraints throughout the operation. However, some embodiments are based on the understanding that RL methods allow the data-driven advantages of RL to be combined with model-based constraint optimization.

[0169] To this end, some embodiments use an RL-based model of the system 102 (e.g., the optimal reduced-order model 406) determined by data-driven adaptation in various predictive model-based algorithms. In some embodiments, an optimizer 1202 is formulated to consider constraints on the control system 102. Some embodiments are based on the recognition that the optimizer 1202 can be a model predictive control algorithm (MPC). MPC is a control method for controlling the system 102 while enforcing constraints. To this end, some embodiments utilize the advantage of MPC in considering constraints in the control of the system 102. Additionally, the Figures 13 to 15 real-time implementation of the device 200 of the control system 102 is described in the detailed description.

[0170] Figure 13 An exemplary real-time implementation of the device 200 for controlling the system 102 is illustrated, where the system 102 is an air conditioning system. In this example, a room 1300 has a door 1302 and at least one window 1304. The temperature and air flow in the room 1300 are controlled by the device 200 via the air conditioning system 102 through a ventilation unit 1306. A set of sensors 1308 are arranged in the room 1300, such as at least one air flow sensor 1308a for measuring the air flow velocity at a given point in the room 1300 and at least one temperature sensor 1308b for measuring the indoor temperature. Other types of setups can be considered, such as a room with multiple HVAC units, or a house with multiple rooms.

[0171] Some embodiments are based on the recognition that the air conditioning system 102 can be described by a physics-based model called the Boussinesq equation as Figure 3 exemplarily shown. However, the Boussinesq equation contains an infinite dimension to solve the Boussinesq equation for controlling the air conditioning system 102. To this end, a model including the ODE 402 and an updated closure model with updated gains is formulated to be in Figures 1 to 12The model described in the detailed description. This model optimally reproduces the dynamics (e.g., aerodynamics) of the air conditioning system 102. Additionally, in some embodiments, the aerodynamic model links the air flow values (e.g., air flow velocity) during operation of the air conditioning system 102 and the temperature of the air conditioned room 1300. To this end, the device 200 optimally controls the air conditioning system 102 to generate an air flow in a regulated manner.

[0172] Figure 14A An exemplary real-time implementation of the device 200 for controlling the system 102 is shown, where the system 102 is a vehicle 1400. The vehicle 1400 can be any type of wheeled vehicle, such as a passenger car, a bus, or an off-road vehicle. Additionally, the vehicle 1400 can be an autonomous vehicle or a semi-autonomous vehicle. For example, some embodiments control the movement of the vehicle 1400. Examples of movement include the lateral movement of the vehicle controlled by the steering system 1404 of the vehicle 1400. In one embodiment, the steering system 1404 is controlled by the controller 1402. Additionally or alternatively, the steering system 1404 can be controlled by the driver of the vehicle 1400.

[0173] In some embodiments, the vehicle can include an engine 1410, which can be controlled by the controller 1402 or other components of the vehicle 1400. In some embodiments, the vehicle can include an electric motor instead of the engine 1410 and can be controlled by the controller 1402 or by other components of the vehicle 1400. The vehicle can also include one or more sensors 1406 to sense the surrounding environment. Examples of the sensors 1406 include rangefinders, such as radar. In some embodiments, the vehicle 1400 includes one or more sensors 1408 to sense its current motion parameters and internal state. Examples of the one or more sensors 1408 include a global positioning system (GPS), an accelerometer, an inertial measurement unit, a gyroscope, an axle rotation sensor, a torque sensor, a deflection sensor, a pressure sensor, and a flow sensor. The sensors provide information to the controller 1402. The vehicle can be equipped with a transceiver 1412, which enables the communication ability of the controller 1402 to communicate with the device 200 in some embodiments through a wired or wireless communication channel. For example, through the transceiver 1412, the controller 1402 receives control commands from the device 200. Additionally, the controller 1402 outputs the received control commands to one or more actuators of the vehicle 1400, such as the vehicle's steering wheel and / or brakes, to control the movement of the vehicle.

[0174] Figure 14BA schematic diagram showing the interaction between a controller 1402 and a controller 1414 of a vehicle 1400 according to some embodiments is shown. For example, in some embodiments, the controller 1414 of the vehicle 1400 is a cruise control 1416 and an obstacle avoidance 1418 that control the rotation and acceleration of the vehicle 1400. In this case, the controller 1402 outputs control commands to the controllers 1416 and 1418 to control the kinematic state of the vehicle. In some embodiments, the controller 1414 further includes an advanced controller, for example, a lane keeping controller 1420 that further processes the control commands of the controller 1402. In both cases, the controller 1414 utilizes the output of the controller 1402 (i.e., the control commands) to control at least one actuator of the vehicle, such as the vehicle's steering wheel and / or brakes, to control the movement of the vehicle. In some embodiments, the movement of the vehicle 1400 can be constrained. Constraints are considered as described in the detailed description of Figure 12 . The constraints can include state constraints in the continuous state space of the vehicle 1400 and control input constraints in the continuous control input space of the vehicle 1400. In some embodiments, the state of the vehicle 1400 includes the position, orientation, and one or a combination of longitudinal speed and lateral speed of the vehicle 1400. The state constraints include one or a combination of speed constraints, lane keeping constraints, and obstacle avoidance constraints.

[0175] In some embodiments, the control input includes one or a combination of lateral acceleration, longitudinal acceleration, steering angle, engine torque, and braking torque. The control input constraints include one or a combination of steering angle constraints and acceleration constraints.

[0176] Figure 15 An exemplary real-time implementation of a device 200 for controlling a system 102 is shown, where the system 102 is an induction motor 1500. In this example, the induction motor 1500 is integrated with the device 200. The device is configured to control the operation of the induction motor 1500 as described in the detailed description of Figures 1 to 12 . In some embodiments, the operation of the induction motor 1500 can be constrained. The constraints include state constraints in the continuous state space of the induction motor 1500 and control input constraints in the continuous control input space of the induction motor 1500. In some embodiments, the state of the motor 1500 includes one or a combination of stator flux, line current, and rotor speed. The state constraints include constraints on the values of one or a combination of stator flux, line current, and rotor speed. In some embodiments, the control input includes the value of the excitation voltage. The control input constraints include constraints on the excitation voltage.

[0177] The above description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Instead, the following description of the exemplary embodiments will provide those skilled in the art with a sufficient description for implementing one or more exemplary embodiments. Various modifications can be envisioned in the functions and arrangements of the elements without departing from the spirit and scope of the disclosed subject matter as set forth in the appended claims.

[0178] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, the embodiments can be practiced without these specific details if understood by those of ordinary skill in the art. For example, systems, processes, and other elements in the disclosed subject matter can be shown in block diagram form as components to avoid obscuring the embodiments with unnecessary details. In other instances, well-known processes, structures, and techniques can be shown without unnecessary details to avoid obscuring the embodiments. Additionally, like reference numerals and designations in the various figures refer to like elements.

[0179] Additionally, the embodiments of the disclosed subject matter can be implemented at least in part manually or automatically. Execution or at least assistance for manual or automatic implementation can be performed by using a machine, hardware, software, firmware, middleware, microcode, a hardware description language, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments for performing the necessary tasks can be stored in a machine-readable medium. The processor can perform the necessary tasks.

[0180] Furthermore, the various methods or processes outlined herein can be encoded as software executable on one or more processors that employ any of a variety of operating systems or platforms. Additionally, such software can be written using any of a large number of suitable programming languages and / or programming tools or scripting tools, and can also be compiled into executable machine language code or intermediate code for execution on a framework or virtual machine. Generally, the functions of program modules can be combined or distributed as needed in the various embodiments.

[0181] The various methods or processes outlined herein can be encoded as software executable on one or more processors that employ any of a variety of operating systems or platforms. Additionally, such software can be written using any of a large number of suitable programming languages and / or programming tools or scripting tools, and can also be compiled into executable machine language code or intermediate code for execution on a framework or virtual machine. Generally, the functions of program modules can be combined or distributed as needed in the various embodiments.

[0182] Embodiments of the present disclosure can be embodied as a method, and examples of the method have been provided. The actions performed as part of the method can be ordered in any suitable way. Thus, embodiments can be constructed in which the actions are performed in an order different from the order illustrated, which can include performing some actions simultaneously, even if those actions are shown as sequential actions in the exemplary embodiments.

[0183] Although the present disclosure has been described with reference to certain preferred embodiments, it is to be understood that various other adaptations and modifications can be made within the spirit and scope of the present disclosure. Accordingly, one aspect of the appended claims is to cover all such variations and modifications that fall within the true spirit and scope of the present disclosure.

Claims

1. An apparatus (100, 200) configured to control the operation of a dynamic system (102) operating continuously in an engineering process and a machine, the apparatus comprising: An input interface (202) configured to receive a state trajectory (216) of the system as a sequence of states; A memory (206) configured to store a model (104, 208a) describing the dynamics of the system (102), the model including a combination of at least one differential equation (108a) and a closure model (108b); A processor (204) configured to: Update the closure model (108b) using reinforcement learning RL having a value function that reduces the difference between the shape of the received state trajectory (216) and the shape of the state trajectory (216) estimated using the model with the updated closure model, wherein the shape of the state trajectory (216) is a sequence of states of the system (102) as a function of time; and Determine a control command based on the model with the updated closure model; and An output interface (218) configured to send the control command to an actuator (220) of the system (102) to control the operation of the system (102); Wherein the differential equation of the model defines a reduced-order model of the system (102), the reduced-order model having a smaller number of parameters than a physical model of the system (102) according to a partial differential equation PDE, and wherein the reduced-order model is an ordinary differential equation ODE, and wherein the updated closure model is a non-linear function of the state of the system (102) that captures the difference in the behavior of the system (102) according to the ODE and the PDE; Wherein the partial differential equation PDE is the Boussinesq equation; Wherein the updated closure model includes a gain, and wherein the processor (204) is configured to determine a gain that reduces the error between the state of the system (102) estimated using the model with the updated closure model including the updated gain and the actual state of the system (102); Wherein the actual state of the system (102) is a measured state.

2. The apparatus (100, 200) according to claim 1, wherein, The processor (204) is configured to initialize the closure model (108b) with a linear function of the state of the system (102) and iteratively update the closure model (108b) with the RL until a termination condition is met.

3. The apparatus (100, 200) according to claim 1, wherein, The actual state of the system (102) is a state estimated using a partial differential equation PDE describing the dynamics of the system (102).

4. The apparatus (100, 200) according to claim 1, wherein, The processor (204) uses extremum seeking to update the gain.

5. The apparatus (100, 200) according to claim 1, wherein, The processor (204) uses Gaussian process-based optimization to update the gain.

6. The apparatus (100, 200) according to claim 1, wherein, The operation of the system (102) is constrained, wherein the RL updates the closure model (108b) without considering the constraint, and wherein the processor (204) uses the model with the updated closure model (108b) subject to the constraint to determine the control command.

7. The apparatus (100, 200) according to claim 6, wherein, The constraints include state constraints in the continuous state space of the system (102) and control input constraints in the continuous control input space of the system (102).

8. The apparatus (100, 200) according to claim 6, wherein, While implementing the constraints, the processor (204) uses model predictive control to determine the control command.

9. The apparatus (100, 200) according to claim 7, wherein, The system (102) is a vehicle controlled to perform one or a combination of lane keeping, cruise control, and obstacle avoidance operations, wherein the state of the vehicle includes the position, orientation, and one or a combination of longitudinal speed and lateral speed of the vehicle, wherein the control input includes one or a combination of lateral acceleration, longitudinal acceleration, steering angle, engine torque, and braking torque, wherein the state constraints include one or a combination of speed constraints, lane keeping constraints, and obstacle avoidance constraints, and wherein the control input constraints include one or a combination of steering angle constraints and acceleration constraints.

10. The apparatus (100, 200) according to claim 7, wherein, The system (102) is an induction motor controlled to perform a task, wherein the state of the motor includes one or a combination of stator flux, line current, and rotor speed, wherein the control input includes the value of the excitation voltage, wherein the state constraints include constraints on the values of one or a combination of the stator flux, the line current, and the rotor speed, wherein the control input constraints include constraints on the excitation voltage.

11. The device (100, 200) according to claim 1, wherein, The system (102) is an air conditioning system (102) that generates an air flow in a regulated environment, wherein the model is an aerodynamic model that links the flow rate value and the temperature value of the air regulated during the operation of the air conditioning system (102).

12. The device (100, 200) according to claim 1, wherein, The RL uses a trained neural network to minimize the value function.

13. A method for controlling the operation of a dynamic system (102) operating continuously in an engineering process and a machine, wherein, The method uses a processor (204) coupled to a memory that stores a dynamic model of the system (102) including a combination of at least one differential equation (108a) and a closure model (108b), the processor (204) being coupled to the stored instructions that, when executed by the processor (204), implement the steps of the method, the method including the steps of: Receiving a state trajectory (216) of the system (102) as a sequence of states; Using reinforcement learning RL to update the closure model, the reinforcement learning RL having a value function that reduces the difference between the shape of the received state trajectory and the shape of the state trajectory (216) estimated using the model with the updated closure model, wherein the shape of the state trajectory (216) is a sequence of states of the system (102) as a function of time; Determining a control command based on the model with the updated closure model; and Sending the control command to an actuator of the system to control the operation of the system (102); Among them, the differential equation of the model defines a reduced-order model of the system (102), the reduced-order model having a smaller number of parameters than the physical model of the system (102) according to the Boussinesq equation, where the Boussinesq equation is a partial differential equation (PDE), and where the reduced-order model is an ordinary differential equation (ODE), and where the updated closure model is a non-linear function of the state of the system (102) that captures the difference in the behavior of the system (102) according to the ODE and the PDE; Among them, the updated closure model includes a gain, and the method further includes the steps of: determining a gain that reduces the error between the state of the system estimated by the model using the updated closure model (108b) having the updated gain included therein and the actual state of the system (102); Among them, the actual state of the system (102) is the measured state.

14. The method according to claim 13, wherein, The operation of the system (102) is constrained, where the RL updates the closure model (108b) without considering the constraint, and where the method further includes the steps of: using the model with the updated closure model (108b) subject to the constraint to determine the control command.

Citation Information

Patent Citations

  • Unmanned surface ship optimal trajectory tracking control method based on reinforced learning method

    CN110018687A

  • System and Method for Controlling Operations of Air-Conditioning System

    US20190293314A1