Apparatus and method for controlling a system with uncertainty in dynamics
By combining the RCMDP framework with fuzzy set and Lyapunov descent method, the control problem of uncertain systems is solved, and the optimal control strategy under uncertainty and constraints is realized, ensuring the unified optimization of system performance and safety.
Patent Information
- Application Number
- CN202180070733.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-16
- Filing Date
- 2021-07-02
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2041-07-02
AI Technical Summary
Existing technologies struggle to effectively control system operation in systems with uncertainties, especially in the presence of dynamic and environmental uncertainties, and fail to meet safety and performance constraints.
We adopt the robust and constrained Markov decision process (RCMDP) framework, and combine robust MDP (RMDP) and constrained MDP (CMDP) to optimize performance and safety costs using fuzzy set and Lyapunov descent method. We design a controller to optimize system dynamics under uncertainty and constraints.
The optimal control strategy for the system under uncertainty and constraints was realized, ensuring that performance and cost are optimized while satisfying safety constraints, simplifying computational complexity and improving the robustness of the control system.
Smart Images

Figure CN116324635B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates generally to control of systems, and more specifically to devices and methods for controlling systems subject to constraints on system operation with uncertainty in their dynamics. BACKGROUND
[0002] In control of systems, a controller, which can be implemented using one or a combination of software or hardware, generates control commands for the system. The control commands direct the operation of the system as needed, for example, the operation follows a desired benchmark profile, or adjusts the output to a specific value. However, many real-world systems, such as autonomous vehicles and robots, need to satisfy constraints when deployed to ensure safe operation. Furthermore, real-world systems are often subject to influences such as non-stationarity, wear and tear, uncalibrated sensors, etc. Such influences lead to uncertainty in the dynamics of the system, and thus make the model of the system uncertain or unknown. In addition, there can be uncertainty in the environment of the system operation. Such uncertainty adversely affects the control of the system.
[0003] For example, a robotic arm manipulating different objects with different shapes and masses results in difficulty in designing an optimal controller for manipulating all objects. Similarly, it is difficult to design an optimal controller for a robotic arm manipulating a known object on different unknown surfaces with different contact geometries due to inherent switching between contact dynamics. Therefore, there is a need for a controller that is able to control a system with uncertainty in system operation. SUMMARY
[0004] It is an object of some embodiments to control a system subject to constraints while having uncertainty in system operation. The uncertainty in system operation can be due to uncertainty in the dynamics of the system, which can be caused by uncertainty in the values of parameters of the system, uncertainty in the environment of the system operation, or both. Therefore, the system can also be referred to as an “uncertain system”. In some embodiments, a dynamic model of the system includes at least one uncertain parameter. For example, a model of an arm of a robotic system moving an object can include uncertainty in the mass of the object carried by the arm. A model for movement of a train can include uncertainty about the friction of the train wheels with the track under current weather conditions.
[0005] Some embodiments are based on the principle of employing Markov Decision Processes (MDPs) to control objectives for constrained but uncertain systems. In other words, some embodiments employ MDPs to control constrained uncertain systems (systems). MDPs are discrete-time stochastic control processes that provide a framework for modeling decisions in situations where outcomes are partly random and partly under the control of a decision maker. MDPs are advantageous because they build on a formal framework that guarantees optimality in terms of expected cumulative cost while accounting for uncertain outcomes of actions. Furthermore, some embodiments are based on the understanding that MDPs are advantageous for a variety of different control scenarios including robotics and automation. To this end, it is an objective of some embodiments to extend MDPs to control systems subject to constraints and uncertainty.
[0006] Some embodiments are based on the recognition that MDPs can be extended to cover uncertainty in system operation in the context of a Robust MDP (RMDP). While MDPs aim to estimate control actions that optimize a cost (herein referred to as a performance cost), RMDPs aim to optimize the performance cost for different instances of system dynamics within a bound of uncertainty that defines operation of the system. For example, while the actual mass of an object carried by an arm of a robotic system can be unknown, a range of possible values that define a bound of uncertainty in operation of the robotic system can be known in advance. The system can have one or more uncertain parameters.
[0007] In many cases, RDMPs optimize the performance cost for the worst possible conditions demonstrated by the uncertainty in system operation. However, RMDPs are not suitable for constrained systems because optimization of the performance cost for the worst possible conditions can violate imposed constraints outside of the optimized performance cost.
[0008] Some embodiments are based on the recognition that MDPs can be extended to handle constraints in operation of the system in the context of a Constrained MDP (CMDP). CMDPs are designed to determine policies for sequential stochastic decision problems where multiple costs are considered simultaneously. Considering multiple costs allows constraints to be included within MDPs. For example, as described above, one optimization cost can be a performance cost while another cost can be a safety cost that manages satisfaction of constraints.
[0009] To this end, it is an object of some embodiments to combine RMDP and CMDP into a common framework of robust and constrained MDP (RCMDP). However, since some principles of MDP are common to both RMDP and CMDP, but some other principles are different and difficult to reconcile, it is challenging to produce a common framework (i.e., RCMDP) by combining RMDP and CMDP. For example, while RMDP and CMDP share many properties in their definitions, some differences can arise when computing the optimal policy. For the assumed model of a system without uncertainty, the optimal policy of CMDP is generally a randomized policy. Therefore, there is a need to consider the uncertainty of the system dynamics in the randomized policy formulation of CMDP in a way that is suitable for RMDP.
[0010] Some embodiments are based on the recognition that the uncertainty of the dynamics of a system can be transformed into the uncertainty of the state transitions of the system. In an MDP, the probability of the process entering its next state s' is influenced by the chosen action. Specifically, it is given by the state transition function P a (s, s'). Thus, the next state s' depends on the current state s and the action a of the decision maker. But the current state s and the action a of the decision maker are conditionally independent of the previous state and action. In other words, the state transitions of an MDP satisfy the Markov property.
[0011] Some embodiments represent the uncertainty of the dynamics of a system as a set of fuzzinesses P For example, the uncertainty of the system dynamics can be represented as a set of fuzzinesses P s,a , which is a set of feasible transition matrices defined for each state s E S and action a E A, i.e., a set of all possible uncertainty models of the system. In the following, P is used to indicate P s,a P s,a collectively for all states s and actions a. In other words, this modification with the set of fuzzinesses allows RMDP to consider the uncertainty of the system dynamics in the performance cost estimate, and allows CMDP to consider the uncertainty of the system dynamics in the constraint imposition (i.e., in the safety cost) in a way that is consistent with the performance cost estimate.
[0012] The performance cost and the safety cost are modified with the set of fuzzinesses, respectively. In particular, the set of fuzzinesses is incorporated into the performance cost to produce a robust performance cost, and into the safety cost to produce a robust safety cost, respectively. Thus, solving (or formulating) the RCMDP means optimizing the performance cost over the set of all possible uncertainty models of the system (set of fuzzinesses) subject to the safety cost that also needs to be satisfied over the set of all possible uncertainty models of the system. In other words, this modification with the set of fuzzinesses allows RMDP to consider the uncertainty of the system dynamics in the performance cost estimate, and allows CMDP to consider the uncertainty of the system dynamics in the constraint imposition (i.e., in the safety cost) in a way that is consistent with the performance cost estimate.
[0013] Accordingly, translating the system's dynamic uncertainty into uncertainty over state transitions of the system allows for unification of optimization of both performance cost and safety cost in a single consistent formulation, i.e., an RCMDP. Moreover, such translation is advantageous because the real or true state transitions, although unknown, are common to both performance cost and safety cost, and such formulation reinforces this consistency. To this end, the RCMDP formulation includes a set of ambiguities to optimize performance cost subject to optimization of safety cost that imposes constraints on system runs.
[0014] Accordingly, some embodiments use a joint multi-functional optimization of both performance cost and safety cost, where state transitions of each of the state and action pairs in performance cost and safety cost are represented by multiple state transitions that capture uncertainty of system runs. This joint optimization introduces interdependence of performance cost and safety cost.
[0015] Additionally, some embodiments perform an unbalanced joint multi-functional optimization, where optimization of performance cost is the primary objective, and optimization of safety cost is the secondary objective. Indeed, satisfying constraints is useless if the task is not performed. Accordingly, some embodiments define optimization of safety cost as a constraint on optimization of performance cost. In this way, optimization of safety cost becomes secondary to optimization of performance cost because safety cost as a constraint does not have an independent optimization objective and only limits the actions taken by the system to perform the task.
[0016] Some embodiments are based on the recognition that optimization of performance cost can benefit from the principle of minimax optimization, while optimization of safety cost can remain generic. Minimax is a decision rule for minimizing the maximum possible loss. In the context of RCMDP, minimax optimization aims to optimize performance cost for the worst-case scenario of dynamic uncertainty parameter values of the system. Because multiple state transitions that capture uncertainty of system runs are included in both performance cost and safety cost, actions that satisfy constraints for the same worst-case values of uncertainty parameters in the secondary optimization of safety cost, as determined by the primary minimax optimization of performance cost for the worst-case values of uncertainty parameters, can also satisfy constraints when real and true values of uncertainty parameters are more favorable for performance of the safety task.
[0017] In other words, if the computed control policy or control action minimizes the performance cost corresponding to the worst possible maximum cost of the set of possible uncertain models of the system, it minimizes the performance cost for any model of the system within the set of possible uncertain models of the system. Similarly, if the computed control policy satisfies the safety constraint bound for the worst possible safety cumulative maximum cost of the set of possible uncertain models of the system, it minimizes the safety cost for any model of the system within the set of possible uncertain models of the system.
[0018] Some embodiments are based on the recognition that constraints on the operation of the system can be imposed as hard constraints that prohibit their violation or soft constraints that discourage their violation. Some embodiments are based on the understanding that the optimization of the safety cost can be used as a soft constraint, which is acceptable for some control applications, but is prohibited in other applications. To this end, for some control applications, it is necessary to impose hard constraints on the operation of the system. In this case, some embodiments impose hard constraints on the optimization of the safety cost compared to imposing constraints on the optimization of the performance cost.
[0019] The performance of the task obtained by the operation of the system is designed. Therefore, the optimization of the performance cost should be constrained. This imposition can contradict the principle of the RMDP because the variables optimized by the optimization of the performance cost are independent of the constraints. In contrast, the optimization of the safety cost optimizes one or more variables that are related to the constraints. Therefore, it is easier to impose hard constraints on the variables that are related to the constraints.
[0020] To this end, in the RCMDP, the optimization of the performance cost is a minimax optimization that optimizes the performance cost for the worst-case scenario of the uncertain parameter values that lead to the uncertainty of the system dynamics, and the optimization of the safety cost optimizes the optimization variables that are subject to hard constraints.
[0021] Some embodiments are based on the recognition that, while the RCMDP is valuable in many robotic applications. However, its practical application remains challenging due to the computational complexity of the RCMDP. Because in many practical applications, the control policy computation of the RCMDP requires solving a constrained linear program with a large number of variables.
[0022] Some embodiments are based on the recognition that the RCMDP solution can be simplified by exploiting Lyapunov theory to render a Lyapunov function and show its decrease. This approach is referred to herein as the Lyapunov descent method. The Lyapunov descent method is advantageous because it allows to iteratively control the system while optimizing the control policy for controlling the system. In other words, the Lyapunov descent method allows to replace the determination of optimal and safe control actions before starting the control with a suboptimal but safe control action that can eventually (i.e., iteratively) converge to the optimal control. This replacement is possible due to the invariance set resulting from the Lyapunov descent method. To this end, the performance cost and the safety cost are optimized with the Lyapunov descent method.
[0023] Some embodiments are based on the recognition that designing a Lyapunov function and making it explicit greatly simplifies, clarifies, and unifies to some extent the convergence theory for optimization. However, designing a Lyapunov function for this constrained environment of RCMDP is challenging. To this end, some embodiments are based on designing a Lyapunov function from a computed auxiliary cost function such that it enforces satisfaction of the safety constraints defined by the safety cost at the current state while reducing the Lyapunov dynamics on the subsequent state transitions. Such an auxiliary cost function explicitly and constructively introduces Lyapunov arguments into the RCMDP framework without the need to solve the constrained control of uncertain systems in its entirety.
[0024] Thus, some embodiments interpret the safety cost with an auxiliary cost function that is configured to enforce satisfaction of the constraints at the current state, which together with the reduction of the Lyapunov dynamics via the Bellman operator on the subsequent state evolution imposed by the suboptimal control policy results in satisfaction of the safety constraints on all state evolutions of every suboptimal control policy. Iteration of this process of computing the auxiliary cost function and the associated suboptimal control policy eventually leads to an optimal control policy that satisfies the constraints.
[0025] According to one embodiment, the auxiliary cost function is a solution of a robust linear programming optimization problem that maximizes the value of the auxiliary cost function that maintains satisfaction of the safety constraints for all possible states of the system with uncertainty of the dynamics. According to an alternative embodiment, the auxiliary cost function is a weighted combination of basis functions with weights determined by a solution of a robust linear programming optimization problem. In some embodiments, the auxiliary cost function is a weighted combination of basis functions defining a deep neural network, the weights of the neural network being determined by a solution of a robust linear programming optimization problem.
[0026] Accordingly, one embodiment discloses a controller for controlling a system having uncertainty in its dynamics subject to constraints on the system's operation, the controller comprising: at least one processor; and a memory having instructions stored thereon that, when executed by the at least one processor, cause the controller to perform the following operations: obtaining historical data of the system's operation, the historical data comprising pairs of control actions and state transitions of the system controlled according to the corresponding control actions; determining, for the system in a current state, a current control action that transitions the system's state from the current state to a next state, wherein the current control action is determined according to a robust and constrained Markov decision process (RCMDP) that optimizes a performance cost of the system's operation using the historical data, the system being subject to an optimization of a safety cost that imposes the constraints on the operation, wherein the state transitions of each of the state and action pairs in the performance cost and the safety cost are represented by a plurality of state transitions that capture the uncertainty of the system dynamics; and controlling the system's operation according to the current control action to change the system's state from the current state to the next state.
[0027] Accordingly, one embodiment discloses a controller for controlling a system having uncertainty in its dynamics subject to constraints on the system's operation, the controller comprising: at least one processor; and a memory having instructions stored thereon that, when executed by the at least one processor, cause the controller to perform the following operations: obtaining historical data of the system's operation, the historical data comprising pairs of control actions and state transitions of the system controlled according to the corresponding control actions; determining, for the system in a current state, a current control action that transitions the system's state from the current state to a next state, wherein the current control action is determined according to a robust and constrained Markov decision process (RCMDP) that optimizes a performance cost of the system's operation using the historical data, the system being subject to an optimization of a safety cost that imposes the constraints on the operation, wherein the state transitions of each of the state and action pairs in the performance cost and the safety cost are represented by a plurality of state transitions that capture the uncertainty of the system dynamics; and controlling the system's operation according to the current control action to change the system's state from the current state to the next state. BRIEF DESCRIPTION OF DRAWINGS
[0028] [ Figure 1A ]
[0029] Figure 1A A schematic diagram illustrating the conception of a robust and constrained Markov decision process (RCMDP) according to some embodiments is shown.
[0030] [ Figure 1B]
[0031] Figure 1B A schematic diagram illustrating principles for considering uncertainty in system dynamics in a robust Markov decision process (RMDP) and a constrained Markov decision process (CMDP) consistent with principles of a Markov decision process (MDP) is shown in accordance with some embodiments.
[0032] [ Figure 1C ]
[0033] Figure 1C A schematic diagram illustrating a conception of an RCMDP including a set of ambiguities is shown in accordance with some embodiments.
[0034] [ Figure 2 ]
[0035] Figure 2 A block diagram of a controller for controlling a system having uncertainty in its dynamics subject to constraints on system operation is shown in accordance with some embodiments.
[0036] [ Figure 3 ]
[0037] Figure 3 A schematic diagram illustrating a set of ambiguities for designing is shown in accordance with some embodiments.
[0038] [ Figure 4 ]
[0039] Figure 4 A schematic diagram illustrating principles of a Lyapunov function is shown in accordance with some embodiments.
[0040] [ Figure 5A ]
[0041] Figure 5A A schematic diagram illustrating a Lyapunov descent-based solution for determining an optimal control policy for an RCMDP is shown in accordance with some embodiments.
[0042] [ Figure 5B ]
[0043] Figure 5B A robust safe policy iteration (RSPI) algorithm for determining an optimal control policy within a set of robust Lyapunov-inducing Markov stationary policies is shown in accordance with some embodiments.
[0044] [ Figure 5C ]
[0045] Figure 5C A robust safe value iteration (RSVI) algorithm for determining an optimal control policy within a set of robust Lyapunov-inducing Markov stationary policies is shown in accordance with some embodiments.
[0046] [ Figure 6 ]
[0047] Figure 6 A schematic diagram illustrating a method for determining an auxiliary cost function according to one embodiment is shown.
[0048] [ Figure 7 ]
[0049] Figure 7 A schematic diagram illustrating a method for determining an auxiliary cost function based on a basis function according to one embodiment is shown.
[0050] [ Figure 8 ]
[0051] Figure 8 A schematic diagram illustrating a robot system integrated with a controller for performing operations according to some embodiments is shown.
[0052] [ Figure 9A ]
[0053] Figure 9A A schematic diagram illustrating a vehicle system including a vehicle controller in communication with a controller employing principles of some embodiments is shown.
[0054] [ Figure 9B ]
[0055] Figure 9B A schematic diagram illustrating interaction between a vehicle controller and other controllers of a vehicle system according to some embodiments is shown.
[0056] [ Figure 9C ]
[0057] Figure 9C A schematic diagram illustrating an autonomous or semi-autonomous controlled vehicle for which control actions are generated using some embodiments is shown.
[0058] [ Figure 10 ]
[0059] Figure 10 A schematic diagram illustrating characteristics of a CMDP-based reinforcement learning (RL) method, a RMDP-based RL method, and a Lyapunov-based robust constrained MDP (L-RCMDP) based RL is shown.
[0060] [ Figure 11 ]
[0061] Figure 11 A schematic diagram illustrating an overview of the RCMDP conceit according to some embodiments is shown. DETAILED DESCRIPTION
[0062] In the following description, for the purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding. It will be apparent, however, to one skilled in the art that the present disclosure can be practiced without these specific details. In other instances, devices and methods are shown in block diagram form in order to avoid obscuring the present disclosure.
[0063] As used in the specification and claims, the terms "for example," "e.g.," "for instance," and "such as," and the verbs "comprising," "having," "including," and their other verb forms, when used in the specification and / or claims, are not used to limit the scope of the present disclosure. The term "based on" means at least partially based on.
[0064] Figure 1A A schematic diagram showing the conception of a robust and constrained Markov decision process (RCMDP) is shown in accordance with some embodiments. It is an objective of some embodiments to control a system 100 that is subject to constraints while having system operation uncertainty. The uncertainty in the operation of the system 100 can be due to dynamic uncertainty of the system 100, which can be caused by uncertainty in parameter values of the system 100, uncertainty in the environment in which the system 100 operates, or both. Thus, the system 100 can also be referred to as an "uncertain system." In some embodiments, the dynamic model of the system 100 includes at least one uncertainty parameter. For example, a model of an arm of a robot system that moves an object can include uncertainty in the mass of the object carried by the arm. A model for train movement can include uncertainty about the friction of train wheels with the track under current weather conditions.
[0065] Some embodiments are based on the objective of controlling a constrained but uncertain system using the principles of a Markov decision process (MDP) 102. In other words, some embodiments use an MDP 102 to control an uncertain system (system 100) that is subject to constraints. An MDP 102 is a discrete-time stochastic control process that provides a framework for modeling decisions in situations where outcomes are partly random and partly under the control of a decision maker. An MDP is advantageous because it builds on a form framework that guarantees optimality in terms of expected cumulative cost while taking into account uncertain outcomes of actions. Furthermore, some embodiments are based on the understanding that an MDP 102 is advantageous for a variety of different control scenarios, including robotics and automated control. To this end, it is an objective of some embodiments to extend an MDP 102 to control a system 100 that is subject to constraints.
[0066] Some embodiments are based on the recognition that MDP 102 can be extended to cover the uncertainty of the operation of system 100 in the context of a robust MDP (RMDP) 104. While MDP 102 aims to estimate a control action that optimizes a cost 106 (herein referred to as a performance cost), RMDP 104 aims to optimize performance cost 106 for different instances of the dynamics of system 100 within the bounds of the uncertainty of the operation of system 100. For example, while the actual mass of an object carried by an arm of a robotic system can be unknown, a range of possible values defining the bounds of the uncertainty of the operation of the robotic system can be known in advance. System 100 can have one or more uncertain parameters.
[0067] In many cases, RDMP 104 optimizes performance cost 106 for the worst possible conditions demonstrated by the uncertainty of the operation of system 100. However, RMDP 104 is not suitable for constrained systems because the optimization of performance cost 106 for the worst possible conditions can violate imposed constraints outside of the optimized performance cost.
[0068] Some embodiments are based on the recognition that MDP 102 can be extended for handling the constraints of the operation of system 100 in the context of a constrained MDP (CMDP) 108. CMDP 108 is designed to determine a policy for a sequential stochastic decision problem in which multiple costs are considered simultaneously. Considering multiple costs allows the inclusion of constraints of MDP 102. For example, as described above, one optimization cost can be performance cost 106, while another cost can be safety cost 110 that manages constraint satisfaction.
[0069] To this end, it is an objective of some embodiments to combine RMDP 104 and CMDP 108 into a common framework of a robust and constrained MDP (RCMDP) 112. However, since many of the principles of MDP 102 are common to both RMDP 104 and CMDP 108, but many other principles are different and difficult to reconcile, it is challenging to produce a common architecture (i.e., RCMDP 112 by combining RMDP 104 and CMDP 108). For example, while RMDP and CMDP share many properties in their definitions, some differences can arise when computing an optimal policy. For an assumed model of system 100 without uncertainty, the optimal policy of CMDP 108 is generally a randomized policy. Therefore, it is necessary to consider the uncertainty of the dynamics of system 100 in the randomized policy formulation of CMDP 108 in a manner suitable for RMDP.
[0070] Figure 1BA schematic diagram showing principles of considering the dynamics' uncertainty 114 of the system 100 in the RMDP 104 and the CMDP 108 in line with the principles of MDPs, according to some embodiments, is shown. Some embodiments are based on the insight that the dynamics' uncertainty 114 of the system 100 can be transformed 116 into an uncertainty of state transitions 118 of the system 100. In MDPs, the probability of a process entering its next state s' is influenced by the chosen action. Specifically, it is given by the state transition function P a (s, s'). Thus, the next state s' depends on the current state s and the decision maker's action a. But the current state s and the decision maker's action a are conditionally independent of previous states and actions. In other words, the state transitions of the MDP 102 satisfy the Markov property.
[0071] Some embodiments represent the dynamics' uncertainty 114 of the system 100 as a set of transition probabilities For example, the dynamics' uncertainty 114 of the system 100 can be represented as a set of fuzziness P s,a s,a , which is a set of feasible transition matrices defined for each state s e S and action a e A, i.e. a set of all possible uncertainty models of the system 100. In the following, P is used to refer to P s,a collectively for all states s and actions a.
[0072] Figure 1C A schematic diagram showing the conception of the RCMDP 112 including the set of fuzziness P 120, according to some embodiments, is shown. The performance cost 106 and the safety cost 110 are modified with the set of fuzziness 120, respectively. In particular, the set of fuzziness 120 is incorporated into the performance cost 106 to yield a robust performance cost, and into the safety cost 110 to yield a robust safety cost, respectively. Thus, solving (or conceiving) the RCMDP 112 means an optimization of the performance cost 106 over the set of all possible uncertainty models of the system 100 (the set of fuzziness 120), subject to the safety cost 110 that also needs to be satisfied over the set of all possible uncertainty models of the system 100. In other words, this modification with the set of fuzziness 120 allows the RMDP 104 to consider the dynamics' uncertainty 114 of the system 100 in the performance cost estimate, and the CMDP 108 to consider the dynamics' uncertainty 114 of the system 100 in the constraint enforcement (i.e. in the safety cost 110) in a way that is consistent with the performance cost estimate.
[0073] Accordingly, converting 116 the dynamics uncertainty 114 of the system 100 into an uncertainty of state transitions of the system 100 allows for unification of optimization of both the performance cost 106 and the safety cost 110 in a single consistent formulation, i.e., the RCMDP. Moreover, such conversion 116 is advantageous because the real or true state transitions, although unknown, are common to both the performance cost 106 and the safety cost 110, and such conversion reinforces this consistency. To this end, the RCMDP 112 formulation includes a set of ambiguities 120 to optimize the performance cost 106 subject to optimization of the safety cost 110 that imposes constraints on the operation of the system 100.
[0074] Accordingly, some embodiments use a joint multi-functional optimization of both the performance cost 106 and the safety cost 110, where the state transitions of each of the state and action pairs in the performance cost 106 and the safety cost 110 are represented by a plurality of state transitions that capture the uncertainty of the operation of the system 100. This joint optimization introduces interdependencies of the performance cost 106 and the safety cost 110.
[0075] Additionally, some embodiments perform an unbalanced joint multi-functional optimization, where optimization of the performance cost 106 is the primary objective, and optimization of the safety cost 110 is the secondary objective. Indeed, satisfying the constraints is useless if the task is not performed. Accordingly, some embodiments define optimization of the safety cost 110 as a constraint on optimization of the performance cost 110. In this way, optimization of the safety cost becomes subordinate to optimization of the performance cost, because the safety cost as a constraint does not have an independent optimization objective and only limits the actions taken by the system 100 to perform the task.
[0076] Moreover, some embodiments determine the current control action 122 of the system 100 according to the RCMDP 112. Specifically, the RCMDP 112 optimizes the performance cost 106 subject to optimization of the safety cost 110 that imposes constraints on the operation of the system 100 to determine the current control action 122.
[0077] Figure 2 A block diagram of a controller 200 for controlling a system 100 having uncertainty in its dynamics subject to constraints on the operation of the system 100 is shown in accordance with some embodiments. The controller 200 is connected to the system 100. The system 100 can be a robotic system, an autonomous vehicle system, a heating, ventilation, and air conditioning (HVAC) system, etc. The controller 200 is configured to acquire, via an input interface 202, historical data of the operation of the system 100, which includes pairs of control actions and state transitions of the system 100 controlled according to the respective control actions.
[0078] The controller 200 can have a plurality of interfaces that connect the controller 200 with other systems and devices. For example, a network interface controller (NIC) 214 is adapted to connect the controller 200 to a network 216 through the bus 212. Through the network 216, the controller 200 obtains historical data 218 of the operation of the system 100, including pairs of control actions or state transitions of the system 100 controlled according to the corresponding control actions.
[0079] The controller 200 includes a processor 204 configured to execute stored instructions, and a memory 206 that stores instructions executable by the processor 204. The processor 204 can be a single core processor, a multi-core processor, a computing cluster, or any number of other configurations. The memory 206 can include random access memory (RAM), read only memory (ROM), flash memory, or any other suitable memory systems. The processor 204 is connected through the bus 212 to one or more input and output devices. In addition, the controller 200 includes a storage device 208 adapted to store different modules that store executable instructions for the processor 204. The storage device 208 can be implemented using a hard disk drive, an optical drive, a thumb drive, an array of drives, or any combination thereof. The storage device 208 is configured to store a set of ambiguities 210 of the RCMDP construct. The set of ambiguities 210 includes a set of all possible uncertain models of the system 100.
[0080] In some implementations, the controller 200 is configured to determine, for the system 100 in a current state, a current control action that transitions the state of the system 100 from the current state to a next state, where the current control action is determined according to the RCMDP that optimizes a performance cost of the operation of the system subject to an optimization of a safety cost that enforces a constraint on the operation. The state transitions of each state and action pair in the performance cost and the safety cost are represented by a plurality of state transitions that capture the uncertainty of the dynamics of the system 100. The controller 200 is further configured to control the operation of the system 100 according to the current control action to change the state of the system 100 from the current state to the next state.
[0081] Additionally, the controller 200 can include an output interface 220. In some implementations, the controller 200 is further configured to submit to a controller of the system 100, via the output interface 220, to operate the system 100 according to the current control action.
[0082] Mathematical formulation of RCMDP
[0083] Consider an RMDP model with a finite number of states S = {1,..., S} and a finite number of actions A = {1,..., A}. Each action a e A is available to the decision maker to take in each state s e S. After taking action a e A in state s e S, the decision maker receives a cost c(s, a) e R and transitions to the next state s' according to true but unknown transition probabilities A set of ambiguities P is defined for each state s e S and action a e A, a set of feasible transition matrices, i.e., a set of all possible uncertain models of the system 100. P is used to cumulatively indicate P s,a .
[0084] Figure 3 A schematic diagram for designing a set of ambiguities is shown in accordance with some embodiments. In one embodiment, a s, a-rectangular ambiguity set is used, which assumes independence between different state-action pairs. A data set D 300 of runs of the system (e.g., system 100) is used to determine the ambiguity set. The data set D 300 can include pairs of control actions and state transitions of the system. In addition, the controller 200 computes a mean of the data set D and uses an L1-norm 302 to define a set of ambiguities 306 around the mean. Specifically, an L1-norm bounded ambiguity set around the nominal transition probabilities 304 on the data set 300 is used to define the ambiguity set 306 as follows:
[0085]
[0086] where ψ s,a ≥ 0 is a budget for allowed deviation. Such a budget can be computed using the Hoeffding bound as:
[0087]
[0088] where n s,a is the number of transitions in the data set D originating from state s and action a, and d is a confidence level.
[0089] In different embodiments, different norms are used to design the ambiguity set 306. For example, one embodiment can use the L2 norm. In another embodiment, the ambiguity set 306 can be designed using the L0 norm. In some other embodiments, the ambiguity set 306 can be designed using the L ∞ norm.
[0090] Alternatively, in some embodiments, the ambiguity set 306 can be defined using data-driven and confidence regions. In another alternative embodiment, the ambiguity set 306 can be defined using a level of likelihood of the probability distribution of the dataset.
[0091] A stationary randomized policy π(· | s) for a state s e S defines a probability distribution over actions a e A, and Π is a set of stationary randomized policies. The robust return g θ for a sampled trajectory ξ and an ambiguity set 306 (P) is defined as:
[0092]
[0093] where ξ = [s0, a0,...]. The random variable g θ (ξ, θ) is defined as the robust value function for that state:
[0094] Furthermore, to accommodate safety constraints, a CMDP is used. Here, the RMDP model is extended by introducing an additional immediate safety constraint cost d(s) e [0, D max ] and an associated constraint budget d0 e R + or safety margin as an upper bound on the expected cumulative constraint cost. For a policy θ, the total robust constraint return h θ for a sampled trajectory ξ and an ambiguity set P is defined as:
[0095]
[0096] where ξ = [s0, a0,...]. The random variable h θ (ξ, θ) is defined as the constraint value function for that state:
[0097] Thus, for an initial state distribution p0 e Δ S the robust return in terms of value functions is defined as: and the robust return for constraint costs, i.e., the robust safety cost, is defined as:
[0098]
[0099] Some implementations are based on the understanding that performance cost optimization can benefit from the principle of minimax optimization, while safety cost optimization can remain generalized. Minimax is a decision rule used to minimize the possible losses in the worst-case (maximum loss) scenario. In the context of RCMDP, minimax optimization aims to optimize the performance cost of the worst-case scenario for the values of dynamic uncertain parameters of system 100. Because multiple state transitions that capture the uncertainty of the operation of system 100 are included in both performance cost and safety cost, actions that satisfy the same worst-case value constraint on the uncertain parameter in the subordinate optimization of safety cost, determined by primary minimax optimization of performance cost for the worst-case value of the uncertain parameter, can also satisfy the constraint when the actual and true values of the uncertain parameter are more favorable to the performance of the safety task.
[0100] Therefore, some implementations formulate the following RCMDP problem:
[0101]
[0102] In other words, some implementations aim to solve the RCMDP problem (1), that is, to optimize the performance cost over the set of all possible uncertain models of system 100 under security constraint 206, which also needs to be satisfied over the set of all possible uncertain models of system 100. According to one implementation, this is ensured by exerting effort on the worst-case performance cost and security cost over the set of all possible uncertain models of system 100. If the calculated control strategy achieves this, then the worst-case performance cost over the set of all possible uncertain models of system 100 is satisfied. Corresponding performance cost Minimizing this minimizes the performance cost of any model of system^100 within the set of possible uncertain models of system^100. Similarly, if the calculated control policy satisfies the worst-case safety cost for the set of possible uncertain models of system^100. Safety cost of safety constraint limit d0 It minimizes the security cost of any model of system 100 within the set of possible uncertain models of system 100.
[0103] Some embodiments are based on the recognition that constraints on the operation of the system 100 can be imposed as hard constraints that prohibit their violation or soft constraints that discourage their violation. Some embodiments are based on the understanding that optimization of safety cost can be used as a soft constraint, which is acceptable for some control applications, but is prohibited in other applications. To this end, for some control applications, it is necessary to impose hard constraints on the operation of the system 100. In this case, some embodiments impose hard constraints on the optimization of safety cost, as opposed to imposing constraints on the optimization of performance cost.
[0104] The performance of the task obtained by the operation of the system 100 is designed to be constrained. Therefore, the optimization of the performance cost should be constrained. This imposition can contradict the principle of RMDP, since the variable optimized by the optimization of the performance cost is not related to the constraint. In contrast, the optimization of the safety cost optimizes one or more variables in relation to the constraint. Therefore, it is easier to impose hard constraints on the variables related to the constraint.
[0105] To this end, in the RCMDP problem (1), the optimization of the performance cost is a minimax optimization that optimizes the performance cost for the worst-case scenario of the uncertain parameter values that lead to the uncertainty of the dynamics of the system 100, and the optimization of the safety cost optimizes the optimization variables subject to the hard constraints.
[0106] Some embodiments are based on the recognition that while the RCMDP formulation (1) is a valuable insight in many robotic applications. However, its practical application remains challenging due to its computational complexity. Because in many practical applications, the control policy computation of the RCMDP requires solving a constrained linear problem with a large number of variables.
[0107] Some embodiments are based on the recognition that the RCMDP solution can be simplified by presenting Lyapunov functions using Lyapunov theory. Figure 4A schematic diagram illustrating the principle of Lyapunov functions according to some embodiments is shown. For a system to be controlled (e.g., system 100) 400, Lyapunov theory allows to design a Lyapunov function 402 for this system. In particular, Lyapunov theory allows to design a positive definite function, e.g., an energy function of the system. Furthermore, it is checked whether the Lyapunov function decreases over time 404, e.g., by testing the time derivative of the Lyapunov function. If the Lyapunov function decreases over time, it can be concluded that the trajectory of the system is bounded 408. If the Lyapunov function does not decrease, the boundedness of the system trajectory is not guaranteed 406. Thus, some embodiments simplify RCMDP by exploiting Lyapunov theory by presenting a Lyapunov function and showing its decrease. This approach is called Lyapunov descent.
[0108] In addition, Lyapunov descent is advantageous because it allows to iteratively control the system while optimizing the control policy for controlling the system. In other words, Lyapunov descent allows to replace the determination of optimal and safe control actions before starting the control with suboptimal but safe control actions that can eventually (i.e., iteratively) converge to the optimal control. This replacement is possible due to the invariance set resulting from Lyapunov descent. To this end, the performance cost and the safety cost are optimized with Lyapunov descent.
[0109] Figure 5A A schematic diagram illustrating a Lyapunov descent based solution for the RCMDP problem (1) for determining an optimal control policy according to some embodiments is shown. Some embodiments are based on the insight that designing and using a Lyapunov function simplifies and unifies the convergence theory for optimization. However, it is challenging to design a Lyapunov function for the constrained environment of RCMDP.
[0110] To this end, some embodiments design a Lyapunov function 504 based on an auxiliary cost function 500. The auxiliary cost function 500 is configured to enforce that the constraints defined by the safety cost 502 are satisfied at the current state while causing the Lyapunov function to decrease along the dynamics of the system 100 on subsequent evolutions of the state transition. Thus, the safety cost 502 is interpreted with the auxiliary cost function 500. The auxiliary cost function 500 explicitly and constructively introduces a Lyapunov variable into the RCMDP formulation (1) without the need to solve the constrained control of uncertain systems in total.
[0111] Thus, for the RCMDP problem given by formulation (1), the Lyapunov function 504 can be given as
[0112] a.
[0113] where f is the auxiliary cost function 500. The Lyapunov function (2) is related to the auxiliary cost function f.
[0114] Further, to determine the optimal control policy based on the Lyapunov function (2), the controller 200 computes a set of robust Lyapunov-induced Markov stationary policies 506. The set of robust Lyapunov-induced Markov stationary policies is defined as
[0115] a.
[0116] where, is the Bellman operator with respect to a policy p from the set of Markov stationary policies, for the robust cost d max is defined as
[0117] i.
[0118] and is defined as
[0119] i.
[0120] where Ξ is the set of initial states. The Bellman operator satisfies a contraction property, which can be written as
[0121] i.
[0122] Thus,
[0123]
[0124] According to the Lyapunov function (2), a feasible solution to the RCMDP problem given by equation (1) can be given as
[0125] 1.
[0126] Equation (3) implies that any control policy computed according to the set of robust Lyapunov-induced Markov stationary policies is a robustly safe policy for the system to be controlled (e.g., system 100).
[0127] Further, the controller 200 determines an optimal control policy within the set of robust Lyapunov-induced Markov stationary policies 508. In one embodiment, the robust Lyapunov-induced Markov stationary policies within the set of robust Lyapunov-induced Markov stationary policies 508 are determined using a robust safe policy iteration (RSPI) algorithm.
[0128] Figure 5BA robust safe policy iteration (RSPI) algorithm for determining an optimal control policy is shown in accordance with some embodiments.
[0129] The RSPI algorithm starts with a feasible but suboptimal control policy π0. Subsequently, the associated robust Lyapunov function is computed. Next, the associated robust cost function c max. is computed. The corresponding robust cost value function is then computed as Furthermore, an intermediate control policy is obtained within the set of robust Lyapunov-induced Markov stationary policies. This process is repeated until a predetermined number of iterations is reached or until the intermediate control policy converges to a stable optimal control policy π * .
[0130] In an alternative embodiment, a robust safe value iteration (RSVI) algorithm is used to determine an optimal control policy within the set of robust Lyapunov-induced Markov stationary policies 508.
[0131] Figure 5C A robust safe value iteration (RSVI) algorithm for determining an optimal control policy is shown in accordance with some embodiments.
[0132] The RSVI algorithm starts with a feasible but suboptimal control policy π0. Subsequently, the associated robust Lyapunov function is computed. Next, the associated robust cost function c max. is computed. The corresponding value function Q k+1. is computed for the associated robust Lyapunov-induced Markov stationary policies. Furthermore, an intermediate control policy is obtained within the set of robust Lyapunov-induced Markov stationary policies. This process is repeated until a predetermined number of iterations is reached or until the control policy converges to a stable optimal control policy π * .
[0133] Figure 6 A schematic diagram for determining an auxiliary cost function 500 is shown in accordance with one embodiment. A robust linear programming optimization problem 600 is solved 602 by the controller 200 to determine the auxiliary cost function 500. The robust linear programming optimization problem 600 is given by
[0134]
[0135] The auxiliary cost function 500 is the solution of the robust linear programming optimization problem 600 given by equation (4), where L f is given by equation (2). The robust linear programming optimization problem 600 maximizes the value of the auxiliary cost function that maintains satisfaction of the safety constraint for all possible states of the system with uncertainty of the dynamics of the system to determine the auxiliary cost function 500.
[0136] Figure 7 An illustration of determining an auxiliary cost function 500 based on basis functions is shown, according to one embodiment. Some embodiments are based on the recognition that a combination of basis functions 700 and optimal weights 706 associated with the basis functions can be used to determine an auxiliary cost function 500. Specifically, the auxiliary cost function 500 is determined using a basis function approximation of The basis function approximation of
[0137]
[0138] where φ i is a basis function 700, and is an optimal weight associated with the basis function φ i . The optimal weight 706 is denoted as According to one embodiment, the optimal weight 706 is computed by solving a robust linear programming optimization problem 704 given by
[0139]
[0140] Thus, the auxiliary cost function 500 is a weighted combination of basis functions with weights determined by the solution of the robust linear programming optimization problem 704.
[0141] Alternatively, in other embodiments, the basis function approximation of can be implemented by a deep neural network (DNN) model. The weights of the deep neural network can be determined by solving a robust linear programming optimization problem.
[0142] The DNN model is used to represent a mapping between a state s and a value of the auxiliary cost function at the state s. The DNN can be any deep neural network, such as a fully connected network, a convolutional network, a residual network, etc. The DNN model is trained by solving an optimal problem given by equation (5) to obtain optimal coefficients of the DNN model, thereby obtaining an approximation of the auxiliary cost function.
[0143] Figure 8 A robotic system 800 integrated with a controller 200 for performing an operation is shown in accordance with some embodiments. A robotic arm 802 is configured to perform an operation including picking up an object 804 of a particular shape while maneuvering between obstacles 806a and 806b. Here, the robotic arm 802 is the system to be controlled, the task of picking up the object 804 is the performance task, and obstacle avoidance is the safety task. In other words, the objective of some embodiments is to control the robotic arm 802 (system) to pick up the object 804 (performance task) while avoiding the obstacles 806a and 806b (safety task). The models of the object 804 or the obstacles 806a and 806b or the robotic arm 802 can be unknown because the model of the robot can be uncertain due to aging and failures (in other words, the dynamics of the robotic arm 802 are uncertain).
[0144] The controller 200 obtains historical data of the operation of the robotic arm 802. The historical data can include pairs of control actions and state transitions of the robotic arm 802 controlled according to the respective control actions. The robotic arm 802 is in a current state. The controller 200 can determine a current control action or control policy in accordance with the RCMDP given by equation (1). The RCMDP given by equation (1) uses the historical data to optimize a performance cost of the task of picking up the object 804 subject to a safety cost imposed on the task of picking up the object 804 (obstacle avoidance). The state transitions of each state and action pair in the performance cost and the safety cost are represented by a plurality of state transitions that capture the uncertainty of the dynamics of the robotic arm 802.
[0145] The controller 200 controls the task of picking up the object 804 in accordance with the determined current control action or control policy to change the state of the system from the current state to a next state. To this end, during the operation of the robotic system 800, the controller 200 ensures that the obstacles 806a and 806b are not hit while picking up the object 804 regardless of the uncertainty on the object 804 or the obstacles 806a and 806b or the robotic arm 802.
[0146] Figure 9A A schematic diagram of a vehicle system 900 including a vehicle controller 902 in communication with a controller 200 that employs the principles of some embodiments is shown. The vehicle 900 can be any type of wheeled vehicle, such as a passenger car, a bus, or a flow station. Further, the vehicle 900 can be an autonomous or semi-autonomous vehicle. For example, some embodiments control the motion of the vehicle 900. Examples of motion include lateral motion of the vehicle controlled by a steering system 904 of the vehicle 900. In one embodiment, the steering system 904 is controlled by the vehicle controller 902. Additionally or alternatively, the steering system 904 can be controlled by a driver of the vehicle 900.
[0147] In some embodiments, the vehicle 900 can include an engine 910, which can be controlled by the vehicle controller 902 or by other components of the vehicle 900. In some embodiments, the vehicle 900 can include an electric motor instead of the engine 910 and can be controlled by the vehicle controller 902 or by other components of the vehicle 900. The vehicle 900 can also include one or more sensors 906 to sense the surrounding environment. Examples of the sensors 906 include distance finders, such as radar. In some embodiments, the vehicle 900 includes one or more sensors 908 to sense its current motion parameters and internal states. Examples of the one or more sensors 908 include a global positioning system (GPS), an accelerometer, an inertial measurement unit, a gyroscope, an axle rotation sensor, a torque sensor, a deflection sensor, a pressure sensor, and a flow sensor. The sensors provide information to the vehicle controller 902. The vehicle 900 can be equipped with a transceiver 910 that enables the vehicle controller 902 to have the ability to communicate with the system 200 of some embodiments through wired or wireless communication channels. For example, through the transceiver 910, the vehicle controller 902 receives control actions from the controller 200.
[0148] Figure 9B A schematic diagram showing the interaction between the vehicle controller 902 of the vehicle 900 and other controllers 912 according to some embodiments is shown. For example, in some embodiments, the controllers 912 of the vehicle 900 are a steering controller 914 and a brake / throttle controller 916 that control the rotation and acceleration of the vehicle 900. In this case, the vehicle controller 902 outputs control commands to the controllers 914 and 916 based on the control actions to control the motion state of the vehicle 900. In some embodiments, the controllers 912 also include high-level controllers, such as a lane-keeping assist controller 918, which further processes the control commands of the vehicle controller 902. In both cases, the controllers 912 utilize the output (i.e., control commands) of the vehicle controller 902 to control at least one actuator of the vehicle 900, such as the steering wheel and / or the brakes of the vehicle 900, in order to control the motion of the vehicle 900.
[0149] Figure 9CA schematic diagram showing an autonomous or semi-autonomous controlled vehicle 920 for which control actions are generated using some embodiments is shown. The controlled vehicle 920 can be equipped with a controller 200. The controller 200 controls the controlled vehicle 920 to keep the controlled vehicle 920 within certain limits of a road 924 and to avoid other non-controlled vehicles, i.e., obstacles 922, for the controlled vehicle 920. For such control, the controller 200 determines control actions according to the RCMDP. In some embodiments, the control actions include commands that specify values for one or a combination of steering angle of wheels of the controlled vehicle 920, rotational speed of the wheels, and acceleration of the controlled vehicle 920. Based on the control actions, the controlled vehicle 920 can, for example, pass another vehicle on the left side 926 or on the right side without hitting the vehicle 926 and the vehicle 922 (obstacle).
[0150] In addition, the RCMDP given by equation (1) can be used for policy transfer from simulation to the real world (Sim2Real). As in practical applications, to mitigate the sample inefficiency of model-free reinforcement learning (RL) algorithms, it is common to train against a simulated environment. The results are then transferred to the real world, often followed by fine-tuning, a process known as Sim2Real. In safety-critical applications, using RCMDP (equation (1)) for policy transfer from simulation to the real world (Sim2Real) can yield benefits in terms of performance and safety guarantees that are robust to model uncertainty.
[0151] Figure 10 A schematic diagram showing characteristics of a CMDP-based RL 1000 method, a RMDP-based RL 1002 method, and a Lyapunov-based robust constrained MDP (L-RCMDP) based RL 1004 is shown. A list of characteristics exhibited by the CMDP-based RL method 1000 is shown in block 1006. The characteristics of the CMDP-based RL method 1000 include, for example, performance cost, safety constraints, given or learned accurate model. However, the CMDP-based RL method 1000 does not exhibit robustness in performance, does not exhibit robustness in safety. A list of characteristics of the RMDP-based RL method 1002 is shown in block 1008. The characteristics of the RMDP-based RL method 1002 include, for example, performance cost, no safety constraints, given or learned uncertain model, and robustness in performance.
[0152] A list of properties exhibited by L-RCMDP-based RL 1004 is shown in block 1010. L-RCMDP-based RL 1004 can correspond to the RCMDP problem given by equation (1). Properties of L-RCMDP-based RL 1004 include, for example, performance cost, safety constraints, given or learned uncertain model, robust performance, robust safety constraints. From L-RCMDP-based RL properties 1010, it can be noted that, in comparison to the properties of CMDP-based RL 1000 method and RMDP-based RL 1002 method, L-RCMDP-based RL properties 1010 exhibit advantageous properties 1012, namely, robust performance and robust safety constraints. Due to such advantageous properties 1012, L-RCMDP-based RL 1004 can seek and guarantee robustness of both performance and safety constraints.
[0153] The properties of each type of RL method define which type of application is suitable for each type of RL method. For example, CMDP-based RL method 1000 can be applied to a constrained system 1014 without uncertainty, such as a robot with an ideal model and an ideal environment with known obstacles. RMDP-based RL method 1002 can be applied to an unconstrained system 1016 with uncertainty, such as a robot with an imperfect model and an imperfect environment without obstacles. L-RCMDP-based RL 1004 can be applied to a constrained system 1018 with uncertainty, such as a robot with an imperfect model and an imperfect environment with obstacles.
[0154] Figure 11 A schematic diagram showing an overview of the RCMDP construct according to some embodiments is shown. Performance cost 1104 is combined with uncertain model set 1100 to form robust performance cost 1106. Specifically, uncertain model set 1100 is incorporated in performance cost 1104 to yield robust performance cost 1106. Furthermore, safety cost 1112 is combined with the same uncertain model set 1100 to form robust safety cost 1110. In particular, uncertain model set 1100 is incorporated in safety cost 1112 to yield robust safety cost 1110. Robust performance cost 1106 together with robust safety cost 1110 constitute RCMDP 1108. The RCMDP construct is explained in detail above with reference to Figure 1A 、 Figure 1B 、 Figure 1C and Figure 3 The RCMDP construct can be solved. Solving RCMDP 1108 can refer to optimizing performance cost 1104 over uncertain model set 1100 subject to optimizing safety cost 1112 over uncertain model set 1100.
[0155] The above description is provided only to illustrate examples of the present disclosure and does not intend to limit the scope, applicability or configuration of the present disclosure. Rather, the above description of the example embodiments will provide enabling descriptions for those skilled in the art to implement one or more example embodiments. It is contemplated that various changes can be made to the function and arrangement of elements without departing from the spirit and scope of the disclosed subject matter as set forth in the appended claims.
[0156] In the above description, specific details are given to provide a thorough understanding of the embodiments. However, those skilled in the art will understand that the embodiments can be practiced without these specific details. For example, the systems, processes, and other elements in the disclosed subject matter can be shown as components in block diagram form in order not to obscure the embodiments in unnecessary detail. In other instances, well-known processes, structures, and techniques can be shown without detailed description in order to avoid obscuring the embodiments. Furthermore, like reference numerals and symbols can be used to designate like elements among the various figures.
[0157] Also, various embodiments can be described as a process, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flow diagram can describe operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations can be re-arranged. A process can be terminated when its operations are completed, but could have additional steps not discussed or included in a figure. Further, not all operations in any particularly described process can occur in every embodiment. A process can correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination can correspond to a return of the function to the calling function or the main function.
[0158] Furthermore, embodiments of the disclosed subject matter can be implemented, at least in part, manually or automatically. Manual or automatic implementation can occur across any combination of devices, hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, the program code or code segments to perform the necessary tasks can be stored in a machine readable medium. A processor(s) can perform the necessary tasks.
[0159] The various methods or processes outlined herein can be coded as software that is executable on one or more processors that employ any one of a variety of operating systems or platforms. Additionally, such software can be written using any of a number of suitable programming languages and / or programming or scripting tools, and also can be compiled as executable machine language code or intermediate code that is executed on a framework or virtual machine. Typically, the functionality of the program modules can be combined or distributed as desired in various embodiments.
[0160] Embodiments of the present disclosure can be implemented as a method, examples of which have been provided. The acts performed as part of the method can be ordered in any suitable way. Accordingly, embodiments in which acts are performed in an order different than illustrated can be constructed, including acts performed simultaneously even though shown as performed sequentially in illustrative embodiments. While the present disclosure has been described with reference to certain preferred embodiments thereof, a worker of ordinary skill in the art will readily appreciate that changes can be made in the spirit and scope of the present disclosure while still falling within the true spirit and scope thereof. Accordingly, reference should be made to the appended claims as indicating the scope of the present disclosure.
Claims
1. A controller for controlling a system subject to constraints on the operation of the system, the system having uncertainty in its dynamics, the controller comprising: at least one processor; and a memory having instructions stored thereon that, when executed by the at least one processor, cause the controller to perform operations of: obtaining historical data of the system's operation, the historical data comprising pairs of control actions and state transitions of the system controlled according to the corresponding control actions; for the system in a current state, determining a current control action that transitions the state of the system from the current state to a next state, wherein the current control action is determined according to a robust and constrained Markov decision process RCMDP that optimizes a performance cost of the system's operation using the historical data as a constraint to impose an optimization of a safety cost of the constraint of the system's operation, wherein the state transitions of each of the state-action pairs of the performance cost and the safety cost are represented by a plurality of state transitions that capture the uncertainty of the dynamics of the system; and controlling the system's operation according to the current control action to change the state of the system from the current state to the next state, the optimization of the performance cost is a minimax optimization that optimizes the performance cost for a worst-case scenario of values of uncertain parameters that cause the uncertainty of the dynamics of the system, the optimization of the safety cost optimizes optimization variables subject to hard constraints.
2. The controller of claim 1, wherein, the performance cost and the safety cost are optimized using Lyapunov descent.
3. The controller of claim 1, wherein, the safety cost is interpreted with an auxiliary cost function configured to enforce satisfaction of the constraints at the current state while making a Lyapunov function decrease along the dynamics of the system at a subsequent evolution of state transitions evolution.
4. The controller of claim 3, wherein, the auxiliary cost function is a solution of a robust linear programming optimization problem that maximizes the value of the auxiliary cost function that keeps the satisfaction of the safety constraints for all possible states of the system with uncertainty of the dynamics.
5. The controller of claim 4, wherein, the auxiliary cost function is a weighted combination of basis functions with weights determined by the solution of the robust linear programming optimization problem.
6. The controller of claim 5, wherein, the auxiliary cost function is a weighted combination of basis functions that are implemented by a neural network whose weights are determined by the solution of the robust linear programming optimization problem.
7. A method for controlling a system with uncertainty in its dynamics subject to constraints on the system's operation, the method comprising the steps of: obtaining historical data of the system's operation, the historical data comprising pairs of control actions and state transitions of the system controlled according to the corresponding control actions; determining, for the system in the current state, a current control action that transitions the state of the system from the current state to a next state, wherein the current control action is determined according to a robust and constrained Markov decision process, RCMDP, that uses the historical data to optimize a performance cost of the operation of the system subject to an optimization of a safety cost of the operation of the system imposing the constraints, wherein a state transition of each of the state and action pairs of the performance cost and the safety cost is represented by a plurality of state transitions that capture the uncertainty of the dynamics of the system; and controlling the operation of the system according to the current control action to change the state of the system from the current state to the next state, the optimization of the performance cost is a minimax optimization of the performance cost for a worst-case scenario of values of uncertain parameters that cause the uncertainty of the dynamics of the system, the optimization of the safety cost optimizes optimization variables subject to hard constraints.
8. The method of claim 7, wherein, the performance cost and the safety cost are optimized using a Lyapunov descent method.
9. The method of claim 7, wherein, the safety cost is interpreted with an auxiliary cost function configured to enforce satisfaction of the constraints at the current state while causing a Lyapunov function to decrease along the dynamics of the system at a subsequent evolution of the state transition evolution.
10. The method of claim 9, wherein, the auxiliary cost function is a solution of a robust linear programming optimization problem that maximizes the value of the auxiliary cost function for satisfaction of the safety constraints for all possible states of the system with uncertainty of the dynamics.
11. The method of claim 10, wherein, the auxiliary cost function is a weighted combination of basis functions with weights determined by the solution of the robust linear programming optimization problem.
12. The method of claim 11, wherein, the auxiliary cost function is a weighted combination of basis functions that are implemented by a neural network, weights of the neural network being determined by the solution of the robust linear programming optimization problem.