Method and system for modeling and controlling a partially measurable system
MC-PILCO addresses the inefficiencies of model-free reinforcement learning and limitations of existing model-based methods by using Gaussian processes and particle-based methods for system dynamics modeling and long-term prediction, achieving improved data efficiency and asymptotic performance.
Patent Information
- Application Number
- JP2023550751
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-12-04
- Filing Date
- 2021-07-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-07-21
AI Technical Summary
Model-free reinforcement learning algorithms require a large number of interactions with the environment to solve tasks, which is inefficient and can cause wear and tear on mechanical systems. Additionally, existing model-based reinforcement learning methods face limitations due to inaccurate models, unimodal distribution assumptions, and restrictive kernel selections.
The Monte Carlo Probabilistic Inference for Learning Control (MC-PILCO) algorithm uses Gaussian processes to model one-step-ahead system dynamics and employs a particle-based method to approximate long-term state distributions, allowing for more flexible kernel choices and improved data efficiency. This approach also incorporates dropout and reparameterization tricks to enhance policy optimization.
MC-PILCO achieves better data efficiency and asymptotic performance compared to state-of-the-art GP-based MBRL algorithms, effectively handling multimodal distributions and improving the ability to escape local minima, thus learning tasks more efficiently and accurately.
Smart Images

Figure 0007699660000042 
Figure 0007699660000043 
Figure 0007699660000044
Abstract
Description
Technical Field
[0001] The present invention generally relates to methods and systems for modeling and controlling partially and fully measurable systems, including mechanical systems.
Background Art
[0002] In recent years, reinforcement learning (RL) has achieved remarkable results in many different environments and shown the potential to provide an automated framework for learning different control applications from scratch. However, model-free RL (MFRL) algorithms may require a huge amount of interaction with the environment to solve the assigned tasks. Data inefficiency poses a limitation on the potential of RL in real-world applications due to the time and cost of interacting with real-world applications. In particular, when dealing with mechanical systems, it is important to learn the task after a minimal amount of possible trials in order to reduce wear and tear and avoid any damage to the system.
[0003] There is a need to develop ways to take into account the modeling and filtering of different components before model learning and before policy optimization.
Summary of the Invention
Problems to be Solved by the Invention
[0004] An object of some embodiments of the present invention is to provide a promising method for overcoming the above limitations, which is model-based reinforcement learning (MBRL), which is based on constructing a predictive model of the environment using data from interactions and using it to plan control actions. MBRL increases data efficiency by using its model to extract more valuable information from the available data.
[0005] Some embodiments of the present invention are based on the recognition that the MBRL method is effective only as long as its models closely resemble the real system. Thus, deterministic models may suffer dramatically from model inaccuracies, and the use of probabilistic models becomes necessary to capture the uncertainty. The Gaussian process (GP) is a class of Bayesian models commonly used in RL methods precisely because of its inherent ability to handle uncertainty and provide principle-based probabilistic predictions. Furthermore, PILCO (Probabilistic Inference for Learning Control) can be a successful MBRL algorithm that uses GP models and gradient-based policy search to achieve substantial data efficiency in solving different control problems in both simulation and real systems. In PILCO, long-term predictions are computed analytically, and the distribution of the next state at each time point is approximated by a Gaussian distribution through moment matching. In this way, the policy gradient is computed in closed form. However, the use of moment matching can bring about two related issues. (i) Moment matching enables the modeling of only unimodal distributions. This fact introduces related limitations associated with the initial conditions, in addition to being a potentially inaccurate assumption in system dynamics. In particular, the limitation on the use of unimodal distributions not only complicates dealing with multimodal initial conditions but is also a potential limitation even when the system initial state is unimodal. For example, when the initial variance is high, the optimal solution may be multimodal due to the dependence on the initial conditions. (ii) The computation of moments has been shown to be tractable only when considering the squared exponential (SE) kernel and a differentiable cost function. In particular, the limitations on kernel selection can be very strict because GPs with SE kernels impose smooth characteristics on the posterior estimator and may exhibit poor generalization characteristics not seen in the data during training.
[0006] Furthermore, some embodiments of the present invention are based on the recognition that PILCO triggered several other MBRL algorithms that sought to improve PILCO in different ways. The limitations due to the use of the SE kernel have been addressed in Deep-PILCO, system evolution can be modeled using Bayesian neural networks, and long-term predictions are calculated by combining particle-based methods and moment matching. The results show that Deep-PILCO requires more interactions with the system to learn the task compared to PILCO. This fact suggests that using neural networks (NNs) may not be advantageous in terms of data efficiency because the amount of parameters required to characterize the model is quite large. A more explicit approach could be to use a probabilistic ensemble of NNs to model the uncertainty in system dynamics. Despite positive results in simulated high-dimensional systems, numerical results show that GPs are more data-efficient than NNs when considering low-dimensional systems such as the inverted pendulum benchmark. An alternative route could be to use a simulator to learn a prior distribution for the GP model before starting the reinforcement learning procedure on the actual system to be controlled. This simulated prior distribution can improve the performance of PILCO in regions of the state space where no data points are available. However, this method requires an accurate simulator that may not always be available to the user. Some issues may arise due to gradient-based optimization and have been addressed in Black-DROPS, which employs gradient-free policy optimization. Some embodiments are based on the recognition that non-differentiable cost functions can be used and the computation time can be improved by parallelizing black-box optimizers. Using this strategy, Black-DROPS achieves similar data efficiency to PILCO but increases the asymptotic performance significantly.
[0007] Furthermore, some embodiments of the present invention are based on the recognition that there are other approaches that focus on overcoming approximations by moment matching to improve the accuracy of long-term predictions. One attempt could be a method that relies on a particle-based approach to calculate the long-term distribution. Based on the current policy and the 1-step-ahead GP model, the evolution of a batch of particles sampled from the initial state distribution can be simulated. Next, the particle trajectories are used to approximate the expected cumulative cost. The policy gradient can be calculated using a certain strategy, and by fixing the initial random seed, the probabilistic Markov decision process (MDP) is transformed into an equivalent partially observable MDP with deterministic transitions. Compared to PILCO, the results were not satisfactory. The low performance was due to the policy optimization method, especially the inability to escape from the numerous minima generated by multimodal distributions. Another particle-based approach could be PIPPS, and the policy gradient is calculated using the so-called reparameterization trick instead of the PEGASUS strategy.
[0008]
Number
[0009] The reparameterization trick has been introduced with successful results in stochastic variational inference (SVI). In contrast to the results obtained in SVI, where only a few samples are required to estimate the gradient, there can be some problems associated with the gradient calculated using the reparameterization trick due to the magnitude of its explosion and random directions. To overcome these problems, they proposed a total propagation algorithm in which the reparameterization trick is combined with the likelihood ratio gradient. This algorithm performs similarly to PILCO and involves some improvements in gradient calculation and performance in the presence of additional noise.
[0010] Some embodiments disclose a model-based reinforcement learning (MBRL) algorithm named Monte Carlo Probabilistic Inference for Learning Control (MC-PILCO). Similar to PILCO, MC-PILCO is a policy gradient algorithm that uses Gaussian processes (GPs) to describe one-step-ahead system dynamics and relies on a particle-based method to approximate the long-term state distribution instead of using moment matching. The gradient of the expected cumulative cost with respect to the policy parameters is obtained by backpropagation on the associated probabilistic computational graph, leveraging the reparameterization trick. Different from PIPPS, which focuses on obtaining accurate estimates of the gradient, the optimization problem can be interpreted as a stochastic gradient descent (SGD) problem. This problem has been deeply studied in the context of neural networks, where over-parameterized models are optimized using noisy estimates of the gradient. Analytical and experimental considerations show that the shape of the adopted cost function and non-linear activation function can dramatically affect the performance of the SGD algorithm. Motivated by the results obtained in this field for previous particle-based approaches, the inventors considered the use of more complex policies and less peaked cost functions, i.e., costs that impose fewer penalties. During policy optimization, the inventors also considered applying dropout to the policy parameters to improve the ability to escape local minima and obtain more implementable policies. The effectiveness of the proposed choices is evaluated and analyzed in simulations. First, to compare MC-PILCO with PILCO and Black-DROPS, a simulated inverted pendulum, a common benchmark system, was considered. The results show that MC-PILCO outperforms both PILCO and Black-DROPS, which can be regarded as state-of-the-art GP-based MBRL algorithms. Second, to evaluate the behavior of MC-PILCO in a higher-dimensional system, it was applied to a simulated UR5 robotic arm. The task considered consists of learning a joint-space controller that can follow a desired trajectory, which was successfully achieved.These results confirm that the reparameterization trick can be effectively used in MBRL, and that the Monte Carlo method does not suffer from the gradient estimation problem as generally claimed in the literature when properly accounting for the cost function, the use of dropout, and complex / rich policies.
[0011] Furthermore, unlike previous studies that combined GP with particle-based methods, the inventors demonstrate the relevant advantages of this strategy, namely the possibility of adopting different kernel functions. Consider the kernel function given by the combination of an SE kernel and a polynomial kernel, as well as the use of semi-parametric models. Results obtained in both simulations and the actual Furuta pendulum show that the use of such kernels significantly increases data efficiency and limits the interaction time required to learn the task.
[0012] Finally, MC-PILCO is applied and analyzed in a partially measurable system and takes the name MC-PILCO4PMS. Different from the simulated environment where the state is usually assumed to be fully measurable, the state of the real system may be partially measurable. For example, in most cases, only the position is directly measured in an actual robotic system, and the velocity is typically calculated by estimators such as numerical differentiation using a state observer, a Kalman filter, and a low-pass filter. In particular, the controller, i.e., the policy, operates on the output of an online state estimator that may introduce significant delays and inconsistencies with respect to the filtered data used during policy training due to noise and real-time calculation constraints. In this regard, the inventors have verified that it is important to distinguish between the state generated by the model, which aims to describe the evolution of the real system state during policy optimization, and the state provided to the policy. In fact, providing model predictions to the control policy corresponds to assuming that the system state is directly measured, which, as mentioned above, is not possible in the real system. This incorrect assumption may undermine the effectiveness of the trained policy for the real system due to the presence of distortions caused by the online state estimator. Therefore, during policy optimization, from the evolution of the system state predicted by the GP model, the inventors calculate the estimated value of the state observed by modeling both the measurement system and the online estimator used in the real system. Next, the estimated value of the observed state is supplied to the policy. In this way, the inventors aim to obtain robustness against the delays and distortions caused by online filtering. The effectiveness of the proposed strategy has been tested in simulations as well as using two real systems, namely, the Furuta pendulum and the ball-and-plate system. The obtained performance confirms the importance of considering the presence of filters in the real system during policy optimization.
Means for Solving the Problems
[0013] Some embodiments of the present invention are based on the recognition that a controller for controlling a system can be provided, including a policy configured to control the system. In this case, the controller may include an interface configured to be connected to the system and obtain an action state and a measurement state via a sensor that measures the system; a memory for storing a computer-executable program module including a model learning module and a policy learning module; and a processor configured to execute steps of the program module. Further, the steps may include an offline modeling step of using a model learning program to generate an offline learning state based on the action state and the measurement state, the model learning module including an offline state estimator and a model learning program, the offline state estimator estimating an offline state estimate value and providing it to the model learning program, the policy learning module including a system model, a sensor model, an online state estimator model, and a policy optimization program, the system model generating a particle state approximating the state of the actual system, the sensor model approximating a particle measurement value approximating a measurement value on the actual system based on the particle state, the online state estimator model being configured to generate a particle online estimate value based on the particle measurement value and optionally a previous particle online estimate value, the policy optimization program generating policy parameters, and the above steps may further include a step of providing the offline state to the policy learning module to generate policy parameters, and a step of updating the policy of the system based on the policy parameters to operate the system.
[0014] According to another embodiment of the present invention, a vehicle control system for controlling the movement of a vehicle is provided. The vehicle control system may include a controller that may include an interface configured to obtain an action state and a measurement state via a sensor connected to the system and measuring the system; a memory for storing a computer-executable program module including a model learning module and a policy learning module; and a processor configured to execute steps of the program module. Further, the steps include an offline modeling step of using a model learning program to generate an offline learning state based on the action state and the measurement state, the model learning module includes an offline state estimator and a model learning module, the offline state estimator estimates an offline state estimate value and provides it to the model learning program, the policy learning module includes a system model, a sensor model, an online state estimator model, and a policy optimization program, the system model generates a particle state approximating the state of the actual system, the sensor model approximates a particle measurement value approximating the measurement value on the actual system based on the particle state, the online state estimator model is configured to generate a particle online estimate value based on the particle measurement value and optionally a previous particle online estimate value, the policy optimization program generates policy parameters, the above steps further include providing the offline state to the policy learning program to generate policy parameters, and updating the policy of the system based on the policy parameters to operate the system, the controller is connected to a motion controller of the vehicle and a vehicle motion sensor for measuring the movement of the vehicle, the control system generates policy parameters based on the measurement data of the movement, and the control system provides the policy parameters to the motion controller of the vehicle to update the policy unit of the motion controller.
[0015] Furthermore, some embodiments of the present invention provide a robot control system for controlling the movement of a robot. The robot control system may include an interface configured to obtain an action state and a measurement state via a sensor connected to the system and measuring the system; a memory for storing computer-executable program modules including a model learning module and a policy learning module; and a processor configured to execute steps of the program modules. Further, the steps may include an offline modeling step of using a model learning program to generate an offline learning state based on the action state and the measurement state, the model learning module including an offline state estimator and a model learning module, the offline state estimator estimating an offline state estimate value and providing it to the model learning program, the policy learning module including a system model, a sensor model, an online state estimator model, and a policy optimization program, the system model generating a particle state approximating the state of the actual system, the sensor model approximating a particle measurement value approximating a measurement value on the actual system based on the particle state, the online state estimator model being configured to generate a particle online estimate value based on the particle measurement value and optionally a previous particle online estimate value, the policy optimization program generating policy parameters, and the above steps further including providing the offline state to a policy learning program to generate policy parameters, and updating the policy of the system based on the policy parameters to operate the system. The controller is connected to an actuator controller of the robot and a sensor configured to measure the state of the robot, the control system generating policy parameters based on the measurement data of the sensor, and the control system providing the policy parameters to the actuator controller of the robot to update the policy unit of the actuator controller.
[0016] The accompanying drawings, which are included to provide a further understanding of the invention, illustrate embodiments of the invention and together with the description serve to explain the principles of the invention.
Brief Description of the Drawings
[0017]
Figure 1A
Figure 1B
Figure 1C
Figure 1D
Figure 1E
Figure 1F
Figure 2A
Figure 2B
Figure 3A
Figure 3B
Figure 4A
Figure 4B
Figure 5A
Figure 5B
Figure 6A
Figure 6B
Figure 6C
Figure 7A
Figure 7B
Figure 8
Figure 9A
Figure 9B
Figure 10A
Figure 10B
Figure 11A
Figure 11B
Figure 12A
Figure 12B
Figure 13A
Figure 13B
Figure 14
Best Mode for Carrying Out the Invention
[0018] Hereinafter, various embodiments of the present invention will be described with reference to the drawings. Note that the drawings are not drawn to scale, and elements having the same structure or function are denoted by the same reference numerals throughout the drawings. Also note that the drawings are only intended to facilitate the description of specific embodiments of the present invention. These are not intended as an exhaustive description of the present invention or as a limitation on the scope of the present invention. Furthermore, aspects described in connection with specific embodiments of the present invention are not necessarily limited to those embodiments and can be implemented in any other embodiment of the present invention.
[0019] According to some embodiments of the present invention, considering two different state estimators has the advantage of providing higher performance when controlling a partially measurable system. A partially measurable system is a system that can directly measure only a subset of the state components and estimate the remaining components by an appropriate state estimator. Partially measurable systems are particularly relevant in real-world applications because they include, for example, mechanical systems (such as vehicle and robot systems) where typically only the position is measured and the velocity is estimated through numerical differentiation or more complex filters. Some embodiments of the present invention model the presence of two different state estimators and include a model learning module, an offline state estimator, a policy learning module, a sensor model, and an online state estimator model. According to some embodiments of the present invention, the presence of the offline state estimator improves the accuracy of the model learning module, and the presence of the sensor model and the online state estimator model improves the performance of the policy learning module.
[0020] FIG. 1A is a block diagram showing a system module 10 controlled via a controller 100 according to an embodiment of the present invention. The components of the system module 10 are shown by a block diagram of a general application example to which the controller 100 according to some embodiments of the present invention can be applied. The controller 100 represents a block diagram of some embodiments of the present invention.
[0021] In the system module 10, components 11, 12, 13, and 14 represent an overview of policy execution and data collection of the actual system 11. Component 11 represents the actual physical system 11 that can be controlled by a controller 100 according to some embodiments of the present invention. Examples of component 11 may be a vehicle system, a robot system, and a suspension system. The actual system 11 may receive a control signal u (action state u ) to move the actual system 11 to a certain state. Then, the state is measured by the sensor 12. The state definition varies according to the actual system 11.
[0022] When the actual system 11 is a vehicle, the state can be the orientation of the vehicle, the steering angle, and the speeds along two axes of the vehicle. When the actual system 11 is a robot system, the state can be the joint positions and joint speeds. When the actual system 11 is a vehicle suspension system, the state can be the displacement of the suspension system from the rest position and the speed of the displacement. The sensor 12 measures the state and outputs a measured value of the state. Most common sensors can measure only some parts of the state and cannot measure all components of the state. For example, positioning sensors such as encoders, potentiometers, proximity sensors, and cameras can measure only the position components of the state; other sensors such as tachometers, laser surface velocimeters, and piezoelectric sensors can measure only the speed components of the state; and other sensors such as accelerometers can measure only the acceleration of the system. None of these sensors outputs the full state, and for this reason, the measured value, which is the output of the sensor 12, is only a part of the state. Therefore, the controller 100 according to some embodiments of the present invention can control the actual system 11 based on dealing with the partially measurable state of the system. The online state estimator 13 takes the measured value as an input, estimates a state called the state estimate, and tries to approximate the parts of the state that do not exist in the measured value. The policy 14 is a controller parameterized by some policy parameters. The policy 14 takes the state estimate as an input and outputs a control signal to control the actual system 11. Examples of policies can be Gaussian processes, neural networks, deep neural networks, proportional-integral-derivative (PID) controllers, and the like.
[0023] The controller 100 includes an interface controller (hardware circuit), a processor, and a memory unit. The processor may be one or more processor units, and the memory unit may be a memory device, a data storage device, etc. The interface controller may be an interface circuit that includes analog / digital (A / D) and digital / analog (D / A) converters for performing signal / data communication with the sensor 12 disposed within the system module 10. Further, the interface controller may include a memory for storing data to be used by the A / D converter or the D / A converter. The sensor 12 is disposed within the system module 10 and measures the state of the actual system 11.
[0024] The controller 100 composed of the model learning module 1300 and the policy learning module 1400 represents some embodiments of the present invention. During policy execution in the system module 10, the measured values and control signals are collected as data. These data are processed by the offline state estimator 131, which filters the data and outputs an offline state estimate. The offline state estimate approximates the state of the actual system 11 that cannot be directly accessed because the sensor 12 outputs only measured values. Examples of the offline state estimator 131 can be a non-causal filter, a Kalman smoother, a central difference velocity approximator, etc. The model learning module (program module) 132 takes the offline state estimate as an input and learns a system model that emulates the actual system. The model learning module 132 can be a Gaussian process, a neural network, a physical model, or any machine learning model. The output of the model learning module 132 is a system model 141 that approximates the actual system 11.
[0025] In the policy learning module 1400, a policy for controlling the system is learned. The components 141, 142, 143 of the policy learning module 1400 approximate the components of the actual system 11, the sensor 12, and the online state estimator 13 of the system module 10, respectively. The component 141 may be a system model 141 configured to approximate the actual system 11, the sensor model 142 approximates the sensor 12, the component 143 approximates the online state estimator 13, and the policy optimization 144 is configured to optimize the policy parameters that define the policy block 14 within the system module 10. The system model 141 is configured to approximate the actual system 11. When a control signal is applied to the system model 141, a particle state is generated by the policy optimization module 144. The particle state generated by the system model 141 is an approximation of the state of the actual system 11. The model of the sensor 142 calculates a particle measurement value, which is an approximation of the measurement value in the system module 10, from the particle state. The model of the online state estimator 143 calculates an online estimate from the particle measurement value and the previous particle online estimate. The online estimate is an approximation of the state estimate in the system module 10. The policy optimization block 144 takes the particle online estimate and the particle measurement value as inputs and learns the optimal policy parameters for the policy 14. During learning (training), the policy optimization 144 generates a control signal to be transmitted to the system model 141. When the learning is completed, the policy parameters can be used to define the policy 14 for controlling the actual system.
[0026] FIG. 1B is a block diagram showing a vehicle motion controller 100B configured to control a vehicle motion system according to an embodiment of the present invention. In this case, the controller 100 can be applied to a vehicle control system 100B for controlling the motion of a vehicle. The vehicle control system 100B including a model learning module and a policy learning module is connected to a vehicle motion sensor that measures the motion of the vehicle and the motion controller of the vehicle, and the control system generates policy parameters based on the measured data of the motion, and the control system provides the policy parameters to the motion controller of the vehicle to update the policy unit of the motion controller.
[0027] The vehicle motion controller 100B may include an interface controller 110B, a processor 120, and a memory unit 130B. The processor 120 may be one or more processor units, and the memory unit 130B may be a memory device, a data storage device, etc. The interface controller 110B may be an interface circuit that may include an analog / digital (A / D) and digital / analog (D / A) converter for signal / data communication with the vehicle motion sensor 1101, the road surface roughness sensor 1102, and the vehicle motion controller 150B of the vehicle. Further, the interface controller may include a memory for storing data to be used by the A / D converter or the D / A converter. The vehicle motion sensor 1101 and the road surface roughness sensor 1102 are arranged on the vehicle to measure the motion state of the vehicle. The vehicle includes a motion controller device / circuit including a policy unit 151B that generates action parameters to control a suspension system 1103 that controls suspension devices 1103-1, 1103-2, 1103-3, and 1103-4. The suspension devices may be 1103-1, 1103-2, 1103-3, … 1103-#N according to the number of wheels. For example, the vehicle motion sensor 1101 may include an acceleration sensor, a positioning sensor, or a global positioning system (GPS) device to measure the motion state of the vehicle. The road surface roughness sensor 1102 may include an acceleration sensor, a positioning sensor, etc.
[0028] The interface controller 110B is also connected to a vehicle sensor 1101 that measures the state of motion of the vehicle. Further, the interface controller 110B may be connected to a road surface roughness sensor 1102 mounted on the vehicle to obtain information on the roughness of the road on which the vehicle is traveling. In some cases, when the vehicle is an electric vehicle, the motion controller 150B may control the individual electric motors that drive the wheels of the vehicle. In some cases, the motion controller 150B may control the rotation of the individual wheels so as to smoothly accelerate or safely decelerate the vehicle in response to the policy parameters generated from the policy learning module 1400B. Further, depending on the design of the vehicle driving operation, the motion controller 150B may control the angle of the wheels in response to the policy parameters generated from the policy learning module 1400B.
[0029] The memory unit 130B may store computer-executable program modules including a model learning module 1300B and a policy learning module 1400B. The processor 120 is configured to execute the steps of the program modules 1300B and 1400B. In this case, the steps may include offline modeling that generates an offline learning state based on the action state (motion state) and measurement state of the vehicle from the vehicle motion sensor 1101, the road surface roughness sensor 1102, or a combination of the vehicle motion sensor 1101 and the road surface roughness sensor 1102 using the model learning module 1300B. The steps may further include providing the offline state to the policy learning module 1400B and updating the policy 151B of the vehicle's motion controller 150B to operate the actuator or the suspension system 1103 based on the policy parameters.
[0030] FIG. 1C is a block diagram showing a robot control system 100C for controlling the motion of a robot according to an embodiment of the present invention. The robot control system 100C is configured to control the actuator system 1203 of the robot.
[0031] In this case, the controller 100 can be applied to a robot control system 100C for controlling the movement of the robot. The robot control system 100C includes a model learning module 1300C and a policy learning module 1400C, is connected to the robot's actuator controller 150C and a sensor 1201 that measures the movement of the robot. The robot control system 100C generates policy parameters based on the measured movement data, and the control system 100C provides the policy parameters to the robot's actuator controller 150C to update the policy unit 151C of the actuator controller.
[0032] The robot control system 100C may include an interface controller 110C, a processor 120, and a memory unit 130C. The processor 120 may be one or more processor units, and the memory unit 130C may be a memory device, a data storage device, etc. The interface controller 110C can be an interface circuit that may include analog / digital (A / D) and digital / analog (D / A) converters to perform signal / data communication with the sensor 1201 and the motion controller 150C of the robot. Further, the interface controller 110C may include a memory for storing data to be used by the A / D converter or the D / A converter. The sensor 1201 is disposed at the joints of the robot (robot arm) or the object collection mechanism (e.g., finger part) to measure the state of the robot. The robot includes an actuator controller (device / circuit) 150C including a policy unit 151C to generate action parameters to control a robot system 1203 that controls combinations of robot arms, handling mechanisms, or arms and handling mechanisms 1203-1, 1203-2, 1203-3, and 1203-#N according to the number of joints or handling fingers. For example, the sensor 1201 may include an acceleration sensor, a positioning sensor, or a global positioning system (GPS) device to measure the motion state of the vehicle. The sensor 1201 may include an acceleration sensor, a positioning sensor, etc.
[0033] Also, the interface controller 110C is connected to a sensor 1201 mounted on the robot for measuring / acquiring the motion state of the robot. In some cases, when the actuator is an electric motor, the actuator controller 150C may control individual electric motors that drive the angle of the robot arm or the handling of an object by the handling mechanism. In some cases, the actuator controller 150C may control the rotation of individual motors arranged on the arm so as to smoothly accelerate or safely decelerate the motion of the robot in response to the policy parameters generated from the policy learning module 1400C. Further, depending on the design of the object handling mechanism, the actuator controller 150C may control the length of the actuator in response to the policy parameters generated from the policy learning module 1400C.
[0034] The memory unit 130C may store computer-executable program modules including a model learning module 1300C and a policy learning module 1400C. The processor 120 is configured to execute the steps of the program modules 1300C and 1400C. In this case, the steps may include offline modeling to generate an offline learning state based on the action state (motion state) of the robot and the measurement state from the sensor 1201 using the model learning module 1300C. These steps further include providing the offline state to the policy learning module 1400C to generate policy parameters, and updating the policy 151C of the motion controller 150C of the robot based on the policy parameters to operate the actuator system 1203.
[0035] FIG. 1D is a schematic diagram showing a particle-based method for policy optimization for controlling a system according to an embodiment of the present invention.
[0036] Component 1301 represents a schematic diagram of the initial state distribution of the system. Components 1302 and 1312 are named particles and are examples of initial conditions sampled according to the initial state distribution. The particle state evolution of particle 1302 is represented by 1303, 1304, and 1305. Starting from 1302, the system model estimates the distribution of particle states in the following steps represented by 1303 in the first step, 1304 in the second step, and 1305 in the third step. The state evolution continues for the number of steps determined when the simulation continues. Similarly, the particle state evolution of particle 1312 is represented by 1313, 1314, and 1315.
[0037] According to the present invention, several embodiments are described as follows. The inventors describe the general problems of the model-based policy gradient method and present a modeling approach for mechanical systems with GP. Furthermore, the inventors present MC-PILCO, their proposed algorithm for fully measurable systems, which details the policy optimization and model learning techniques employed. The inventors analyze several aspects that affect the performance of MC-PILCO, such as cost shaping, dropout, and kernel selection. In addition, the inventors compare MC-PILCO with PILCO and Black-DROPS using a simulated inverted pendulum benchmark system, test the performance of MC-PILCO using a simulated UR5 robot, and also test the advantages of the particle-based approach when dealing with different distributions of initial conditions. The inventors present an extension of the algorithm MCPILCO to systems with partially measurable states, which is here called MC-PILCO4PMS. Experiments are shown as examples, and the controller according to the present invention is applied to an actual Furuta pendulum and a ball-and-plate system.
[0038] Model-Based Policy Gradient
[0039]
Number
[0040]
Number
[0041]
Number
[0042] The model-based approach for learning a policy generally consists of a sequence of several trials; that is, it attempts to solve the desired task. Each trial consists of three main stages: · Model learning: Using the data collected from all previous dialogues, construct a model of the system dynamics (in the first iteration, in some cases apply random search control to collect data); · Policy update: Optimize the policy to minimize the cost J(θ) according to the current model; · Policy execution: Apply the current optimized policy to the system and save data for model improvement.
[0043] The model-based policy gradient method uses the learned model to predict the state evolution when the current policy is applied. These predictions are used to estimate J(θ) and its gradient in order to update the policy parameter θ according to the gradient descent approach
[0044]
Number
[0045] and are used for this purpose. GPR and One-Step Ahead Prediction In this section, we discuss how the inventors use Gaussian process regression (GPR) for model learning. The inventors focus on three aspects, namely, some background concepts regarding GPR, the description of model prediction for one step ahead, and finally, the inventors discuss long-term prediction focusing on two possible strategies, namely, moment matching and particle-based methods.
[0046]
Number
[0047] Here, the scaling factor λ and the matrix Λ are kernel hyperparameters that can be estimated by maximizing the marginal likelihood. Typically, Λ is assumed to be diagonal, and the diagonal elements are named length scales.
[0048]
Number
[0049]
Number
[0050] Long-term Prediction Using GP Dynamics Model
[0051]
Number
[0052] Unfortunately, it is intractable to compute the exactly predicted distribution in (8). There are various ways to approximately solve this, and here we discuss two main approaches, namely, moment matching adopted by PILCO, and the particle-based method which is the strategy followed in this discussion.
[0053] Moment Matching
[0054]
Mathematics
[0055] Finally, the procedure is repeated for each time step in the prediction horizon to compute the subsequent probability distribution. For details regarding the calculation of the first and second moments, moment matching offers the advantage of providing a closed-form solution for handling uncertainty propagation through the GP dynamics model. Thus, in this setting, it is possible to analytically compute the policy gradient from long-term predictions. However, as already mentioned above, the Gaussian approximation performed in moment matching is also the cause of two main weaknesses: (i) The calculation of the two moments is performed assuming the use of an SE kernel, which can lead to poor generalization properties in data not seen during training. (ii) Moment matching only allows modeling of unimodal distributions, which may be too limiting an approximation of real system behavior.
[0056] Particle-based method
[0057]
Mathematics
[0058] MC-PILCO The following presents the algorithm proposed for a fully measurable system. MC-PILCO relies on GPR for model learning and estimates the cumulative cost predicted from particle trajectories propagated through the learned model according to the Monte Carlo sampling method. The policy gradient is obtained from the sampled particles, and the reparameterization trick is utilized to optimize the policy. This approach is very flexible, enabling the use of any type of kernel for the GP and providing a more reliable approximation of the system's behavior. MC-PILCO consists, broadly speaking, of the iteration of three main steps, namely, updating the GP model, updating the policy parameters, and executing the policy on the system. Then, the policy update consists of three steps and is repeated up to N opt times:
[0059]
Number
[0060] ●The following discusses the model learning step and the policy optimization step in more depth.
[0061] Model learning program Here, the model learning framework considered in MC-PILCO is described. The inventors start by presenting the proposed one-step-ahead prediction model. Next, the selection of the kernel function is explained. Finally, the inventors briefly discuss the strategies adopted for model hyperparameter optimization and reducing the computational cost.
[0062] One-step-ahead model
[0063]
Number
[0064]
Number
[0065] Kernel function Regardless of the GP mechanical model structure adopted, one of the advantages of the particle-based policy optimization method is the possibility of selecting any kernel function without constraints. Therefore, the inventors considered different kernel functions, for example, to model the evolution of physical systems. However, the reader may consider a custom kernel function suitable for their own application example. ● Squared Exponential (SE). The SE kernel described in (2) represents a standard choice adopted in many different considerations. ● SE + Polynomial (SE+P (d) ). Recall that the sum of kernels is still a kernel, and the kernel given by the sum of the SE and polynomial kernels was also considered. In particular, the multiplicative polynomial (MP) kernel, which is an improvement of the standard polynomial kernel, was used. The MP kernel of degree d is defined as the product of d linear kernels, i.e.,
[0066]
Equation
[0067]
Equation
[0068] The basic principle behind this kernel is as follows: k PI encodes the prior information given by physics, and k SE compensates for the mechanical components not modeled in k PI .
[0069] Model Optimization and Reduction Techniques In MC-PILCO, the GP hyperparameters are optimized by maximizing the marginal likelihood (ML) of the training samples. Previously, the inventors found that the computational cost of particle prediction scales with the square of the number of samples n, imposing a significant computational burden when n is high. In this context, it is essential to implement strategies to limit the computational load of prediction. Several solutions have been proposed in the literature. The inventors implemented the procedure proposed by the authors for an online importance sampling strategy. After optimizing the GP hyperparameters by ML maximization, the samples in D are downsampled to a subset
[0070]
Number
[0071] and then used to compute the predictions. This procedure first initializes D with the first sample in D and then iteratively computes the GP estimates for all the remaining samples in D as training samples using D r and then uses D r . If the uncertainty of the estimate is higher than a threshold β (i) , each sample in D is either added to D r or discarded. The GP estimator is updated each time a sample is added to D r . The trade-off between the reduction in computational complexity and the severity of the introduced approximation is adjusted by tuning β (i) . The higher β (i) , the fewer samples in D r . On the other hand, using too high a value of β (i) may compromise the accuracy of the GP prediction.
[0072] Policy optimization program Here, we present the policy optimization strategy adopted in MC-PILCO. We begin by describing the general policy structure under consideration. Later, we show how to utilize backpropagation and the reparameterization trick to estimate the policy gradient from particle-based long-term predictions. Finally, we explain how to implement dropout in this framework.
[0073] Policy structure In all the experiments presented in this discussion, we considered an RBF network policy with outputs bounded by a properly scaled hyperbolic tangent function. We refer to this function as the squashed-RBF-network, which
[0074]
Number
[0075] Gradient calculation
[0076]
Number
[0077]
Number
[0078] Dropout
[0079]
Number
[0080] The use of a probabilistic policy during policy optimization makes it possible to increase the entropy of the distribution of particles. This property increases the probability of visiting low-cost regions and escaping from minima. Furthermore, the inventors have also verified that dropout can mitigate problems related to exploding gradients. This is probably due to the fact that, in order to compute the gradient, the average of several different values of w is used, rather than a single value of w, i.e., different policy functions are used to obtain regularization of the gradient estimate.
[0081]
Number
[0082]
Number
[0083] The first case occurs when the optimization reaches a minimum, but the high variance means that the particle trajectories cross regions of the work space where the uncertainty of the GP prediction is high. In both cases, the inventors are interested in testing the policy on the real system in order to verify, in the first case, whether the reached configuration solves the task and, in the second case, to collect data for which the prediction is uncertain and thus improve the model accuracy. The algorithm MC-PILCO with dropout is summarized in pseudo-code in Figure 1F.
[0084] The inventors conclude the discussion on policy optimization by reporting the optimization parameters used in all the proposed experiments in Figure 1E, unless stated otherwise explicitly. However, it is worth mentioning that some adaptation may be required in other settings depending on the problem considered.
[0085] Ablation experiments In what follows, we analyze several aspects that affect the performance of MC-PILCO, such as the shape of the cost function, the use of dropout, kernel selection, and the probabilistic model employed, namely, the full-state or velocity-integrated dynamics model. The purpose of the analysis is to verify the choices made in the proposed algorithm MC-PILCO and show the impact they have on the control of dynamical systems. MC-PILCO is implemented in Python and utilizes the automatic differentiation feature of the PyTorch library; the code is publicly available. The inventors considered the swing-up of a simulated inverted pendulum, a classical benchmark problem, to conduct ablation experiments. The system and the experiments are described below. The physical characteristics of the system are the same as those of the system used in PILCO: the mass of both the cart and the rod is 0.5 [kg], the length of the rod is L = 0.5 [m], and the friction coefficient between the cart and the ground is 0.1.
[0086]
Number
[0087]
Number
[0088] All comparisons are in a Monte Carlo simulation consisting of 50 experiments. Each experiment consists of 5 trials, each 3 seconds long. The random seed varies in each experiment to account for different explorations and initializations of the policy, as well as different realizations of measurement noise. The performance of the learned policy is evaluated using the following cost
[0089]
Number
[0090] Cost shaping The first test concerns the performance obtained by varying the length scale of the cost function in (19). Reward shaping is an important aspect of RL known, and here we analyze it for MCPILCO. In FIGS. 2A and 2B,
[0091]
Number
[0092] we compare the evolution of the cumulative cost obtained therein and report the observed success rate. The set of the latter length scales defines a more selective cost as the function shape becomes more distorted. In both cases, we adopted the speed integration model together with the SE kernel and dropout was not used during policy optimization.
[0093]
Number
[0094] This fact suggests that the use of a too selective cost function may significantly reduce the probability of converging to the solution. The reason is that for small values of the length scale, when the policy parameters are far from a good configuration, c(x t) becomes very sharp, resulting in an almost zero gradient and potentially increasing the probability of getting stuck at the minimum. Instead, a larger length scale value may still promote the existence of a non-zero gradient far from the goal and facilitate the policy optimization procedure. These observations have already been made in PILCO, but we did not encounter difficulties in using a small length scale such as 0.25 in (20). This may be due to the analytical calculation of the policy gradient made possible by moment matching and the different optimization algorithms used. On the other hand, the value of the length scale does not seem to affect the accuracy of the learned solution. To confirm this, in Figure 6C, in rows 3 - 4, we report the average distance from the target state obtained by the policy that succeeded during the last second of the episode in trial 5. No significant difference can be observed regarding the accuracy in reaching the goal.
[0095] Dropout In this experiment, we compared the results obtained with or without using dropout during policy optimization. In Figures 3A and 3B, we compare the evolution of the cumulative cost obtained in the two cases and show the success rate obtained.
[0096] In both scenarios, we adopted a speed integration model with an SE kernel and a cost function with length scales (l θ = 3, l p = 1). When using dropout, MC-PILCO learned the optimal solution in 94% of the experiments at trial 4 and somehow obtained the optimal solution for all random seeds by trial 5. Instead, without dropout, even in the last trial, the optimal policy is not always found. Note that when not using dropout, the upper bound of the cumulative cost in the last two trials is higher and the task cannot always be solved. Furthermore, rows 2 - 4 in Figure 6C show that the use of dropout also helps to reduce the cart positioning error at the end of the swing (both regarding the mean and the standard deviation).
[0097] Empirically, the inventors have found that dropout not only helps to stabilize the learning process and more consistently find better solutions, but can also improve the accuracy of the learned policy.
[0098] Kernel function In this test, the results obtained using either the SE kernel or the SE+P (2) kernel were compared. In both cases, a speed integration model was adopted, the cost function was defined with length scales (l θ =3, l p =1), and dropout was used. Figures 4A and 4B show that SE+P (2) converges to the optimal solution more quickly than SE. For the SE+P (2) kernel, the algorithm learns in 90% of the cases at trial 3 and obtains a 100% success rate at trial 4. On the other hand, when using the SE kernel, the task is only solved at trial 5 for all random seeds. This can be explained by the ability of the more structured kernel to learn the correct dynamics of the system well even in regions of the state-action space where there are no available data points. In fact, some parts of the dynamics of the inverted pendulum system are polynomial functions of the GP input
[0099]
Number
[0100] and the structure of SE+P (2) improves the data efficiency of model learning. Speed integration model In this test, the inventors compared the performance obtained by the proposed speed integration dynamics model and by a standard full-state model. In both cases, the SE kernel was selected, the cost function was defined with length scales (l θ =3, l pIt was defined by (1), and dropout was used. FIGS. 5A and 5B show that the speed integration model obtains better performance with a narrower confidence interval and a better success rate in Trials 2 and 3. In contrast, during the last two trials, the success rate of the full-state model is slightly better. In the full-state model, the position and velocity are learned independently, but it should be recalled that in the speed integration model, the position is calculated as the integral of the velocity under a constant acceleration assumption. The speed integration model can then reduce the uncertainty in long-term prediction and facilitate learning about the counterparts when a small number of data points are collected. In fact, the full-state model may face some difficulties in learning the relationship between the position and each velocity from a limited amount of data. This reduction in uncertainty may explain the narrower confidence intervals observed during the first trial of the experiment. On the other hand, when sufficient data points have been collected (Trials 4 and 5), the improvement in accuracy obtained by the full-state model is not very significant. Even with comparable performance, the choice of the speed integration model is justified as it halves the number of GPs to be learned, and thus this structure also improves the computation time.
[0101] Experiments in Simulation In the following, two simulated systems are considered. First, MC-PILCO is tested on an inverted pendulum system and compared with other policy gradient algorithms, namely PILCO and Black-DROPS. In the same environment, the inventors tested the ability of MC-PILCO to handle bimodal probability distributions. Second, MC-PILCO learns a controller in the joint space of a UR5 robotic arm, which is considered as an example of a higher DoF system.
[0102] Inverted Pendulum: Comparison with Other Methods The inventors tested PILCO, Black-DROPS, and MC-PILCO on the aforementioned inverted pendulum system. In MC-PILCO, since all three algorithms have the same kernel function, the cost function (19) was set with a length scale (l θ = 3, l p=1) and the SE kernel. The results of the cumulative cost are reported in FIGS. 6A and 6B. MC-PILCO achieved the best performance both in terms of transient and convergence, and by trial 5, learned a way to swing up the pendulum with a 100% success rate. In each and every trial, MC-PILCO obtained a cumulative cost with a lower median and lower variability. On the other hand, the policy in PILCO showed poor convergence characteristics with a success rate of only 42% after all 5 trials. Black-DROPS is better performing than PILCO, but obtained worse results than MC-PILCO in each and every trial, and the success rate at trial 5 is only 86%. MC-PILCO is SE+P (2) Considering the kernel, it is desired to recall that even better performance can be obtained. FIGS. 6C, the results of rows 1-2-6-7 also show that the policy learned by MC-PILCO is more precise in reaching the target.
[0103] Pendulum: Handling of bimodal distributions One of the main advantages of particle-based policy optimization is the ability to handle multimodal state evolution. This is impossible when applying methods based on moment matching such as PILCO. The inventors considered a very high variance σ 2 p =0.5 at the initial cart position to cope with having an unknown initial position of the cart (although restricted within a reasonable range), and verified this advantage by applying both PILCO and MC-PILCO to a simulated pendulum system. The goal is a situation where the policy has to solve the task regardless of the initial conditions and needs to have a bimodal behavior to be optimal. It should be noted that the situation described may be relevant to several practical applications. The inventors maintained the same settings used in previous pendulum experiments and changed the initial state distribution to a zero-mean Gaussian distribution with a covariance matrix diag([0.5,10 -4 ,10 -4 ,10 -4 ). MC-PILCO used a length scale (l θ =3, lp Optimize at \(= 1\). The inventors started from 9 different initial trolley positions \((-2, -1.5, -1, -0.5, 0, 0.5, 1, 1.5, 2 [m])\) and tested the policies learned by two algorithms. Previously, the inventors struggled to get PILCO to consistently converge to a solution and observed that the high variance in the initial conditions highlighted this problem. Nevertheless, to enable comparison, the inventors carefully selected the random number seeds for which PILCO converged to a solution in this particular scenario. Figures 7A and 7B show the results of the experiment. MC-PILCO can handle the initial high variance. It learns a bimodal policy that pushes the trolley in two opposite directions depending on the initial position of the trolley, stabilizing the system in all experiments. In contrast, the policy of PILCO cannot control the inverted pendulum for all the starting conditions tested. Its strategy is to always push the trolley in the same direction, and when the trolley starts far from the zero position, it cannot stabilize the system. The state evolution under the policy of MC-PILCO is bimodal, but PILCO cannot find this type of solution due to the unimodality approximation implemented by moment matching.
[0104] In this example, the inventors found that when starting from a unimodal state distribution with high variance, due to the dependence on the initial conditions, multimodal state evolution can be the optimal solution. In other cases, multimodality can be directly implemented by the existence of multiple possible initial conditions that would be poorly modeled by a single unimodal distribution. MC-PILCO can handle all these situations thanks to its particle-based method for long-term prediction. Similar results were obtained when considering a bimodal initial distribution. Due to space constraints, the inventors do not report the results obtained, but the experiments are available in the code of the supplementary materials.
[0105] UR5 Joint Space Controller: High DoF Application Example
[0106]
Number
[0107] [Number]
[0108] The inventors assumed full state observability with the measured value perturbed by white noise with a standard deviation of 10 -3 . The initial state distribution is a Gaussian distribution with a standard deviation of 10
[0109] [Number]
[0110] centered around -3 . The policy optimization parameters are the same as those reported in Figure 1E, except for n s = 400 and δ s = 0.05, which implement more restrictive end conditions.
[0111] In Figures 9A and 9B, the inventors report the trajectory followed by the end effector in each trial, along with the desired trajectory. MC-PILCO, with a PD controller, significantly improved the high tracking error obtained after only 2 trials (corresponding to 8 seconds of interaction with the system). The learned control policy followed the reference trajectory of the end effector with a mean error of 0.65 ± 0.69 [mm] (confidence interval calculated as 3 × standard deviation) and a maximum error of 1.08 [mm].
[0112] MC-PILCO for Partially Measurable Systems Below, we discuss the application of MC-PILCO to systems where the state is partially measurable, i.e., systems where the state is observable, but only some components of the state can be directly measured and the rest must be inferred from the measurements. For the sake of brevity, we introduce the problem by discussing the case of a mechanical system where only the position can be measured (the velocity cannot), but a similar consideration can be carried out for any partially measurable system with an observable state. Next, we describe MC-PILCO for Partially Measurable Systems (MC-PILCO4PMS), a modified version of MC-PILCO proposed to handle such settings. The algorithm MC-PILCO4PMS is verified in simulation as a proof of concept.
[0113] MC-PILCO4PMS
[0114]
Number
[0115] In particular, it is valuable to distinguish between online-computed and offline-computed estimates. The former is provided to the control policy to determine the system control input and needs to take into account real-time constraints, i.e., the velocity estimate is causal and the calculation must be performed within a given interval. The latter does not need to handle such constraints. As a result, the offline estimate can take into account non-causal information and limit delays and distortions, and thus can be more accurate.
[0116] In this regard, the inventors verified that it is relevant to distinguish between the particle state prediction calculated by the model and the data provided to the policy during policy optimization. In fact, the GP should simulate the actual system dynamics regardless of the additional noise given by the sensing instrumentation, and thus it needs to operate with the most accurate available estimates; delays and distortions may impair the accuracy of long-term predictions. On the other hand, directly providing the state of the particles calculated using the GP to the policy during policy optimization corresponds to training the policy by directly assuming the available access to the system state, which, as mentioned above, is not possible in the considered settings. In fact, a significant difference between the state of the particles and the state estimate calculated online during the application of the policy to the actual system may impair the effectiveness of the policy. This approach is typically distinguished from standard MBRL approaches where the effect of the online state estimator is not considered during training.
[0117] To address the above problems, the inventors introduced MC-PILCO4PMS, a modified version of MC-PILCO. In MC-PILCO4PMS, the inventors propose the following two additional aspects for MC-PILCO.
[0118] Calculation of GP training data using an offline state estimator
[0119]
Number
[0120] The estimation of the state using a Kalman smoother is given by a general equation that relates the state space model to position, velocity, and acceleration. The advantage of this technique is to utilize the correlation between position and velocity and increase regularization.
[0121] Simulation of an online estimator
[0122]
Number
[0123] Simulation results Here, the present inventors test the relevance of modeling the presence of an online estimator using the simulated inverted pendulum system, adding the assumption of emulating real-world experiments. The present inventors considered the same physical parameters and the same initial conditions as described for the inverted pendulum system above, but assumed that only the position of the cart and the angle of the rod are measured. The present inventors assume that in the real world, the standard deviation is 3·10 -3A possible measurement system with additive Gaussian independent and identically distributed noise was modeled. To obtain a reliable estimate of the speed, samples were collected at 30 [Hz]. The online estimate of the speed was calculated by causal numerical differentiation followed by a first-order low-pass filter with a cut-off frequency of 7.5 [Hz]. The speed used to train the GP was derived using a central difference formula. The effectiveness of MC-PILCO4PMS against MC-PILCO was verified in this system. The exploration data was collected using a random exploration policy. To avoid dependence on initial conditions such as policy initialization and exploration data, the same random seed was fixed in both experiments. In FIGS. 11A and 11B, the inventors report the results of Monte Carlo simulations over 400 runs. In FIG. 11A, the final policy is applied to the learned model (ROLLOUT), and in FIG. 11B, it is applied to the inverted pendulum system (TEST). The two policies behave similarly when applied to the model, but can all be tested offline, and the results obtained by testing the policy in the inverted pendulum system are significantly different. MC-PILCO4PMS solves the task in all 400 trials. In contrast, in some trials, MC-PILCO does not solve the task due to delays and inconsistencies introduced by the online filter and not considered during policy optimization. The inventors believe that these considerations regarding how to manipulate the data during model learning and policy optimization may be beneficial for other MBRL algorithms different from MC-PILCO.
[0124] Experiments using the exemplary system In the following, the inventors test MC-PILCO4PMS when applied to a real system. In particular, the inventors experimented with two benchmark systems, namely, Furuta's pendulum (FIG. 12A) and ball-and-plate (FIG. 12B). These are just a few examples of real systems to which some embodiments of the present invention may be applied. Other examples of real systems can be robotic manipulators, vehicles, and suspension systems.
[0125] Furuta's Pendulum The Furuta Pendulum (FP) is a popular benchmark system used in nonlinear control and reinforcement learning. This system consists of two revolute joints and three links. The first link, called the base link, is fixed and perpendicular to the ground. The second link, called the arm, rotates parallel to the ground, and the axis of rotation of the last link, the pendulum, is parallel to the main axis of the second link (see Figure 12A). The FP is an underactuated system, since only the first joint is actuated. In particular, in the considered FP, the horizontal joint is actuated by a DC servo motor and the two angles are measured at 4096 [ppr] by optical encoders.
[0126]
number
[0127]
number
[0128] MC-PILCO4PMS managed to learn how to swing the Furuta pendulum in all cases. It did so in Trial 6 with Kernel SE, Kernel SE+P (2) Trial 4 with,and trial 3 with the SP kernel were successful.,These experimental results confirm the higher data efficiency of,more structured kernels, and the advantage that MC-PILCO4PMS,offers by allowing arbitrary kinds of kernel,functions.
[0129]
number
[0130] Bowl and Plate
[0131]
number
[0132] The trial length is 3 seconds and the sampling frequency is 30 [Hz]. The measurements provided by the camera are very noisy and cannot be used directly to estimate velocity from position. The inventors
[0133]
Number
[0134] used a Kalman smoother for offline filtering of. In the control loop, instead, the inventors used a Kalman filter to estimate the ball state online from the noisy position measurements. When simulating the online estimator during policy optimization, the inventors tried both perturbing the predicted particle positions with some additive noise and not perturbing them. The inventors obtained similar performance in both cases, which may be due to the fact that the Kalman filter can effectively filter out the white noise added to the particles.
[0135]
Number
[0136] According to some embodiments of the present invention, the proposed framework can use Gaussian processes (GPs) to derive a probabilistic model of the system dynamics and update policy parameters through gradient-based optimization; the optimization utilizes the reparameterization trick and relies on Monte Carlo methods to approximate the expected cumulative cost. Compared to similar algorithms proposed in the past, the Monte Carlo method has worked by focusing on two aspects, namely, (i) the appropriate selection of the cost function, and (ii) the introduction of exploration during policy optimization through the use of dropout. The inventors compared MC-PILCO with two state-of-the-art GP-based MBRL algorithms, PILCO and Black-DROPS. MC-PILCO outperforms both algorithms and exhibits better data efficiency and asymptotic performance. The results obtained in simulations confirm the effectiveness of the proposed solution and show the relevance of the two aforementioned aspects when optimizing a policy that combines the reparameterization trick with a particle-based approach. Furthermore, the inventors investigated two advantages of a particle-based approximation for moment matching employed in PILCO, namely, the possibility of using structured kernels such as polynomial and semi-parametric kernels, and the ability to handle multimodal distributions. In particular, the results obtained in simulations and using real systems show that the use of structured kernels can increase data efficiency and reduce the interaction time required to learn the task. Some embodiments show systems with partially measurable states that are particularly relevant in practical applications. Furthermore, some embodiments may provide a modified algorithm called MC-PILCO4PMS, where the inventors verified the importance of considering the state estimator used in real systems during policy optimization. Some results are shown for different simulated scenarios, namely, the inverted pendulum and the robotic manipulator, and also on real systems, such as the Furuta pendulum and the ball-and-plate setup.
[0137] The above-described embodiments of the present invention can be realized in any of a number of ways. For example, the embodiments may be realized using hardware, software, or a combination thereof. When realized in software, the software code can be executed on any suitable processor or collection of processors, regardless of whether it is provided on a single computer or distributed among multiple computers. Such a processor may be realized as an integrated circuit having one or more processors within the integrated circuit component. However, the processor may be realized using any suitable form of circuitry. Also, embodiments of the present invention may be embodied as a method for which examples are provided. The acts performed as part of the method may be ordered in any suitable manner. Thus, although shown as consecutive acts in the exemplary embodiments, embodiments may be constructed in which some acts are performed simultaneously, including in a different order than that of the example. Furthermore, the use of ordinal numbers such as "first," "second," etc. to modify claim elements in the claims does not, in and of itself, imply a priority, precedence, or order of one claim element over another claim element, or an order in time in which the acts of a method are performed, but is used only as a label to distinguish a claim element having a particular name from another element of the same name (in the absence of the use of the ordinal number) to distinguish those claim elements.
[0138] Although the present invention has been described by way of examples of preferred embodiments, it should be understood that various other adaptations and modifications can be made within the spirit and scope of the present invention. Therefore, the purpose of the claims is to encompass all such variations and modifications that are within the true spirit and scope of the present invention.
Claims
1. A controller for controlling the system, including a policy configured to control the system, An interface connected to the system, configured to obtain a measurement state including at least one of the position, velocity, and acceleration of an object included in the system via a control signal for moving the system to a predetermined state and a sensor for measuring the system, A memory for storing a computer-executable program module including a first model learning module and a policy learning module, A processor configured to execute steps of the program module, the steps including: An offline modeling step including generating an offline learning state based on the control signal and the measurement state using a model learning program, The model learning program is configured to consider a squared exponential (SE) kernel, a multiplicative polynomial (MP) kernel, and / or a semi-parametric (SP) kernel for modeling the system. The SE kernel, MP kernel, and SP kernel are configured to use inputs of Gaussian process regression and Gaussian process for model learning. The inputs include the control signal and an input with time, measured by the sensor, The first model learning module includes an offline state estimator and a second model learning module. The offline state estimator estimates an offline state and provides the offline state to the second model learning module. The policy learning module includes a Monte Carlo (MC)-based particle generation for estimating a cumulative cost predicted from a particle trajectory propagated through a learned model, and a model of an online state estimator configured to generate the online state estimate based on the particle measurement and a previous particle online estimate. The steps further include: A step in which the second model learning module provides the offline learning state to the policy learning module, and the policy learning module generates policy parameters using the online state estimate of the particles, Updating the policy of the system based on the updated policy parameters and operating the system. A controller including the above steps.
2. The second model learning module learns the behavior of the system using the speed integration model with the SE kernel, and / or the second model learning module generates the offline learning state and provides it to the policy learning module, and / or the policy learning module includes a policy optimization program, and the policy optimization program performs policy optimization based on the offline learning state from the second model learning module and generates the policy parameter. The controller according to claim 1.
3. When the policy learning module includes the policy optimization program, the policy optimization program performs policy optimization based on the offline state from the first model learning module and generates the policy parameter. The policy learning module includes a system model, and the system model generates a particle state based on a previous particle state and the control signal. The controller according to claim 2.
4. The policy learning module includes a sensor model configured to generate the particle measurement value based on the particle state. The controller according to claim 3.
5. The policy learning module includes a policy optimization unit configured to generate the policy parameter based on the particle measurement value and the particle online estimate. The policy optimization unit provides the policy parameter to update the policy unit of the system, and / or the policy optimization unit includes a dropout method and an early stopping strategy configured to improve the policy parameter generated by the policy optimization unit, and / or the offline state estimator is formed from a non-causal filter, a Kalman smoother or a central difference velocity approximator. The controller according to claim 1.
6. A vehicle control system for controlling the movement of a vehicle, comprising A vehicle control system comprising the controller of claim 1, wherein the controller is connected to a motion controller of the vehicle and a vehicle motion sensor for measuring the motion of the vehicle, the vehicle control system generates the policy parameter based on the measurement data of the motion, and the vehicle control system provides the policy parameter to the motion controller of the vehicle to update the policy unit of the motion controller.
7. The motion controller is configured to control the suspension of the vehicle and / or The motion controller is configured to control an actuator of the vehicle, the vehicle control system according to claim 6.
8. The second model learning module is configured to generate the offline learning state and provide it to the policy learning module, and the policy learning module generates the policy parameter, the vehicle control system according to claim 6.
9. The policy learning module includes a system model and a sensor model, and the sensor model generates the particle measurement value based on the particle state, the vehicle control system according to claim 8.
10. The model of the online state estimator is configured to generate the particle online estimated value based on the particle measurement value and the previous particle online estimated value, the vehicle control system according to claim 9.
11. A robot control system for controlling the motion of a robot, Comprising the controller of claim 1, the controller is connected to an actuator controller of the robot and a sensor configured to measure the state of the robot, the robot control system generates the policy parameter based on the measurement data of the sensor, and the robot control system provides the policy parameter to the actuator controller of the robot to update the policy unit of the actuator controller.
12. The actuator controller is configured to control at least one actuator of the robot, and / or the actuator controller is configured to control a plurality of actuators of the robot, and / or the first model learning module in the controller includes the offline state estimator and the second model learning module, and the offline state estimator estimates the offline state and provides the offline state to the second model learning module. The robot control system according to claim 11.
Citation Information
Patent Citations
controller
JP2008537271A