System and method for controlling a mechanical system having multiple degrees of freedom

A machine learning-based inverse dynamics model using Gaussian process regression effectively addresses the challenge of complex dynamics in robotic manipulators by modeling energy correlations, enhancing control accuracy and simplifying model formulation.

JP2026503817APending Publication Date: 2026-01-29MITSUBISHI ELECTRIC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025568407
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-07-17
Filing Date
2024-05-10
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Formulating an accurate inverse dynamics model for mechanical systems, such as robotic manipulators, is challenging due to parameter uncertainties and the inability to model complex dynamics like motor friction and joint elasticity, leading to inaccurate control.

Method used

Employing a machine learning-based inverse dynamics model using Gaussian process regression (GPR) to capture correlations between torques of different actuators by modeling the energy of the mechanical system, rather than individual torques, with a full covariance matrix that includes non-zero elements.

Benefits of technology

Enables accurate control of mechanical systems with multiple degrees of freedom by capturing mutual effects between torques, improving control accuracy and simplifying model formulation without requiring extensive physical knowledge of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026503817000001_ABST
    Figure 2026503817000001_ABST
Patent Text Reader

Abstract

The present disclosure provides a system and method for controlling a mechanical system having different actuators and multiple degrees of freedom to track a reference trajectory for performing a task. The system includes a memory configured to store an energy-based inverse dynamics model trained using machine learning to map states of the different actuators to corresponding torques of the different actuators, the energy-based inverse dynamics model being configured to model the energy of the mechanical system using a Gaussian process regression (GPR) process having a matrix that captures correlations between the torques of the different actuators. The system further includes a processor configured to process the states of the different actuators using the energy-based inverse dynamics model to generate torque values ​​for the different actuators of the mechanical system and control the mechanical system based on the generated torque values.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates generally to control systems for mechanical manipulators, and more particularly to systems and methods for controlling robotic manipulators having multiple degrees of freedom and different types of actuators to track a reference trajectory to perform a task. [Background technology]

[0002] A mechanical system, e.g., a robotic manipulator, is configured to track a reference trajectory to perform a task. The task may correspond, for example, to moving an object to a target position or to an assembly task. To control the mechanical system to track the reference trajectory, a dynamic model of the mechanical system is utilized. A dynamic model is a mathematical model that includes equations that define the dynamics of a mechanical system. Because an inverse dynamic model can control a mechanical system, some techniques use an inverse dynamic model of the mechanical system to control the mechanical system. The inverse dynamic model expresses the torque of a joint as a function of the position, velocity, and acceleration of the joint. To control the mechanical system, it is desirable to formulate an accurate inverse dynamic model.

[0003] However, formulating an accurate inverse dynamics model is difficult. For example, the accuracy of an inverse dynamics model is often limited by parameter uncertainties. Furthermore, it is difficult to model the complex dynamics of a mechanical system, such as the friction of a motor or the elasticity of a joint. Therefore, a method for formulating an accurate inverse dynamics model is needed to perform accurate control of a mechanical system. Summary of the Invention

[0004] An objective of some embodiments is to control a mechanical system using an inverse dynamics model trained through machine learning. In this disclosure, the forward dynamics model connects control commands and current states of different actuators with transition states of the different actuators achieved by executing the control commands. In this disclosure, the inverse dynamics model maps the current states and transition states of the different actuators to corresponding torques of the different actuators. The mechanical system is a robotic manipulator having different actuators and multiple degrees of freedom to track a reference trajectory for performing a task. For example, the robotic manipulator is configured to track a reference trajectory for performing a task of moving an object to a target position. The robotic manipulator includes joints and links. Each joint is actuated by an actuator, such as an electric motor. The action applied by the actuator moves the link attached to the joint.

[0005] Generally, generating an accurate physical model of inverse dynamics is difficult and time-consuming. The accuracy of such models is often limited by parameter uncertainty and an inability to describe certain complex dynamics typical of real systems, such as motor friction or joint elasticity. Therefore, in some embodiments, a machine learning algorithm is used to train the inverse dynamics model. An inverse dynamics model trained using a machine learning algorithm is called a data-driven model. In contrast to physical models, data-driven models do not require knowledge of the mechanical system and are trained directly based on experimental data. Despite their ability to approximate complex nonlinear dynamics, data-driven models suffer from low data efficiency and poor generalization properties. That is, the trained model is trained only in the vicinity of the training trajectory, requiring a large number of extrapolated samples.

[0006] Some embodiments are based on the recognition that training inverse dynamics models is difficult due to the lack of a structure for inverse dynamics models that cannot describe the physical principles governing the dynamics of multiple degrees of freedom of mechanical systems. For example, robotic manipulators have joints that define multiple degrees of freedom, and the joint dynamics are correlated with each other. Therefore, there may be complex and nonlinear correlations between the joint torques required to track a reference trajectory. Learning such correlations through training is difficult.

[0007] Some embodiments are based on the recognition that correlations can be learned using Gaussian process regression (GPR). For example, some Gaussian process-based solutions use GPR to model n torque components for each degree of freedom using n independent Gaussian processes (GPs), each including a model-based component for the mean function or covariance. As a result, the covariance matrix of such a GPR model is diagonal or block-diagonal, indicating that the correlations between the torques of different actuators are not captured.

[0008] Some embodiments are based on the recognition that such deficiencies are caused, at least in part, by attempts to model the torque itself as a Gaussian process. However, some embodiments are based on the recognition that Gaussian processes can be designed to model the energy of a mechanical system. Modeling the energy, as opposed to modeling individual torques, can capture the mutual effects that the torques of different actuators have on each other, thereby learning the correlations between the torques of different actuators. As a result, a covariance matrix capturing the correlations between the torques of different actuators is a full matrix with nonzero elements both inside and outside the diagonal.

[0009] To this end, some embodiments of the present disclosure provide an inverse dynamics model that models the energy of a mechanical system using GPR with complete prior and posterior covariance matrices that capture the correlation between the torques of different actuators. The inverse dynamics model is trained using machine learning to map the states of different actuators to corresponding torques of the different actuators. According to one embodiment, the inverse dynamics model processes the states of the different actuators to generate torque values ​​for the different actuators of the mechanical system. The mechanical system is then controlled based on the torque values. For example, control commands are determined based on the torque values ​​of the different actuators. The determined control commands are then applied to the different actuators to control the mechanical system.

[0010] Accordingly, one embodiment discloses a controller for controlling a mechanical system having different actuators and multiple degrees of freedom to track a reference trajectory for performing a task. The controller includes a memory configured to store an energy-based inverse dynamics model trained using machine learning to map states of the different actuators to corresponding torques of the different actuators. The energy-based inverse dynamics model is configured to model the energy of the mechanical system using a GPR process having a covariance matrix that captures correlations between the torques of the different actuators. The covariance matrix is ​​a full matrix that includes non-zero elements. The controller further includes a processor configured to generate torque values ​​for the different actuators of the mechanical system by processing the states of the different actuators using the energy-based inverse dynamics model, and control the mechanical system based on the generated torque values ​​for the different actuators of the mechanical system.

[0011] Accordingly, another embodiment discloses a method for controlling a mechanical system having different actuators and multiple degrees of freedom to track a reference trajectory for performing a task, the method using a processor connected to a memory storing an energy-based inverse dynamics model trained using machine learning to map states of the different actuators to corresponding torques of the different actuators. The energy-based inverse dynamics model is configured to model the energy of the mechanical system using a GPR process having a covariance matrix that captures correlations between the torques of the different actuators, the covariance matrix being a full matrix including non-zero elements. The processor is connected to stored instructions that, when executed by the processor, perform steps of the method. The method includes generating values ​​of the torques of the different actuators of the mechanical system by processing the states of the different actuators with the energy-based inverse dynamics model, and controlling the mechanical system based on the generated values ​​of the torques of the different actuators of the mechanical system.

[0012] Accordingly, yet another embodiment is a non-transitory computer-readable storage medium storing a program executable by a processor to perform a method for controlling a mechanical system having different actuators and multiple degrees of freedom to track a reference trajectory to perform a task, the storage medium storing an energy-based inverse dynamics model trained using machine learning to map states of the different actuators to corresponding torques of the different actuators, the energy-based inverse dynamics model being configured to model the energy of the mechanical system using a GPR process having a covariance matrix that captures correlations between the torques of the different actuators, the covariance matrix being a full matrix with non-zero elements. When executed by the processor, the program performs method steps including generating values ​​of the torques of the different actuators of the mechanical system by processing the states of the different actuators with the energy-based inverse dynamics model, and controlling the mechanical system based on the generated values ​​of the torques of the different actuators of the mechanical system. [Brief explanation of the drawings]

[0013] [Figure 1A] FIG. 1 is a block diagram illustrating a method for controlling a mechanical system according to some embodiments of the present disclosure. [Figure 1B] FIG. 1 illustrates a robotic manipulator, according to some embodiments of the present disclosure. [Figure 1C] FIG. 2 illustrates a diagonal covariance matrix in accordance with some embodiments of the present disclosure. [Figure 1D] FIG. 2 illustrates a full covariance matrix in accordance with some embodiments of the present disclosure. [Figure 2] FIG. 1 is a block diagram illustrating a method for formulating an inverse dynamics model in accordance with some embodiments of the present disclosure. [Figure 2 Cont] FIG. 10 is a block diagram illustrating a method for formulating an inverse dynamics model (continued), according to some embodiments of the present disclosure. [Figure 3]FIG. 10 is a block diagram illustrating training of hyperparameters of a Lagrangian polynomial kernel according to one embodiment of the present disclosure. [Figure 4A] 1 is a flow chart illustrating a method for estimating kinetic energy of a mechanical system, according to some embodiments of the present disclosure. [Figure 4B] 1 is a flow diagram illustrating a method for estimating potential energy of a mechanical system, according to some embodiments of the present disclosure. [Figure 5] 1 is a flow chart illustrating a method for detecting anomalies based on estimated kinetic and potential energies of a mechanical system, according to some embodiments of the present disclosure. [Figure 6] FIG. 1 illustrates a motion plan that consumes a minimal amount of energy to perform a task, according to some embodiments of the present disclosure. [Figure 7] FIG. 1 is a schematic diagram illustrating a computing device for implementing the controller and methods of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0014] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. It will be apparent to one skilled in the art that one or more embodiments may be practiced without these specific details. Additionally, devices and methods are shown as block diagrams in order to avoid obscuring the present disclosure.

[0015] As used in this specification and claims, the terms "for example," "for example," "such as," and "comprises," "has," "includes," and other verb forms thereof, when used in conjunction with a list of one or more components or other items, should be construed as open-ended, meaning not to exclude other additional components or items from this list. The term "based on" means based at least in part on. It should further be understood that the phraseology and terminology used herein are for descriptive purposes and should not be regarded as limiting. Any headings used herein are for convenience only and have no legal or limiting effect.

[0016] 1A is a block diagram 100 illustrating a controller 101 for controlling a mechanical system 109 in accordance with some embodiments of the present disclosure. The controller 101 is communicatively coupled to the mechanical system 109. The controller 101 includes a processor 103 and a memory 105. The processor 103 may be a single-core processor, a multi-core processor, a computing cluster, or any number of other configurations. The memory 105 may include random access memory (RAM), read-only memory (ROM), flash memory, or any other suitable memory system. Furthermore, in some embodiments, the memory 105 may be implemented using a hard drive, an optical drive, a thumb drive, a drive array, or any combination thereof.

[0017] In one embodiment, the mechanical system 109 corresponds to a robotic manipulator with different actuators and multiple degrees of freedom to track a reference trajectory to perform a task. An exemplary robotic manipulator is described below with reference to FIG. 1B.

[0018] FIG. 1B illustrates a robotic manipulator 113 according to some embodiments of the present disclosure. For example, the robotic manipulator 113 is configured to track a reference trajectory 115 for performing a task of moving an object 117 to a target position 119. In another example, the robotic manipulator 113 is configured to track a reference trajectory for performing an assembly task. The robotic manipulator 113 includes joints, such as joint 121a, joint 121b, and joint 121c. The robotic manipulator 113 further includes links, such as link 123a, link 123b, and link 123c. Each joint is actuated by an actuator, such as an electric motor. Actions applied by the actuators move the links attached to the joints. The movement may be rotational or translational depending on the type of joint. The type of joint may be a revolute joint or a prismatic joint. The actuator receives a desired position of the joint as input and outputs an action, which may be a current, a torque, or a quantity that can be converted into a torque that moves the joint to the desired position.

[0019] Additionally, the different actuators are equipped with position sensors (e.g., encoders) that can measure the current position of the joint. In some embodiments, the state of the different actuators is defined as the position of the joint. In some other embodiments, the state of the different actuators is defined as a combination of the joint position and the joint velocity.

[0020] An objective of some embodiments is to control a mechanical system 109 using an inverse dynamics model trained via machine learning. In this disclosure, the forward dynamics model connects control commands and current states of different actuators with the transition states of the different actuators achieved by executing the control commands. In this disclosure, the inverse dynamics model maps the current states and transition states of the different actuators to corresponding torques of the different actuators.

[0021] Generally, generating accurate physical models of inverse dynamics is difficult and time-consuming. The accuracy of such models is often limited by parameter uncertainties and an inability to describe certain complex dynamics typical of real systems, such as motor friction or joint elasticity. Therefore, in some embodiments, machine learning is used to train the inverse dynamics model. An inverse dynamics model trained using a machine learning algorithm is called a data-driven model. In contrast to physical models, data-driven models do not require knowledge of the mechanical system 109 and are learned directly from experimental data. Despite their ability to approximate complex nonlinear dynamics, data-driven models suffer from low data efficiency and poor generalization properties. That is, the trained model is trained only in the vicinity of the training trajectory, requiring a large number of extrapolated samples.

[0022] Some embodiments are based on the recognition that training an inverse dynamics model is difficult due to the lack of a structure for the inverse dynamics model that cannot describe the physical principles governing the dynamics of multiple degrees of freedom of the mechanical system 109. For example, a robotic manipulator 113 has joints that define multiple degrees of freedom, and the joint dynamics are correlated with each other. Thus, there can be complex, nonlinear correlations between the joint torques required to track a reference trajectory. Learning such correlations is difficult.

[0023] Some embodiments are based on the recognition that correlations can be learned using Gaussian process regression (GPR). For example, some Gaussian process-based solutions use GPR to model n torque components for each degree of freedom using n independent GPs, each with a model-based component for the mean function or covariance. As a result, the covariance matrix of such a GPR model is diagonal or block-diagonal, indicating that correlations between torques of different actuators are not captured. An example of such a covariance matrix is ​​shown in FIG. 1C below.

[0024] 1C is an exemplary schematic diagram illustrating a diagonal covariance matrix 125 according to some embodiments of the present disclosure. The diagonal elements 127 of the diagonal covariance matrix 125, i.e., a, b, and c, are non-zero elements, and the off-diagonal elements 129 are zero. Because the off-diagonal elements 129 of the diagonal covariance matrix 125 are zero, the diagonal covariance matrix 125 indicates that the correlation between the torques of different actuators is not captured.

[0025] Some embodiments are based on the recognition that such deficiencies are caused, at least in part, by attempts to model the torque itself as a Gaussian process. However, some embodiments are based on the recognition that a Gaussian process can be designed to model the energy of a mechanical system 109. Modeling the energy, as opposed to modeling individual torques, can capture the mutual effects that the torques of different actuators have on each other, thereby learning the correlations between the torques of different actuators. As a result, a covariance matrix capturing the correlations between the torques of different actuators is a full matrix with nonzero elements both inside and outside the diagonal.

[0026] To this end, some embodiments of the present disclosure provide an energy-based inverse dynamics model 107 that models the energy of a mechanical system 109 using GPR with a covariance matrix that captures the correlation between the torques of different actuators. The covariance matrix is ​​a complete matrix that includes non-zero elements.

[0027] 1D illustrates a full covariance matrix 131 according to some embodiments of the present disclosure. The diagonal elements 133, i.e., a1, b2, and c3, and the off-diagonal elements 135, i.e., a2, a3, b1, b3, c1, and c2, of the full covariance matrix 131 are non-zero. Because the elements of the full covariance matrix 131 are non-zero, the full covariance matrix 131 captures the correlation between the torques of different actuators.

[0028] In one embodiment, the energy-based inverse dynamics model 107 is stored in the memory 105. The energy-based inverse dynamics model 107 is trained using machine learning to map different actuator states to corresponding torques of the different actuators. Hereinafter, the energy-based inverse dynamics model 107 will be referred to as the inverse dynamics model 107.

[0029] The processor 103 is configured to generate torque values ​​for different actuators of the mechanical system 109 by processing the states of the different actuators using the inverse dynamics model 107. For example, the states of the different actuators are applied as inputs to the inverse dynamics model 107. The inverse dynamics model 107 generates torque values ​​for the different actuators of the mechanical system 109 based on the states of the different actuators. The processor 103 is further configured to control the mechanical system 109 based on the torque values. For example, in one embodiment, the processor 109 determines control commands based on the torque values ​​of the different actuators. The determined control commands are then applied to the different actuators. The control commands modify the states of the different actuators to track the reference trajectory. The modified states of the different actuators are input to the feedback controller 111. The feedback controller 111 is configured to accurately track the reference trajectory by correcting the torque values ​​to compensate for errors in the position and velocity of the joints based on the modified states of the different actuators.

[0030] The formulation of the inverse dynamics model 107 used to generate the torque values ​​is described below with reference to FIG.

[0031] 2 is a flow diagram illustrating a method 200 for formulating the inverse dynamics model 107 according to some embodiments of the present disclosure. Generally, GPR can learn the inverse dynamics of a mechanical system by modeling the energy of the mechanical system 109 using different types of kernels. Some embodiments use physics-informed energy kernels that define the energy of the mechanical system 109. The energy kernels improve the accuracy of the inverse dynamics model 107. For example, the energy kernel defines the energy using a kinetic energy function and a potential energy function. Such kernels advantageously take into account the law of conservation of energy of the mechanical system 109.

[0032]

number

[0033]

number

[0034]

number

number

number

[0035]

number

[0036]

number

[0037]

number

[0038]

number

[0039]

number

[0040]

number

[0041]

number

[0042]

number

[0043]

number

[0044]

number

[0045]

number

[0046]

number

[0047]

number

[0048]

number

[0049]

number

[0050]

number

[0051]

number

[0052] According to one embodiment, one or more hyperparameters of the Lagrangian polynomial kernel are determined based on a machine learning process, as described below with reference to FIG.

[0053]

number

[0054]

number

[0055] According to some embodiments, the inverse dynamics model 107 is a multiple-input multiple-output (MIMO) torque estimator model that generates torques for different actuators based on the joint positions, velocities and accelerations of multiple degrees of freedom of the mechanical system 109. In other words, the joint positions, velocities and accelerations of multiple degrees of freedom are applied as inputs to the MIMO torque estimator, which outputs torques for different actuators.

[0056] Additionally or alternatively, in some embodiments, the inverse dynamics model 107 is used to estimate the kinetic energy and potential energy from the torque measurements. For example, the processor 103 is configured to estimate the potential energy of the mechanical system 109 and the kinetic energy of the mechanical system 109 by processing the states of different actuators using the inverse dynamics model 107. The estimation of potential energy is described in more detail below with reference to FIG. 4A, and the estimation of potential energy is described in more detail below with reference to FIG. 4B.

[0057]

number

[0058]

number

[0059] At block 403 of the method 400, the processor 103 is configured to calculate a posterior probability distribution of the kinetic energy.

[0060]

number

[0061]

number

[0062]

number

[0063] In block 411 of the method 407, the processor 103 is configured to calculate a posterior probability distribution of the potential energy.

[0064]

number

[0065]

number

[0066]

number

[0067] Some embodiments are based on the recognition that the estimated kinetic and potential energy can be used to detect anomalies in the mechanical system 109 during operation (e.g., while performing a task) of the mechanical system 109. Anomaly detection based on the estimated kinetic and potential energy of the mechanical system 109 is described below with reference to FIG.

[0068] FIG. 5 is a flow chart illustrating a method 500 for detecting anomalies based on estimated kinetic and potential energies of a mechanical system 109, according to some embodiments of the present disclosure.

[0069] In block 501, the processor 103 is configured to estimate the kinetic and potential energy of the mechanical system 109, as described above with reference to Figures 4A and 4B.

[0070] The processor 103 is also configured to compare the estimated kinetic energy and the estimated potential energy with corresponding thresholds. In other words, the processor 103 is configured to compare the estimated kinetic energy with a first threshold and compare the estimated potential energy with a second threshold. Based on such comparisons, an anomaly is detected. For example, in block 503, the processor 103 is configured to determine whether the estimated kinetic energy and the estimated potential energy are greater than the corresponding thresholds. If the estimated kinetic energy and the estimated potential energy are greater than the first threshold and the second threshold, respectively, it is inferred that an anomaly has been detected in block 505. If the estimated kinetic energy and the estimated potential energy are not greater than the first threshold and the second threshold, respectively, it is inferred that an anomaly has not been detected in block 507.

[0071] Alternatively, in some embodiments, the processor 103 is configured to determine whether the estimated kinetic energy and estimated potential energy are less than a first threshold and a second threshold, respectively. If the estimated kinetic energy and estimated potential energy are less than the first threshold and the second threshold, respectively, it is inferred that an anomaly is detected. If the estimated kinetic energy and estimated potential energy are not less than the first threshold and the second threshold, respectively, it is inferred that an anomaly is not detected. If the estimated kinetic energy is greater than the first threshold and the estimated potential energy is less than the second threshold, it is inferred that the fault detection is a special anomaly detection case that depends only on a potential energy fault. If the estimated potential energy is greater than the second threshold and the estimated kinetic energy is less than the first threshold, it is inferred that the fault detection is a special anomaly detection case that depends only on a kinetic energy fault.

[0072] Further, in some embodiments, the estimated potential energy and the estimated kinetic energy are used to adjust the control commands of the mechanical system 109. For example, a passive controller can be designed to control the mechanical system 109 based on the estimated potential energy and the estimated kinetic energy. Passive controllers are effective in controlling mechanical systems, but require an accurate energy model to define the Hamiltonian or Lagrangian function of the mechanical system 109. Some embodiments can estimate an accurate energy model that results in an accurate passive controller.

[0073] In some other embodiments, a trajectory for performing a task is tracked based on estimated inverse dynamics. For example, the processor 103 is configured to determine a controller capable of tracking a trajectory for performing a task based on torques that are output from an inverse dynamics model. The task may be an assembly task, a polishing task, or a pick-and-place task.

[0074] Additionally, in one embodiment, the processor 103 is further configured to determine a motion plan that consumes a minimum amount of energy to perform the task based on the estimated kinetic energy and the estimated potential energy. The motion plan may include a trajectory for performing the task. Such an embodiment is described below with reference to FIG. 6.

[0075] FIG. 6 illustrates a motion plan 609 for the robotic manipulator 113 to consume a minimum amount of energy to perform a task, according to some embodiments of the present disclosure. The motion plan 609 is a trajectory tracked by the robotic manipulator 113 to perform the task. An objective of some embodiments is to configure the robotic manipulator 113 to perform the task of inserting an object 601 into a hole 603. The motion plan is determined for the robotic manipulator 113 to perform the task. Different motion plans, such as motion plan 605 and motion plan 607, may be determined to perform the task. However, it is desirable to determine a motion plan that consumes a minimum amount of energy to perform the task. The processor 103 is configured to determine the motion plan 609 that consumes a minimum amount of energy to perform the task based on the estimated kinetic energy and estimated potential energy. Furthermore, the processor 103 is configured to control the robotic manipulator 113 according to the motion plan 609 to perform the task with a minimum amount of energy consumption. For example, the processor 103 is configured to determine control commands for different actuators of the robotic manipulator 113 based on the motion plan 609. The control commands are applied to the different actuators of the robotic manipulator 113. The control commands cause the robotic manipulator 113 to follow the motion plan 609 to perform a task.

[0076] 7 is a schematic diagram illustrating a computing device 700 for implementing the controller 101 and method of the present disclosure. The computing device 700 includes a power supply 701, a processor 703, a memory 705, and a storage device 707, all connected to a bus 709. A high-speed interface 711, a low-speed interface 713, a high-speed expansion port 715, and a low-speed connection port 717 may also be connected to the bus 709. A low-speed expansion port 719 is also connected to the bus 709. An input interface 721 may also be connected to an external receiver 723 and an output interface 725 via the bus 709. The receiver 727 may also be connected to an external transmitter 729 and a transmitter 731 via the bus 709. An external memory 733, an external sensor 735, a machine 737, and an environment 739 may also be connected to the bus 709. One or more external input / output devices 741 may also be connected to the bus 709. A network interface controller (NIC) 743 may be adapted to connect to a network 745 via bus 709. Data or other data may be rendered external to computing device 700 on a third party display device, a third party imaging device, and / or a third party printing device.

[0077] The memory 705 may store instructions executable by the computing device 700 and any data that may be utilized by the methods and systems of the present disclosure. The memory 705 may include random access memory (RAM), read-only memory (ROM), flash memory, or any other suitable memory system. The memory 705 may be a volatile and / or non-volatile memory unit. The memory 705 may also be another form of computer-readable medium, such as a magnetic disk or an optical disk.

[0078] The storage device 707 may be configured to store supplemental data and / or software modules used by the computing device 700. The storage device 707 may include a hard drive, an optical drive, a thumb drive, a drive array, or any combination thereof. Additionally, the storage device 707 may include a computer-readable medium such as a floppy disk drive, a hard disk drive, an optical disk drive, a tape drive, a flash memory or other similar solid-state memory device, or an array of devices including devices in a storage area network or other configuration. The instructions may be stored on an information carrier. When executed by one or more processing devices (e.g., processor 703), the instructions perform one or more of the methods described above.

[0079] Optionally, computing device 700 can be connected via bus 709 to a display or user interface (HMI) 747 configured to connect computing device 700 to a display device 749 and a keyboard 751. Display device 749 can include, among other things, a computer monitor, a camera, a television, a projector, or a mobile device. In some implementations, computing device 700 can include a printer interface for connecting to a printing device, which can include, among other things, a liquid inkjet printer, a solid ink printer, a large-scale commercial printer, a thermal printer, a UV printer, or a dye-sublimation printer.

[0080] The high-speed interface 711 manages bandwidth-intensive operations for the computing device 700, and the low-speed interface 713 manages less bandwidth-intensive operations. This allocation of functionality is merely exemplary. In some implementations, the high-speed interface 711 may be connected to memory 705, a user interface (HMI) 747, a keyboard 751 and a display 749 (e.g., via a graphics processor or accelerator), and a high-speed expansion port 715 that may accept various expansion cards via a bus 709. In one implementation, the low-speed interface 713 is connected to storage device 707 and a low-speed expansion port 717 via a bus 709. The low-speed expansion port 717, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may be connected to one or more input / output devices 741. The computing device 700 may be connected to a server 753 and a rack server 755. The computing device 700 may be implemented in several different forms. For example, the computing device 700 may be implemented as part of a rack server 755 .

[0081] The present disclosure provides a controller 101 configured to control a mechanical system 109 using an inverse dynamics model 107. The inverse dynamics model 107 models the energy of the mechanical system 109. The energy modeling captures the mutual effects of torques of different actuators. Thus, the inverse dynamics model 107 enables accurate control of the mechanical system 101. Furthermore, formulating the inverse dynamics model 107 requires minimal physical information about the mechanical system 109. Therefore, formulating the inverse dynamics model 107 is simpler. Additionally or alternatively, the inverse dynamics model 107 can be used to estimate the kinetic energy of the mechanical system 109 and the potential energy of the mechanical system 109.

[0082] The description provides exemplary embodiments only and is not intended to limit the scope, application, or configuration of the present disclosure. Rather, the following description of exemplary embodiments will provide those skilled in the art with an enabling description for implementing one or more exemplary embodiments. Various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the subject matter as set forth in the appended claims.

[0083] In the following description, specific details are given to provide a thorough understanding of the embodiments. However, those skilled in the art will understand that the embodiments can be practiced without these specific details. For example, systems, processes, and other elements in the disclosed subject matter may be shown as components in block diagrams so as not to obscure the embodiments in unnecessary detail. Also, well-known processes, structures, and techniques may be shown without unnecessary detail so as not to obscure the embodiments. Furthermore, like reference numbers and names in the various drawings refer to like elements.

[0084] Each embodiment may also be described as a process, which is depicted as a flowchart, flow diagram, data flow diagram, structure diagram, or block diagram. Although a flowchart may describe operations as a sequential process, many operations may be performed in parallel or simultaneously. The order of operations may also be changed. A process may be terminated when its operations are completed, but the process may include additional steps not discussed or shown. Furthermore, not all operations within a specifically described process need be included in all embodiments. A process may be a method, a function, a procedure, a subroutine, a subprogram, etc. When a process is a function, the termination of the function corresponds to the function returning to the calling function or the main function.

[0085] Furthermore, embodiments of the disclosed subject matter may be implemented, at least in part, manually or automatically. The manual or automatic implementation may be implemented, or at least assisted, by machine, hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented by software, firmware, middleware, or microcode, the program code or code segments to perform the necessary tasks may be stored on a machine-readable medium. A processor may perform the necessary tasks.

[0086] The various methods or steps outlined herein may be coded as software executable on one or more processors employing any one of a variety of operating systems or platforms. Furthermore, such software may be written using any of a number of suitable programming languages ​​and / or programming or scripting tools, and compiled as executable machine language code or intermediate code that runs on a framework or virtual machine. Typically, the functionality of the program modules may be combined or distributed in various embodiments as desired.

[0087] The embodiments of the present disclosure may be embodied as methods, which are provided by way of example. The operations performed as part of the method may be ordered in any suitable manner. Thus, embodiments may be constructed that perform operations in an order different from the operations performed sequentially in the exemplary embodiments, and that may include performing some operations simultaneously.

[0088] Furthermore, embodiments of the present disclosure and the functional operations described herein may be implemented in digital electronic circuitry, tangible computer software or firmware, computer hardware, or one or more combinations thereof, including the structures disclosed in this disclosure and their structural equivalents. Some embodiments of the present disclosure may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier and executed by or controlling the operation of a data processing device. Furthermore, the program instructions may be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical signal generated by encoding information for transmission to a suitable receiver for execution by a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or one or more combinations thereof.

[0089] The term "computing system" includes all types of equipment, devices, and machines for processing data, such as a programmable processor, computer, or multiple processors or computers. The device may also be or include special-purpose logic circuitry, such as an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). In addition to hardware, the device may include code that creates an execution environment for a computer program, such as code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof.

[0090] A computer program (which may be called or written as a program, software, software application, module, software module, script, or code) can be written in any programming language, including compiled or interpreted, declarative or procedural, and can be used as a stand-alone program or in any form as a module, component, subroutine, object, or other unit suitable for use within a computing environment. A computer program can, but need not, correspond to a file in a file system. A program may be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program involved, or in multiple synchronous files (e.g., a file storing one or more modules, subprograms, or portions of code).

[0091] A computer program can be implemented running on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communications network. Computers suitable for executing computer programs may, by way of example, be based on general-purpose or special-purpose microprocessors or both, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from a read-only memory or a random-access memory or both. Some elements of a computer are the central processing unit for executing or carrying out instructions and one or more memory units for storing instructions and data.

[0092] Typically, a computer also includes one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, and / or is operatively coupled to transmit and receive data from these mass storage devices. However, a computer need not have these devices. Additionally, a computer may include another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive).

[0093] To provide for user interaction, embodiments of the subject matter described herein may be implemented on a computer that includes a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, a keyboard, and a pointing device, e.g., a mouse or trackball, by which the user can provide input to the computer. Other types of devices may also be used to provide for user interaction. For example, feedback provided to the user may be any type of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback, and input from the user may be received in any form, including acoustic, speech, or tactile input. The computer may also interact with the user by sending and receiving documents to and from devices used by the user, e.g., by sending web pages to a web browser on a user client device in response to a request received from the web browser.

[0094] Embodiments of the subject matter described herein can be implemented in a computing system that includes a back-end component, e.g., a data server, or includes a middleware component, e.g., an application server, or includes a front-end component, e.g., a client computer with a graphical user interface, or a web browser through which a user can interact with an implementation of the subject matter described herein, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), e.g., the Internet.

[0095] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically exchange information through a communication network. The relationship of clients and servers depends upon computer programs running on the respective computers and having a client-server relationship to each other.

[0096] Although the invention has been described with reference to preferred embodiments, it should be understood that various other modifications and variations can be made within the spirit and scope of the invention. It is therefore intended in the appended claims to cover all such variations and modifications which fall within the true spirit and scope of the invention.

Claims

1. 1. A controller for controlling a mechanical system having different actuators and multiple degrees of freedom to track a reference trajectory for performing a task, the controller comprising: a memory configured to store an energy-based inverse dynamics model trained using machine learning to map states of the different actuators to corresponding torques of the different actuators, the energy-based inverse dynamics model being configured to model the energy of the mechanical system using a Gaussian process having a covariance matrix that captures correlations between the torques of the different actuators, the covariance matrix being a full matrix including non-zero elements; and the controller: a processor; The processor: generating values ​​of the torques of the different actuators of the mechanical system by processing the states of the different actuators with the energy-based inverse dynamics model; a controller configured to control the mechanical system based on the produced values ​​of the torques of the different actuators of the mechanical system.

2. The controller of claim 1 , wherein the energy-based inverse dynamics model is defined by a Lagrangian polynomial kernel based on at least a Lagrangian operator, a kernel function of kinetic energy of the mechanical system, and a kernel function of potential energy of the mechanical system.

3. The controller of claim 2 , wherein the Lagrangian operator is defined based on a set of partial differential equations.

4. The controller of claim 3 , wherein the Lagrangian operator maps a Lagrangian function of the mechanical system to the torques of the different actuators.

5. The controller of claim 4 , wherein the Lagrangian function is defined based on a difference between the kinetic energy of the mechanical system and the potential energy of the mechanical system.

6. The controller of claim 2 , wherein the kernel function of the kinetic energy of the mechanical system is a polynomial function in a space defined by a trigonometric transformation of the state of the mechanical system.

7. The controller of claim 2 , wherein the kernel function of the potential energy of the mechanical system is a polynomial function in a space defined by a trigonometric transformation of the state of the mechanical system.

8. one or more hyperparameters of the Lagrangian polynomial kernel are learned based on a machine learning algorithm; The controller of claim 2 , wherein the machine learning algorithm uses maximization of marginal likelihood.

9. 2. The controller of claim 1, wherein the energy-based inverse dynamics model is a multiple-input multiple-output (MIMO) torque estimator model that generates torques for the different actuators based on positions of multiple joints of the mechanical system and velocities and accelerations of the multiple degrees of freedom.

10. 3. The controller of claim 2, wherein the processor is further configured to estimate potential energy of the mechanical system and kinetic energy of the mechanical system by processing the states of the different actuators with the energy-based inverse dynamics model.

11. The controller of claim 10 , wherein the processor is further configured to compare the estimated kinetic energy and the estimated potential energy with corresponding thresholds.

12. The controller of claim 11 , wherein the processor is further configured to detect an anomaly in the mechanical system based on the comparison of the estimated kinetic energy and the estimated potential energy with the corresponding thresholds.

13. The controller of claim 10 , wherein the processor is further configured to determine a motion plan that consumes a minimal amount of energy to perform the task based on the estimated kinetic energy and the estimated potential energy.

14. 1. A method for controlling a mechanical system having different actuators and multiple degrees of freedom to track a reference trajectory to perform a task, the method using a processor connected to a memory that stores an energy-based inverse dynamics model trained using machine learning to map states of the different actuators to corresponding torques of the different actuators, the energy-based inverse dynamics model being configured to model the energy of the mechanical system using a Gaussian Process Regression (GPR) process having a covariance matrix that captures correlations between the torques of the different actuators, the covariance matrix being a full matrix including non-zero elements, the processor being connected to stored instructions that, when executed by the processor, perform steps of the method, the method comprising: generating values ​​of the torques of the different actuators of the mechanical system by processing the states of the different actuators with the energy-based inverse dynamics model; and controlling the mechanical system based on the produced values ​​of the torques of the different actuators of the mechanical system.

15. 15. The method of claim 14, wherein the energy-based inverse dynamics model is defined by a Lagrangian polynomial kernel based on at least a Lagrangian operator, a kernel function of the kinetic energy of the mechanical system, and a kernel function of the potential energy of the mechanical system.

16. The method of claim 15 , wherein the Lagrangian operator is defined based on a set of partial differential equations.

17. The method of claim 15 , wherein the Lagrangian operator maps a Lagrangian function of the mechanical system to the torques of the different actuators.

18. The method of claim 16 , wherein the Lagrangian function is defined based on the difference between the kinetic energy of the mechanical system and the potential energy of the mechanical system.

19. The method of claim 15 , wherein the kernel function of the kinetic energy of the mechanical system is a polynomial function in a space defined by a trigonometric transformation of the state of the mechanical system.

20. one or more hyperparameters of the Lagrangian polynomial kernel are learned based on a machine learning algorithm; The method of claim 15 , wherein the machine learning algorithm uses maximization of marginal likelihood.

21. 1. A non-transitory computer-readable storage medium storing a program executable by a processor to perform a method for controlling a mechanical system having different actuators and multiple degrees of freedom to track a reference trajectory to perform a task, the storage medium storing an energy-based inverse dynamics model trained using machine learning to map states of the different actuators to corresponding torques of the different actuators, the energy-based inverse dynamics model being configured to model the energy of the mechanical system using a Gaussian Process (GPR) process having a covariance matrix that captures correlations between the torques of the different actuators, the covariance matrix being a full matrix including non-zero elements, the program, when executed by the processor, performing steps of the method, the method comprising: generating values ​​of the torques of the different actuators of the mechanical system by processing the states of the different actuators with the energy-based inverse dynamics model; and controlling the mechanical system based on the generated values ​​of the torques of the different actuators of the mechanical system.