Robot, control method, and training method

WO2026204441A1PCT designated stage Publication Date: 2026-10-01SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/009787
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2026-03-12
Publication Date
2026-10-01

Smart Images

  • Figure JP2026009787_01102026_PF_FP_ABST
    Figure JP2026009787_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present technology relates to a robot, a control method, and a training method that make it possible to easily improve the precision of actions performed by the robot. This robot comprises a plurality of sensors, each of which detects acceleration and angular velocity, the robot acting in accordance with action information generated by imparting, to a learning model, observation values observed by the robot, the observation values including the acceleration and the angular velocity detected by the plurality of sensors. The learning model is trained using, as training data, observation values observed by the robot, the observation values including acceleration and angular velocity obtained through physical simulation in which the plurality of sensors of the robot are reproduced. The present technology can be applied to, e.g., a real-world robot or the like.
Need to check novelty before this filing date? Find Prior Art

Description

Robot, control method, and learning method

[0001] The present technology relates to a robot, a control method, and a learning method, and particularly relates to a robot, a control method, and a learning method that can easily improve the accuracy of the robot's behavior, for example.

[0002] There is a learning method called sim-to-real, in which content learned by a robot in a simulation environment is applied to a real robot, which is a real-world robot, thereby eliminating or reducing the need for the real robot to learn in the real world.

[0003] According to sim-to-real, data can be collected faster than in the real world, and damage to the real robot and the real-world environment during learning can be suppressed. Furthermore, according to sim-to-real, by optimizing behavior (motion) control policies (behavior control means) such as those of a neural network based on reinforcement learning theory, the behavioral performance of the robot can be appropriately elicited for various control objectives.

[0004] However, in sim-to-real, when the difference between the simulation environment and the real world is unignorably large, it may be difficult to apply the learning result obtained in the simulation environment to the real robot. For example, in the real world, friction may occur in drive units such as motors (actuators) of the robot, and delay in the movement of the drive units may occur. When a learning result obtained in a simulation environment that does not reflect such friction and delay of the drive unit is applied to a real robot, it becomes difficult for the real robot to perform appropriate behavior.

[0005] Accordingly, there is a method in which a robot learns behavior (motion) for accomplishing a task in a simulation environment, and learns control values of a drive unit necessary for realizing the behavior learned in the simulation environment in the real world (see, for example, Patent Document 1).

[0006] Japanese Patent No. 7398373

[0007] Currently, there is a demand for proposals for technologies that can easily improve the accuracy of the actions of actual robots.

[0008] This technology was developed in light of these circumstances and aims to easily improve the accuracy of the actions of actual robots.

[0009] The robot of this technology is equipped with multiple sensors, each of which detects acceleration and angular velocity. The robot acts according to the behavioral information generated by providing the observed values, including the acceleration and angular velocity detected by the multiple sensors, to a learning model that takes the observed values ​​as input and outputs behavioral information representing the robot's actions. This learning model is trained using observed values, including the acceleration and angular velocity obtained from a physical simulation that reproduces the multiple sensors of the robot, as training data.

[0010] In the robot of this technology, actions are taken according to the action information generated by providing the observed values, including acceleration and angular velocity detected by the multiple sensors, to a learning model that outputs action information representing the actions of the robot, which has been trained using observed values, including acceleration and angular velocity, obtained by a physical simulation that reproduces the multiple sensors, as training data. The learning model takes the observed values ​​as input and outputs action information representing the actions of the robot.

[0011] The control method of this technology is a robot control method that includes generating behavioral information by providing the observed values, including acceleration and angular velocity detected by the multiple sensors, to a learning model that outputs behavioral information representing the robot's actions, which is trained using observed values ​​observed by the robot as training data, and which takes the observed values ​​as input. The learning model is trained using the observed values, including acceleration and angular velocity detected by the multiple sensors, which are each equipped with multiple sensors that detect acceleration and angular velocity. The method also includes controlling the robot's actions according to the behavioral information.

[0012] In the control method of this technology, the observed values ​​observed by the robot, including acceleration and angular velocity obtained from a physical simulation that reproduces the multiple sensors of a robot each equipped with multiple sensors that detect acceleration and angular velocity, are used as training data to train a learning model that outputs behavioral information representing the robot's actions by taking the observed values ​​as input. By providing the observed values, including acceleration and angular velocity detected by the multiple sensors, the behavioral information is generated. Then, the robot's actions are controlled according to the behavioral information.

[0013] The learning method of this technology includes generating observed values ​​observed by the robot, including acceleration and angular velocity detected by the multiple sensors, by performing a physical simulation that reproduces the multiple sensors of a robot, each equipped with multiple sensors that detect acceleration and angular velocity, and training a learning model that uses the observed values ​​as training data and outputs behavioral information representing the robot's actions, taking the observed values ​​as input.

[0014] In the learning method of this technology, a physical simulation is performed that reproduces the multiple sensors of a robot, each equipped with multiple sensors that detect acceleration and angular velocity. This generates observed values ​​observed by the robot, including the acceleration and angular velocity detected by the multiple sensors. Then, using these observed values ​​as training data, a learning model is trained that takes these observed values ​​as input and outputs behavioral information representing the robot's actions.

[0015] The control method and learning method can be implemented by having a computer execute a program. The program can be provided by recording it on a recording medium or by transmitting it via a transmission medium.

[0016] This is a diagram illustrating the overview of this technology. This is a block diagram showing an example configuration of a learning device that performs learning of the learning model 12. This is a flowchart illustrating an example of the learning process of the learning device 40 for learning the learning model 12. This is a block diagram showing an example configuration of a first embodiment of a robot system to which this technology is applied. This is a flowchart illustrating an example of the action control process that controls the actions of the actual robot 10 by the robot control device 60. This is a block diagram showing an example configuration of a second embodiment of a robot system to which this technology is applied. This is a diagram illustrating the overview of a third embodiment of a robot system to which this technology is applied. This is a block diagram showing an example of the electrical configuration of a robot system 130. This is a block diagram showing an example configuration of one embodiment of a computer to which this technology is applied.

[0017] <Overview of this technology>

[0018] Figure 1 is a diagram illustrating the overview of this technology.

[0019] This technology relates to actual robots.

[0020] In Figure 1, the actual robot 10 has at least a plurality of acceleration / angular velocity sensors 11, each detecting (translational) acceleration and angular acceleration. The plurality of acceleration / angular velocity sensors 11 are attached to different parts (locations) of the actual robot 10. A single acceleration / angular velocity sensor 11 may be a set of physically separate sensors, such as a set of an acceleration sensor that detects acceleration and an angular velocity sensor that detects angular velocity, or it may be a single sensor such as an IMU (inertial measurement unit) that detects both acceleration and angular acceleration. In this embodiment, as the acceleration / angular velocity sensor 11, for example, an IMU that detects (measures) the acceleration and angular velocity at the mounting position in the IMU's local coordinate system is used.

[0021] The actual robot 10 acts according to the action information (commands to the drive unit of the actual robot 10 obtained from it (drive unit commands)) that represents the actions of the actual robot 10, supplied from the learning model 12. The actual robot 10 also supplies the learning model 12 with various observed values ​​observed by the actual robot 10, including acceleration and angular velocity detected by multiple acceleration / angular velocity sensors 11.

[0022] The learning model 12 constitutes an action control strategy (means) for controlling the actions of the actual robot 10. The learning model 12 can be composed of, for example, a general neural network such as a CNN (convolutional neural network), or a machine learning model such as a GAN (Generative Adversarial Network) or a VAE (Variational Autoencoder). When the learning model 12 is given observed values ​​observed by the actual robot 10, including acceleration and angular velocity detected by multiple acceleration / angular velocity sensors 11 of the actual robot 10, as input, it generates (calculates) action information representing the actions of the actual robot 10 and outputs it. The action information is, for example, a vector whose elements are values ​​corresponding to the joint positions of each joint of the actual robot 10.

[0023] As input to the learning model 12, in addition to the acceleration and angular velocity detected by the multiple acceleration / angular velocity sensors 11 of the actual robot 10, any observed values ​​(any information obtainable by the actual robot 10) observable by the actual robot 10 can be used. For example, physical quantities detected by any sensors other than the multiple acceleration / angular velocity sensors 11 of the actual robot 10 can be included as input to the learning model 12.

[0024] The learning model 12 is trained using sim-to-real.

[0025] In other words, during the training of the learning model 12, the sensors of the actual robot 10, including multiple acceleration / angular velocity sensors 11, are reproduced on a virtual robot 20, which is a virtual robot corresponding to the actual robot 10, in the physical simulation. The multiple acceleration / angular velocity sensors 11 reproduced in the physical simulation are referred to as the virtual sensor 21.

[0026] In the physical simulation, observed values ​​(time series) observed by the virtual robot 20 (and by extension, the actual robot 10), including acceleration and angular velocity detected by the virtual sensor 21, are generated and collected.

[0027] Then, the learning model 12 is trained, for example, using machine learning such as reinforcement learning (including deep learning), with the observed values ​​obtained from the physical simulation used as training data.

[0028] The machine learning method for learning model 12 can be arbitrarily selected from reinforcement learning, supervised learning, and other methods. Furthermore, the machine learning of learning model 12 can be performed by combining multiple machine learning methods.

[0029] For example, by using observed values ​​obtained from a physical simulation as training data, reinforcement learning can be performed on the first learning model, and supervised learning (distillation) of the second learning model can be performed using the first learning model as the teacher. That is, supervised learning of the second model can be performed using the behavioral information obtained by inputting observed values ​​into the first learning model and the observed values ​​input into the first learning model as training data. In this case, the trained second model is used as the learning model 12 to control the actual robot 10. Since the first learning model is trained using observed values ​​obtained from a physical simulation as training data, it can be said that the training of the second model as the learning model 12, which is trained using such a first learning model as the teacher, is also indirectly performed using observed values ​​obtained from a physical simulation as training data.

[0030] For example, if the actual robot 10 is a walking robot and an IMU acting as an acceleration / angular velocity sensor 11 is attached to the toe of the walking robot, the acceleration and angular velocity of the toe can be directly detected. In this case, the learning model 12 can be trained using the acceleration and angular velocity of the toe detected by the (virtual) IMU reproduced by the physical simulation. Reproducing the IMU in a physical simulation is easier than modeling the motors, gears, and other drive components that drive the actual robot 10.

[0031] In physical simulations, it is desirable to reproduce the characteristics of the IMU (and the acceleration and angular velocity it detects) as accurately as possible. These characteristics include, for example, the observation noise of the acceleration and angular velocity detected by the IMU, the IMU offset (the mounting position of the IMU, the attitude of the IMU, and the offset of the acceleration and angular velocity detected by the IMU), the upper and lower limits of the measurable values, and the delay.

[0032] The actions of the actual robot 10 are controlled by the learning model 12, which has been trained using the sim-to-real method described above. In other words, the actual robot 10 acts according to the action information generated by inputting various observed values ​​observed by the actual robot 10, including acceleration and angular velocity detected by multiple acceleration / angular velocity sensors 11, into the learning model 12.

[0033] In the actual robot 10, the learning model 12 generates behavioral information (inference), but does not require retraining (fine-tuning). That is, if the learning model 12 is trained using observed values, including acceleration and angular velocity, detected by virtual sensors 21 that reproduce multiple acceleration / angular velocity sensors 11 of the actual robot 10 in a physical simulation, the actual robot 10 can perform appropriate actions (actions expected of the actual robot 10) without retraining using observed values ​​actually observed by the actual robot 10. However, it is possible to retrain the learning model 12 using observed values ​​actually observed by the actual robot 10.

[0034] Here, the performance of controlling the behavior of a real robot using a trained model developed through sim-to-real simulation, and consequently the accuracy of the real robot's behavior (how closely the real robot's behavior matches the expected behavior), depends on the accuracy with which the physical model of the real robot is reproduced in the physical simulation.

[0035] Improving the accuracy of a real robot's actions requires, in particular, accurate modeling of the robot's drive components, such as motors and gears, in physical simulations. If the modeling of the drive components is highly accurate, a learning model, such as a neural network, can accurately estimate the state of the drive components, including the forces acting on them, from the time history (time series) of sensor information detected by sensors that sense the drive components. Furthermore, by combining this with information detected by the IMU in the robot's torso, it can acquire the ability to estimate the contact situation between the robot and the real-world environment. This capability of the neural network makes it possible, for example, to realize a real robot capable of walking on various uneven terrains.

[0036] If a real robot has force sensors (force / torque sensors) that detect (linear) force and torque as sensors for sensing the drive unit, or if the drive unit uses planetary gears with a reduced speed ratio and a brushless DC motor, the drive unit can be modeled with high accuracy.

[0037] In other words, if a real robot has force sensors, the drive unit can be modeled with high accuracy by training a neural network to estimate the torque generated by the drive unit using the time series of torque detected by the force sensors (torque sensors). When a planetary gear with a reduction ratio and a brushless DC motor are used in the drive unit, the torque generated by the drive unit and the current value supplied to the drive unit have a nearly linear relationship, so the drive unit can be modeled with high accuracy.

[0038] However, force sensors are more expensive than IMUs. Furthermore, the appropriate drive unit for a real robot varies depending on the robot's size, weight, application, and manufacturing cost. Only a limited number of real robots can use planetary gears with a low gear ratio or high-power brushless DC motors that can be used even with a low gear ratio as their drive unit.

[0039] When sim-to-real learning is applied to a real robot whose drive mechanism cannot be modeled with high precision, the effects of friction and backlash in the drive mechanism are not reflected in the learning model. As a result, the behavior of the real robot estimated in such a learning model will differ significantly from the real robot's behavior in the real world, leading to a decrease in the performance of the robot's behavior control and, consequently, the accuracy of the robot's actions. In other words, if there is a large difference in the behavior of the drive mechanism between the physical simulation and the real world, there will also be a large difference in the robot's behavior between the physical simulation and the real world.

[0040] For example, if the accuracy of the drive unit model is low, the learning model will incorrectly estimate the state of the drive unit (such as the force acting on it) from sensor information detected by sensors that sense the drive unit. If the state of the drive unit is incorrectly estimated, the learning model will also incorrectly predict the state of contact between the actual robot and the (real-world) environment, increasing the likelihood that the actual robot, whose behavior is controlled using the learning model, will perform inappropriate actions that are far removed from appropriate actions (expected actions).

[0041] Therefore, if the accuracy of the drive unit modeling is low, reinforcement learning of the learning model can be performed using, for example, environmental randomization, assuming that the accuracy of the drive unit modeling is low. However, in this case, the actual robot whose behavior is controlled using the learning model will not perform inappropriate actions due to erroneous predictions by the learning model, but it will only be able to perform conservative actions that do not depend on the state of contact with the environment. As a result, for example, a walking robot may be able to walk on flat ground but will not be able to maintain balance well on uneven ground.

[0042] Accordingly, for example, there is a method in which a learning model is trained through sim-to-real in two stages. In this method, a learning model that has acquired conservative (calm) behaviors through sim-to-real is used to cause a real robot to perform actions in the real world, thereby collecting observation values obtained by the real robot and other necessary real-world data. After correcting the parameters of the physical simulation to match the real-world data, the learning model is trained again via sim-to-real, whereby the learning model acquires more aggressive behaviors. A conservative behavior is, for example, a low-speed behavior or a simple behavior (e.g., a behavior that causes (almost) no change in the center of gravity, the contact state with the environment, or the like). An aggressive behavior is a high-speed (agile) behavior or a complex behavior.

[0043] However, in the method of training a learning model through sim-to-real in two stages, behaviors of the real robot for which high accuracy is guaranteed are limited to conservative behaviors for which data is collected by causing the real robot to act in the real world.

[0044] That is, in the method of training a learning model through sim-to-real in two stages, even if the final objective is to cause the real robot to perform agile behaviors, it is first necessary to collect real-world data for conservative behaviors and correct the parameters of the physical simulation to match the real-world data. For the physical simulation after parameter correction, although high-accuracy simulation can be performed for conservative behaviors, high-accuracy simulation is not necessarily achievable for agile behaviors. When low-accuracy simulation is performed for agile behaviors, the behavior of the real robot estimated by the learning model will differ greatly from the actual behavior of the real robot in the real world, leading to a decrease in the performance of motion control of the real robot, and consequently, a decrease in the accuracy of the real robot's actions.

[0045] Furthermore, in the method of performing training of a learning model by sim-to-real in two stages, the computation time required for reinforcement learning mainly using large-scale computing resources such as a GPU (graphics processing unit) increases significantly. For example, in the method of performing training of a learning model by sim-to-real in two stages, it is simply necessary to perform twice as much reinforcement learning as compared with normal sim-to-real.

[0046] In addition, in the method of performing training of a learning model by sim-to-real in two stages, it is necessary to pay attention to the collection of real-world data, and if there is a problem with the collection of real-world data, the behavior of the real robot estimated by the learning model will differ more greatly from the actual behavior of the real robot in the real world.

[0047] In the present technology, a plurality of acceleration / angular velocity sensors 11 are attached to different parts of the real robot 10. Then, for the real robot 10, behavior control is performed by the learning model 12 trained by sim-to-real, using observation values including acceleration and angular velocity detected by the plurality of acceleration / angular velocity sensors 11 (virtual sensors 21 that reproduce the sensors in physical simulation).

[0048] The real robot 10 is different from, for example, a manipulator not provided with an IMU, for example, a manipulator fixed to the environment where behavior control is performed by a learning model trained by sim-to-real using observation values that do not include acceleration and angular velocity detected by an IMU.

[0049] Furthermore, the real robot 10 is also different from an autonomously moving robot such as a walking robot where only one IMU is provided only on the body, for example, and behavior control is performed by a learning model trained by sim-to-real using observation values including acceleration and angular velocity detected by the one IMU (further including the attitude of the IMU estimated from the acceleration and angular velocity).

[0050] Multiple acceleration / angular velocity sensors 11 are attached to different parts of the actual robot 10. Therefore, the acceleration and angular velocity of multiple parts to which the acceleration / angular velocity sensors 11 are attached can be directly detected (measured) and used for controlling the actions of the actual robot 10.

[0051] In the actual robot 10, any number of acceleration / angular velocity sensors 11 from the multiple acceleration / angular velocity sensors 11 can be attached to parts that can come into contact with the environment, for example, if the actual robot 10 is a (bipedal) walking robot, to the left and right toes that come into contact with the ground or floor as the environment. This allows for the rapid and accurate acquisition of information regarding contact with the environment, which can then be used to control the actions of the actual robot 10.

[0052] In a real robot whose behavior is controlled by a learning model trained using sim-to-real simulation, information regarding contact with the environment can be estimated, for example, by installing only one IMU (Inertial Measurement Unit) in the torso, and estimating the acceleration and angular velocity detected by that single IMU, as well as the attitude of the IMU estimated from that acceleration and angular velocity, and the state of the drive unit estimated from sensor information detected by a sensor that senses the drive unit. However, in this case, accurate modeling of the drive unit in physical simulation is necessary to accurately estimate information regarding contact with the environment.

[0053] On the other hand, in the case of the actual robot 10, if the part to which the acceleration / angular velocity sensor 11 is attached is a part that can come into contact with the environment, information regarding the contact between that part and the environment (for example, whether the feet are touching the ground) can be obtained from the acceleration and angular velocity detected by the acceleration / angular velocity sensor 11 attached to that part. In this case, accurate modeling of the acceleration / angular velocity sensor 11, such as an IMU, is required in the learning of the sim-to-real learning model 12, but accurate modeling of an IMU is extremely easy compared to accurate modeling of the drive unit. Therefore, the accuracy of the actions of the actual robot 10 can be easily improved.

[0054] Furthermore, one method for controlling the behavior of a walking robot is not reinforcement learning, but rather using the ZMP (zero moment point) equation, which describes the mechanical conditions used by the walking robot to maintain balance. However, while the ZMP equation method uses acceleration data from multiple points on the walking robot, it does not use angular velocity, and requires sensors at the toes to detect contact with the environment. Moreover, the ZMP equation method has many constraints, limiting its applicability and lacking versatility.

[0055] <Learning device>

[0056] Figure 2 is a block diagram showing an example configuration of a learning device that performs training on the learning model 12.

[0057] In Figure 2, the learning device 40 has a simulation unit 41 and a learning unit 42, and performs learning of the learning model 12 using sim-to-real.

[0058] The simulation unit 41 performs a physical simulation that reproduces the multiple acceleration / angular velocity sensors 11 of the actual robot 10. When the actual robot 10 (or the corresponding virtual robot 20) performs various actions, the simulation unit 41 generates observed values ​​(time series) observed by the actual robot 10, including the acceleration and angular velocity detected by the multiple acceleration / angular velocity sensors 11 (or virtual sensors 21 that model them), and supplies these values ​​to the learning unit 42.

[0059] The learning unit 42 uses the observed values ​​from the simulation unit 41 as training data for the learning model 12, and performs machine learning such as deep learning on the learning model 12, which takes the observed values ​​as input and outputs behavioral information representing the actions of the actual robot 10. As described above, the machine learning method for the learning model 12 can be arbitrarily selected from reinforcement learning, supervised learning, etc. Furthermore, the machine learning of the learning model 12 can be performed by combining multiple machine learning methods.

[0060] Figure 3 is a flowchart illustrating an example of the learning process for the learning model 12 of the learning device 40.

[0061] In step S11, the simulation unit 41 performs a physical simulation that reproduces various sensors, including multiple acceleration / angular velocity sensors 11 of the actual robot 10. This simulation generates observed values ​​observed by the actual robot 10, including the acceleration and angular velocity detected by the multiple acceleration / angular velocity sensors 11 when the actual robot 10 performs various actions. The simulation unit 41 supplies the observed values ​​to the learning unit 42, and the process proceeds from step S11 to step S12.

[0062] In step S12, the learning unit 42 uses the observed values ​​from the simulation unit 41 as training data to perform machine learning on the learning model 12, which takes the observed values ​​as input and outputs behavioral information, and the learning process is completed.

[0063] <Robot System>

[0064] Figure 4 is a block diagram showing an example configuration of a first embodiment of a robot system to which this technology is applied.

[0065] In Figure 4, the robot system 50 includes a physical robot 10 and a robot control device 60. Note that in the robot system 50, the robot control device 60 can be included within the physical robot 10.

[0066] The actual robot 10 is, for example, a multi-joint robot having multiple joints, and has M (>1) IMUs 51-1 to 51-M as multiple acceleration / angular velocity sensors 11, joint encoders 52, other sensors 53, an external input interface 54, and motors 55.

[0067] The M IMUs 51-1 to 51-M are attached to different parts of the actual robot 10. Each IMU 51-m (m = 1, 2, ..., M) detects the acceleration and angular velocity of the part (location) to which it is attached, and supplies these to the robot control device 60 as (part of) the observed values ​​observed by the actual robot 10.

[0068] The joint encoder 52 is attached to each joint, for example, and detects the position of the joint (joint position) and the speed of joint movement (joint velocity), and supplies these to the robot control device 60 as (part of) the observed values ​​observed by the actual robot 10.

[0069] Other sensors 53 are any sensors other than the IMU. Examples of other sensors 53 include force sensors, tactile sensors, image sensors, depth sensors, etc. The other sensors 53 detect a predetermined physical quantity and supply sensor information (other sensor measurement value) representing that predetermined physical quantity to the robot control device 60 as (part of) the observed value observed by the actual robot 10.

[0070] The external input interface 54 receives external commands, such as commands from a user, and supplies them to the robot control device 60 as (part of) the observed values ​​observed by the actual robot 10. For example, if the actual robot 10 is a mobile robot, the external commands may represent the target speed of movement, the target position after movement, etc.

[0071] The motor 55 is, for example, attached to each joint and is a drive unit that drives each joint. As a command to control the motor 55, for example, a current value, a voltage value, or a torque value to be generated by the motor 55 can be used. In this embodiment, the current value to be supplied to the motor 55 is used as the command to control the motor 55.

[0072] In addition to the information output by the IMUs 51-1 to 51-M, joint encoders 52, other sensors 53, and the external input interface 54, the observed values ​​of the actual robot 10 also include information obtained from outside and inside the robot system 50. For example, the most recent (previous) action information obtained by the robot control device 60 can be used as an observed value of the actual robot 10.

[0073] Furthermore, the observed values ​​observed by the actual robot 10 include not only the information itself output by the IMUs 51-1 to 51-M, joint encoders 52, other sensors 53, and the external input interface 54, but also other information obtained by applying an arbitrary algorithm to that information. For example, feature quantities extracted from arbitrary information output by the IMUs 51-1 to 51-M, joint encoders 52, other sensors 53, and the external input interface 54 using a general neural network such as a CNN (convolutional neural network) or an encoder such as a transformer can be used as observed values ​​observed by the actual robot 10.

[0074] The robot control device 60 has M attitude angle estimation units 61-1 to 61-M, the same number as the IMUs 51-1 to 51-M, a history storage unit 62, a joint position command calculation unit 63, and a motor current command calculation unit 64. The robot control device 60 generates behavioral information by providing observed values ​​observed by the actual robot 10 to the learning model 12 learned by sim-to-real in the learning device 40, and functions as a robot control unit that controls the behavior of the actual robot 10 according to that behavioral information.

[0075] The attitude angle estimation unit 61-m is supplied with the acceleration and angular velocity of the part to which the IMU 51-m is attached from the IMU 51-m. The attitude angle estimation unit 61-m uses the acceleration and angular velocity from the IMU 51-m to estimate the attitude angle of the IMU 51-m (for example, the angles of pitch, roll, and yaw) for which the acceleration and angular velocity have been detected, and supplies it to the history storage unit 62 as (part of) the observed values ​​observed by the actual robot 10.

[0076] The history storage unit 62 sequentially stores the observed values ​​supplied to the robot control device 60, including the attitude angle of the IMU 51-m as an observed value supplied from the attitude angle estimation unit 61-m. The history storage unit 62 can store only the observed value t at the most recent time t, or it can store the observed values ​​t, t-Δt, ..., t-NΔt for N+1 time points up to time t-N, which is N (>0) time points back from the most recent time point t. Δt represents the generation (update) interval of the action information.

[0077] In Figure 5, the observed value t at time 1 includes acceleration and angular velocity output by each of the M IMUs 51-1 to 51-M, attitude angles output by each of the M attitude angle estimation units 61-1 to 61-M, joint position and joint velocity output by the joint encoder 51, sensor information (other sensor measurements) output by other sensors 53, external commands output by the external input interface 54, and recent action information.

[0078] The joint position command calculation unit 63 has a learning model 12 and a post-action processing unit 71. Using the observed values ​​stored in the history storage unit 62, it generates (calculates) joint position commands representing the joint positions of each joint of the actual robot 10 in order to cause the actual robot 10 to perform a predetermined action, and supplies these commands to the motor current command calculation unit 64.

[0079] The learning model 12 in the joint position command calculation unit 63 is a learned model that has been trained using sim-to-real learning in the learning device 40 shown in Figure 2. The joint position command calculation unit 63 generates action information representing the actions of the actual robot 10 by providing the observed values ​​stored in the history storage unit 62 to the learning model 12, and supplies it to the post-action processing unit 71. The action information is, for example, a vector whose elements are values ​​corresponding to the joint positions of each joint of the actual robot 10.

[0080] Furthermore, in generating behavioral information, if, for example, the (translational) velocity of the actual robot 10 can be estimated from the sensor information output by the other sensor 53, that velocity can be used. In addition, in generating behavioral information, if, for example, a map of the area around the actual robot 10 can be generated by the robot control device 60 using SLAM (simultaneous localization and mapping) or the like, or can be obtained from an external source, that map can be used.

[0081] The post-action processing unit 71 converts the action information from the learning model 12 into joint position commands and supplies them to the motor current command calculation unit 64. For example, the post-action processing unit 71 converts the action information into joint position commands by multiplying it by a constant value and / or adding it. The post-action processing unit 71 can also apply an exponential moving average filter or a moving average filter to the action information or joint position commands to generate joint position commands that suppress abrupt changes in joint position.

[0082] The motor current command calculation unit 64 has a calculation unit 81 and generates (calculates) a motor current command that represents the current value to be supplied to the motor 55 of each joint in order to move the joint to the joint position represented by the joint position command from the joint position command calculation unit 63 (or its post-action processing unit 71), and supplies it to the motor 55.

[0083] The calculation unit 81 receives joint position commands from the joint position command calculation unit 63, as well as the joint position and joint velocity of each joint from the joint encoder 52.

[0084] The calculation unit 81 sets a target joint position for each joint using the joint position command from the joint position command calculation unit 63, the current joint position and joint velocity from the joint encoder 52, and generates a motor current command that represents the current value to be supplied to the motor 55 in order to move the current joint position to the target joint position. The calculation unit 81 supplies the corresponding motor current command to the motor 55 of each joint. In most cases, the motor current command is generated at an interval shorter than the interval at which the action information is generated.

[0085] Figure 5 is a flowchart illustrating an example of the action control process used by the robot control device 60 to control the actions of the actual robot 10.

[0086] In step S21, the joint position command calculation unit 63 of the robot control device 60 generates action information representing the actions of the actual robot 10 by providing the observed values ​​observed by the actual robot 10, which are stored in the history storage unit 62 and include acceleration and angular velocity detected by each of the M IMUs 51-1 to 51-M, to the learning model 12, and the process proceeds to step 22.

[0087] In step S22, the robot control device 60 controls the actual robot 10 so that it performs the action indicated by the action information. Specifically, in the robot control device 60, the post-action processing unit 71 of the joint position command calculation unit 63 converts the action information into joint position commands and supplies them to the motor current command calculation unit 64. In the motor current command calculation unit 64, the calculation unit 81 sets a target joint position for each joint using the joint position indicated by the joint position command from the joint position command calculation unit 63 and the current joint position and joint velocity from the joint encoder 52, and generates a motor current command that represents the current value to be supplied to the motor 55 in order to move the current joint position to the target joint position. The calculation unit 81 supplies the corresponding motor current command to the motor 55 of each joint. The motor 55 of each joint is driven according to the motor current command from the calculation unit 81, thereby moving the corresponding joint, and the actual robot 10 performs the action indicated by the action information.

[0088] The generation of action information in step S21 and the control of the actual robot 10 according to the action information in step S22 are performed repeatedly.

[0089] Figure 6 is a block diagram showing an example configuration of a second embodiment of a robot system to which this technology is applied.

[0090] In the figures, parts corresponding to those in Figure 4 are denoted by the same reference numerals, and their explanations will be omitted as appropriate below.

[0091] In Figure 6, the robot system 110 includes a real robot 10 and a robot control device 60. Therefore, the robot system 110 is configured in the same way as the robot system 50 in Figure 4.

[0092] However, in the robot system 110, the mounting locations for some of the M IMUs 51-1 to 51-M, which serve as multiple acceleration / angular velocity sensors 11, are specified to particular parts.

[0093] In other words, in the robot system 110, for example, the actual robot 10 is a bipedal walking robot, and the IMU 51-1 is attached to its torso. Furthermore, IMUs 51-2 and 51-3 are attached to parts of the walking robot, which is the actual robot 10, that can come into contact with the environment, such as the tips of the left and right feet that come into contact with the ground.

[0094] In various real-world environments, the walking robot 10 needs to perform agile balance actions depending on the state of contact with the environment (how it makes contact with the environment), for example, the state of contact between its left and right feet and the ground.

[0095] In walking robots, if a drive unit that can be modeled with high accuracy using physical simulation is used, reinforcement learning of the learning model using sim-to-real allows the learning model to accurately estimate the state of the drive unit and the state of contact between the walking robot and the environment, enabling it to control agile balancing actions. However, if a walking robot does not use (or cannot use) a drive unit that can be modeled with high accuracy using physical simulation, the walking robot will only be able to perform conservative actions, or the learning model will make errors in estimating the state of the drive unit and the state of contact between the walking robot and the environment, resulting in reduced accuracy of actions and making it difficult to perform agile balancing actions.

[0096] Meanwhile, in the robot system 110, IMUs 51-2 and 51-3 are attached to the left and right toes of the walking robot, which is the actual robot 10. The learning model 12 is trained using sim-to-real learning, with the observed values, including the acceleration and angular velocity of the left and right toes, output by the IMUs 51-2 and 51-3. As a result, the learning model 12 can accurately (and quickly) estimate the state of contact between the left and right toes and the environment, and the walking robot, which is the actual robot 10, can be controlled by utilizing the state of contact between the left and right toes and the environment, which is thus accurately estimated.

[0097] In training the sim-to-real learning model 12 using observed values ​​including the acceleration and angular velocity of the left and right toes output by IMUs 51-2 and 51-3, accurate modeling of the IMUs is easy, thus easily improving the accuracy of the actions of the actual robot 10.

[0098] In other words, in the learning of the sim-to-real learning model 12 using observed values ​​including the acceleration and angular velocity of the left and right toes output by IMUs 51-2 and 51-3, it is not necessary to model the motor 55 as the drive unit. Furthermore, without depending on an accurate model of the motor 55, IMUs 51-2 and 51-3 can quickly and accurately detect the acceleration and angular velocity as state variables resulting from the contact between the left and right feet, to which IMUs 51-2 and 51-3 are attached, and the environment. Moreover, according to the learning model 12, which has been learned using these observed values ​​including acceleration and angular velocity, the actual robot 10 can quickly and accurately perform the actions it should take in response to the contact between the left and right feet and the environment.

[0099] Figure 7 shows an overview of a third embodiment of a robot system to which this technology is applied.

[0100] In the figures, parts corresponding to those in Figure 4 are denoted by the same reference numerals, and their explanations will be omitted as appropriate below.

[0101] In Figure 7, the robot system 130 includes a real robot 10 and a robot control device 60. Therefore, the robot system 130 is configured in the same way as the robot system 50 in Figure 4.

[0102] However, in the robot system 130, the actual robot 10 is concretely represented as a manipulator (robot arm).

[0103] The actual robot 10 has a base 131, an arm (link) 132, a flexible part 133, and an end effector (end tool) 134.

[0104] The base 131 supports the arm 132. The arm 132 has one or more joints, for example, three joints, and one end is fixed to the base 131. The other end of the arm 132, i.e., the wrist portion of the manipulator, is provided with a flexible part 133, and an end effector 134, which is the end of the manipulator's hand, is attached to the flexible part 133. The flexible part 133 is any flexible mechanism, such as an elastic body like a spring or rubber, or a damper.

[0105] In the actual robot 10, some of the M IMUs 51-1 to 51-M, which serve as multiple acceleration / angular velocity sensors 11, such as IMU 51-1, are attached to the end effector 134, which is a part that can come into contact with the environment.

[0106] Figure 8 is a block diagram showing an example of the electrical configuration of the robot system 130.

[0107] In the figures, parts corresponding to those in Figure 4 are denoted by the same reference numerals, and their explanations will be omitted as appropriate below.

[0108] In Figure 8, the robot system 130 includes a real robot 10 and a robot control device 60, and is configured similarly to the robot system 50 in Figure 4.

[0109] However, in the robot system 130, one of the M IMUs 51-1 to 51-M, which serve as multiple acceleration / angular velocity sensors 11, for example, IMU 51-1, is mounted on the end effector 134.

[0110] In the actual robot 10 configured as described above, for example, the end effector 134 can grasp an object (workpiece), and by transferring, rotating, tilting, etc., the object can perform tasks such as tightening screws or fitting parts.

[0111] The end effector 134 and the object it is gripping may come into contact with the environment (other objects) during operation. Since the end effector 134 is attached to the flexible part 133, if the end effector 134 or the object it is gripping comes into contact with the environment, the flexible part 133 absorbs the impact of the contact with the environment. Therefore, the impact on the end effector 134, the object it is gripping, and the environment can be suppressed.

[0112] As described above, the actual robot 10, which has a flexible part 133 at the wrist of the manipulator, can perform dexterous manipulation tasks while suppressing the impact when it comes into contact with the environment.

[0113] However, it is difficult to measure the state variables (mutation and mutation rate) of the flexible part 133, and therefore it is also difficult to measure the state (position, orientation, and velocity) of the end effector 134, which is the end effector part of the manipulator, on the side of the flexible part 133 that is not the side of the arm 132.

[0114] For example, one method involves attaching a force sensor to the end effector 134, estimating the state variables (position, posture, and velocity) of the end effector 134 using the state variables (position and velocity) of some joints of the arm 132 and the force and torque (observed values) detected by the force sensor on the end effector 134, and then using these state variables of the end effector 134 to train a sim-to-real learning model 12.

[0115] However, this method requires force sensors. Force sensors are more expensive than IMUs, and their use may be difficult from a manufacturing cost perspective.

[0116] Therefore, an IMU 51-1 is attached to the end effector 134, and the acceleration and angular velocity detected by the IMU 51-1 are utilized, and the learning model 12 can be trained using sim-to-real learning with the observed values ​​including the acceleration and angular velocity.

[0117] As described above, the IMU 51-1 attached to the end effector 134 directly detects (measures) the acceleration and angular velocity of the end effector 134, and by using these observed values, including the acceleration and angular velocity, to train the sim-to-real learning model 12, it is possible to obtain a learning model 12 that enables the actual robot 10 to perform highly accurate actions by controlling its behavior while estimating the state of the end effector 134 on the end-effector side of the flexible part 133. Therefore, the accuracy of the actions of the actual robot 10, which is a manipulator, can be improved using an inexpensive IMU without using an expensive force sensor.

[0118] Furthermore, among the M IMUs 51-1 to 51-M that serve as multiple acceleration / angular velocity sensors 11, IMUs 51-2 to 51-M, excluding IMU 51-1, can be attached, for example, to each part of the arm 132 that is separated by a joint.

[0119] Furthermore, the flexible portion 133 can be provided at the wrist portion between the arm 132 and the end effector 134, or at any portion between the arm 132 and the end effector 134. The IMU 51-1 is attached to the tip side of the end effector 134 from the flexible portion 133.

[0120] In addition, the flexible portion 133 can be provided in multiple locations.

[0121] <Description of a computer using this technology>

[0122] The series of processes described above for the learning device 40 and the robot control device 60 can be executed by hardware or by software. When the series of processes are executed by software, the programs that make up the software are installed on a computer. Here, the computer includes computers built into dedicated hardware, as well as general-purpose personal computers, for example, that can perform various functions by installing various programs.

[0123] Figure 9 is a block diagram showing an example of the hardware configuration of a computer that executes the series of processes described above using a program.

[0124] In a computer, the processing circuit 901, ROM (Read Only Memory) 902, and RAM (Random Access Memory) 903 are interconnected by a bus 904.

[0125] An input / output interface 905 is further connected to the bus 904. An input unit 906, an output unit 907, a storage unit 908, a communication unit 909, and a drive 910 are connected to the input / output interface 905.

[0126] The input unit 906 may include physical or virtual operating means that the user operates to input information, such as a keyboard, mouse, or touch panel, as well as means that the user inputs information through voice, eye gaze, etc. Furthermore, the input unit 906 may include sensors for inputting various physical quantities to the computer. For example, the input unit 906 may include sensors that acquire physical quantities such as light (including infrared light other than visible light) or sound, such as a camera or microphone. Also, for example, the input unit 906 may include sensors that acquire other physical quantities such as temperature, moisture content, acceleration, distance, etc. The output unit 907 may include means that present information to the user by stimulating the user's perception, such as a display, speaker, or haptic device. The storage unit 908 is composed of a hard disk, non-volatile or volatile memory, etc., and stores various types of information (including programs). The communication unit 909 is a network interface, etc., and performs wired or wireless communication with the outside. The drive 910 drives removable media 911 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory.

[0127] The processing circuit 901 includes a processor that executes programs such as a CPU (Central Processing Unit) and a DSP (Digital Signal Processor). The processing circuit 901 (its processor) performs the above-described series of processes by loading the program stored in the storage unit 908 into the RAM 903 via the input / output interface 905 and the bus 904 and executing it. The processing circuit 901 can output the processing results of the series of processes from the output unit 907 via the bus 904 and the input / output interface 905 as needed. The processing circuit 901 can also store the processing results in the storage unit 908 or transmit them from the communication unit 909.

[0128] The program executed by the computer (processing circuit 901) can be provided by recording it on a removable medium 911, such as a package medium. The program can also be provided via wired or wireless transmission media, such as a local area network, the internet, or digital satellite broadcasting.

[0129] In a computer, a program can be installed in the storage unit 908 via the input / output interface 905 by inserting a removable media 911 into the drive 910. Alternatively, a program can be received by the communication unit 909 from another device, such as a server, via a wired or wireless transmission medium, and installed in the storage unit 908. Furthermore, programs can be pre-installed in the ROM 902 or the storage unit 908.

[0130] The programs executed by the computer may be programs that are processed chronologically in the order described herein, or they may be programs that are processed in parallel or at necessary times, such as when a call is made.

[0131] The processes that a computer performs according to a program do not necessarily have to follow the order described in the flowchart. In other words, the processes that a computer performs according to a program include processes that are executed in parallel or individually (e.g., parallel processing and object-based processing).

[0132] The program may be processed by a single computer (processor), or it may be processed in a distributed manner by multiple computers. Furthermore, the program may be transferred to a remote computer and executed there.

[0133] When the computer executes a program, the above-described series of processes are performed, the memory unit 908 functions as a history memory unit 62. The processing circuit 901 (its processor) executes a program, and functions as a simulation unit 41, a learning unit 42, attitude angle estimation units 61-1 to 61-M, a joint position command calculation unit 63, and a motor current command calculation unit 64.

[0134] In this specification, a system means one component or a collection of multiple components (devices, modules (parts), etc.). Therefore, one or more components of a computer, for example, only the processor, or a combination of the processor and memory (for example, only the processing circuit 901, or a combination of the processing circuit 901 to the bus 904, etc.), constitute a system. Regarding a collection of multiple components, it is not necessary whether all components reside in the same enclosure. Therefore, multiple devices housed in separate enclosures and connected via a network, or a single device containing multiple modules within a single enclosure, are all systems. Furthermore, for example, the entire computer, or a combination of a computer and other devices such as a server (not shown), also constitute a system.

[0135] It should be noted that the embodiments of this technology are not limited to those described above, and various modifications are possible without departing from the spirit of this technology.

[0136] For example, this technology can be configured as cloud computing, where a single function is shared and processed collaboratively by multiple devices via a network.

[0137] Furthermore, each step described in the flowchart above can be performed by a single device, or it can be divided and performed by multiple devices.

[0138] Furthermore, if a single step includes multiple processes, those processes can be executed by a single device or shared among multiple devices.

[0139] Furthermore, the effects described herein are merely illustrative and not limiting, and other effects may also occur.

[0140] Furthermore, this technology can take the following configuration.

[0141] <1> A robot that acts according to behavioral information generated by providing the observed values, including the acceleration and angular velocity detected by the multiple sensors, to a learning model that takes the observed values ​​as input and outputs behavioral information representing the robot's actions, which is trained using observed values, including the acceleration and angular velocity obtained by a physical simulation that reproduces the multiple sensors of the robot equipped with the multiple sensors, as training data. <2> The robot according to <1>, further comprising a robot control unit that generates the behavioral information by providing the observed values ​​to the learning model and controls the robot's actions according to the behavioral information. <3> The robot according to <1> or <2>, wherein one or more of the multiple sensors are attached to a part that can come into contact with the environment. <4> The robot according to <3>, which is a walking robot, wherein the sensor is attached to the toe. <5> The robot according to <3>, which is a manipulator, wherein a part of the manipulator from the arm to the end effector is a flexible part, and the sensor is attached to the tip side of the end effector from the flexible part. <6> The robot according to <5>, wherein the space between the arm and the end effector is the flexible part, and the sensor is attached to the end effector. <7> The robot according to any one of <1> to <6>, wherein each of the plurality of sensors includes an acceleration sensor for detecting acceleration and an angular velocity sensor separate from the acceleration sensor for detecting angular velocity. <8> A method for controlling the robot, comprising: generating behavior information by providing the observed values, including acceleration and angular velocity detected by the plurality of sensors, to a learning model that outputs behavior information representing the robot's actions using the observed values ​​as learning data, and which has been trained using the observed values ​​obtained by a physical simulation that reproduces the plurality of sensors of a robot each equipped with a plurality of sensors for detecting acceleration and angular velocity; and controlling the robot's actions according to the behavior information.<9> The control method according to <8>, wherein the robot is a robot in which one or more of the plurality of sensors are attached to a part that can come into contact with the environment. <10> The control method according to <9>, wherein the robot is a walking robot in which the sensor is attached to the toes. <11> The control method according to <9>, wherein the robot is a manipulator, and a part of the manipulator from the arm to the end effector is a flexible part, and the sensor is attached to the tip side of the end effector from the flexible part. <12> The control method according to <11>, wherein the space between the arm and the end effector is the flexible part, and the sensor is attached to the end effector. <13> The control method according to any one of <8> to <12>, wherein each of the plurality of sensors includes an acceleration sensor for detecting acceleration and an angular velocity sensor separate from the acceleration sensor for detecting angular velocity. <14> A learning method comprising: generating observed values ​​observed by a robot, including acceleration and angular velocity detected by the multiple sensors, by performing a physical simulation that reproduces the multiple sensors of a robot, each equipped with multiple sensors that detect acceleration and angular velocity; and training a learning model that uses the observed values ​​as training data and outputs behavioral information representing the robot's actions with the observed values ​​as input. <15> The learning method according to <14>, wherein the robot is a robot in which one or more of the multiple sensors are attached to a part that can come into contact with the environment. <16> The learning method according to <15>, wherein the robot is a walking robot in which the sensor is attached to the toes. <17> The learning method according to <15>, wherein the robot is a manipulator, and a part of the manipulator from the arm to the end effector is a flexible part, and the sensor is attached to the tip side of the end effector from the flexible part. <18> The learning method according to <17>, wherein the space between the arm and the end effector is the flexible part, and the sensor is attached to the end effector.<19> The learning method according to any one of <14> to <18>, wherein each of the plurality of sensors includes an acceleration sensor for detecting acceleration and an angular velocity sensor separate from the acceleration sensor for detecting angular velocity.

[0142] 10 Actual robot, 11 Acceleration / angular velocity sensor, 12 Learning model, 20 Virtual robot, 21 Virtual sensor, 40 Learning device, 41 Simulation unit, 42 Learning unit, 50 Robot system, 51-1 to 51-M IMU, 52 Joint encoder, 53 Other sensors, 54 External input interface, 55 Motor, 60 Robot control device, 61-1 to 61-M Posture angle estimation unit, 62 History storage unit, 63 Joint position command calculation unit, 64 Motor current command calculation unit, 71 Post-action processing unit, 81 Calculation unit, 110, 130 Robot system, 131 Base, 132 Arm, 133 Flexible part, 134 End effector, 901 Processing circuit, 902 ROM, 903 RAM, 904 Bus, 905 Input / Output Interface, 906 Input Section, 907 Output Section, 908 Storage Section, 909 Communication Section, 910 Drive, 911 Removable Media

Claims

1. A robot that acts according to behavioral information generated by providing the observed values, including acceleration and angular velocity detected by the multiple sensors, to a learning model that outputs behavioral information representing the robot's actions, which is trained using observed values, including acceleration and angular velocity obtained by a physical simulation that reproduces the multiple sensors of the robot equipped with the multiple sensors, as training data, and which takes the observed values ​​as input.

2. The robot according to claim 1, comprising a robot control unit that generates the behavior information by providing the observed values ​​to the learning model and controls the robot's behavior according to the behavior information.

3. The robot according to claim 1, wherein one or more of the plurality of sensors are attached to a part that can come into contact with the environment.

4. The walking robot according to claim 3, wherein the sensor is attached to the toe.

5. The robot according to claim 3, wherein a part of the manipulator from the arm to the end effector is a flexible part, and the sensor is attached to the tip side of the end effector from the flexible part.

6. The robot according to claim 5, wherein the space between the arm and the end effector is the flexible part, and the sensor is attached to the end effector.

7. The robot according to claim 1, wherein each of the plurality of sensors includes an acceleration sensor for detecting acceleration and an angular velocity sensor, separate from the acceleration sensor, for detecting angular velocity.

8. A method for controlling a robot, comprising: generating behavioral information by providing the observed values, including acceleration and angular velocity detected by the multiple sensors, to a learning model that outputs behavioral information representing the robot's actions, which has been trained using observed values ​​observed by the robot as training data, and which takes the observed values ​​as input and outputs behavioral information representing the actions of the robot, to the learning model; and controlling the actions of the robot according to the behavioral information.

9. The control method according to claim 8, wherein the robot is a robot in which one or more of the plurality of sensors are attached to a part that can come into contact with the environment.

10. The control method according to claim 9, wherein the robot is a walking robot in which the sensor is attached to the toes.

11. The control method according to claim 9, wherein the robot is a manipulator, a portion of the manipulator from the arm to the end effector is a flexible part, and the sensor is attached to the tip side of the end effector from the flexible part.

12. The control method according to claim 11, wherein the space between the arm and the end effector is the flexible portion, and the sensor is attached to the end effector.

13. The control method according to claim 8, wherein each of the plurality of sensors includes an acceleration sensor for detecting acceleration and an angular velocity sensor, separate from the acceleration sensor, for detecting angular velocity.

14. A learning method comprising: generating observed values ​​observed by a robot, including acceleration and angular velocity detected by a plurality of sensors, by performing a physical simulation that reproduces the plurality of sensors of a robot, each of which has a plurality of sensors that detect acceleration and angular velocity; and training a learning model that uses the observed values ​​as training data and outputs behavioral information representing the actions of the robot, taking the observed values ​​as input.

15. The learning method according to claim 14, wherein the robot is a robot in which one or more of the plurality of sensors are attached to a part that can come into contact with the environment.

16. The learning method according to claim 15, wherein the robot is a walking robot in which the sensor is attached to the toes.

17. The learning method according to claim 15, wherein the robot is a manipulator, a portion of the manipulator from the arm to the end effector is a flexible part, and the sensor is attached to the tip side of the end effector from the flexible part.

18. The learning method according to claim 17, wherein the space between the arm and the end effector is the flexible portion, and the sensor is attached to the end effector.

19. The learning method according to claim 14, wherein each of the plurality of sensors includes an acceleration sensor for detecting acceleration and an angular velocity sensor, separate from the acceleration sensor, for detecting angular velocity.