Trained model generation device, control device, trained model generation method, and trained model generation program
A combined simulation approach using physical and robot simulations with reinforcement learning efficiently generates a learned model for robots to apply forces, addressing the challenge of obtaining accurate reaction force data and enhancing task performance.
Patent Information
- Application Number
- PCT/JP2025/000574
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-19
- Filing Date
- 2025-01-09
- Publication Date
- 2025-07-24
AI Technical Summary
Existing methods struggle to efficiently generate a learned model for robots to perform tasks involving applying forces to objects, as obtaining accurate reaction force data is laborious and difficult, whether from real-world objects or virtual simulations.
A system that combines physical and robot simulations to generate a learned model by simulating reaction forces and robot operations, using reinforcement learning to adjust parameters based on real-world data, allowing efficient generation of a control policy for robots to apply forces effectively.
Enables the efficient generation and utilization of a learned model for robots to perform tasks like cutting or gripping, even with irreversibility, by accurately simulating and learning from both physical and robot simulations, thereby improving task execution.
Smart Images

Figure JP2025000574_24072025_PF_FP_ABST
Abstract
Description
Trained model generation device, control device, trained model generation method, and trained model generation program
[0001] The present disclosure relates to a trained model generation device, a control device, a trained model generation method, and a trained model generation program.
[0002] Conventionally, a robot system that stably acquires information about a workpiece according to the situation has been known (see, for example, Japanese Patent Application Laid-Open No. 2023-168240). This robot system performs a physical simulation in which a virtual workpiece is allowed to freely fall into a virtual container, virtually captures images of the virtual workpieces bulk-stacked in the virtual container using a virtual imaging device to acquire image data, and then acquires a trained model through machine learning using the image data as input data.
[0003] However, it is difficult to teach a robot to perform actions that apply force to an object. For example, consider a task in which a robot holds a knife and uses it to cut an object. In this case, when a robot cuts an object with the knife, it needs to learn things like the angle at which the knife should be placed on the object, how much force should be applied to the knife, and how fast the knife should be moved.
[0004] When a robot learns to move in a way that takes the above points into consideration, it needs to acquire data on the reaction force that occurs when a force is applied to an object. In this case, for example, the reaction force data must be used to train a trained model using a machine learning algorithm, and the robot must then use the trained model to control its own movement.
[0005] However, after applying a force to an object and obtaining reaction force data, the object may change and it may be difficult to obtain reaction force data from the object again. For example, as mentioned above, if the object is cut once, it is difficult to obtain reaction force data thereafter.
[0006] In such cases, obtaining reaction force data from the real world for training a robot is difficult and time-consuming, requiring the preparation of a huge number of objects. On the other hand, even if virtual reaction force data is obtained using a robot simulator, it is difficult to obtain accurate reaction force data. Therefore, in the above-described cases, there is a problem in that reaction force data when a force is applied to an object cannot be efficiently obtained, and a trained model to be used when a robot applies a force to an object cannot be efficiently generated.
[0007] In the above-mentioned JP 2023-168240 A, when a robot is made to learn the operation of grasping an object, a physical simulation is performed in which virtual workpieces are made to freely fall into a virtual container, and the state of bulk stacking of the virtual workpieces is virtually reproduced. For this reason, JP 2023-168240 A is a technology for virtually reproducing the state of bulk stacking of virtual workpieces, and does not make the robot learn using data on the reaction force when the robot applies force to an object.
[0008] The present disclosure has been made in consideration of the above points, and aims to efficiently generate a trained model to be used when a robot performs a task of applying force to an object.
[0009] In order to achieve the above-mentioned object, the trained model generation device according to the present disclosure is a trained model generation device that includes: a simulation unit that executes a physical simulation that simulates a reaction force when a force is applied to a virtual object, and a robot simulation that simulates the operation of a virtual robot that applies a force to the virtual object; and a generation unit that generates a trained model that outputs behavioral data of a real robot corresponding to the virtual robot when state data representing an environment in which the real robot operates is input, based on data obtained while the physical simulation and the robot simulation are being executed.
[0010] In addition, the trained model generation method disclosed herein is a trained model generation method in which a computer performs processing to execute a physical simulation that simulates a reaction force when a force is applied to a virtual object, and a robot simulation that simulates the operation of a virtual robot that applies a force to the virtual object, and generates a trained model that outputs behavioral data of a real robot corresponding to the virtual robot when state data representing the environment in which the real robot operates is input, based on data obtained while the physical simulation and the robot simulation are being executed.
[0011] In addition, the trained model generation program of the present disclosure is a trained model generation program for causing a computer to execute a process that executes a physical simulation that simulates a reaction force when a force is applied to a virtual object, and a robot simulation that simulates the operation of a virtual robot that applies a force to the virtual object, and generates a trained model that outputs behavioral data of a real robot corresponding to the virtual robot when state data representing the environment in which the real robot operates is input, based on data obtained while the physical simulation and the robot simulation are being executed.
[0012] The trained model generation device, trained model generation method, and trained model generation program disclosed herein can efficiently generate a trained model used when a robot executes a task of applying a force to an object. Furthermore, the control device disclosed herein can execute the task of applying a force to an object using the efficiently generated trained model.
[0013] 1 is a diagram for explaining a control system of this embodiment. FIG. 2 is a diagram for explaining an outline of this embodiment. FIG. 3 is a diagram for explaining an outline of this embodiment. FIG. 4 is a diagram for explaining FDCC. FIG. 5 is a diagram for explaining the relationship between reinforcement learning and a compliance controller in this embodiment. FIG. 6 is a block diagram showing a schematic configuration of a control system of this embodiment. FIG. 7 is a block diagram showing a hardware configuration of a control device according to this embodiment. FIG. 8 is a flowchart showing the flow of a trained model generation process in this embodiment. FIG. 9 is a flowchart showing the flow of a control process in this embodiment. FIG. 10 is a diagram showing the results of this embodiment. FIG. 11 is a diagram showing the results of this embodiment.
[0014] An example of an embodiment of the present disclosure will be described below with reference to the drawings. In this embodiment, a control system equipped with a control device according to the present disclosure will be described as an example. Note that the same reference numerals are used in the drawings to designate identical or equivalent components and parts. Furthermore, the dimensions and proportions of the drawings are exaggerated for the sake of explanation and may differ from the actual proportions.
[0015] FIG. 1 is a diagram illustrating the present embodiment. As shown in FIG. 1, a control system 10 of the present embodiment includes a real robot 12R and a control device 14. The real robot 12R includes an end effector 16R. The control system 10 efficiently generates a trained model used when the real robot 12R executes a task of applying a force to a real object BR. For example, the force application task targeted by the control system 10 of the present embodiment is an irreversible task. For example, as shown in FIG. 1, the real robot 12R of the control system 10 holds a real knife KR, and the real robot 12R executes the task of cutting the real object BR using the real knife KR by operating the end effector 16R in accordance with a control signal output from the control device 14. The task of cutting the real object BR is an irreversible task in that once the real object BR is cut, the same part of the real object BR cannot be cut again.
[0016] 2 is a diagram illustrating an overview of this embodiment. As shown in FIG. 2, in this embodiment, a trained model is generated to be used when a real robot 12R in the real world Re executes a task of cutting a real object BR using a real knife KR. In this case, the control system 10 of this embodiment executes two types of simulations Sim.
[0017] 2 is a virtual object BS that is a real object BR that is virtually realized on a computer. Similarly, a virtual robot 12S is a virtual robot 12R that is virtually realized on a computer.
[0018] The control system 10 of this embodiment executes a physical simulation CutSim that simulates the reaction force from the virtual object BS when a force is applied to the virtual object BS, and a robot simulation RoboSim that simulates the operation of the virtual robot 12S that processes the virtual object BS.
[0019] Currently widely used robot simulations are simulations that simulate the internal state of a virtual robot 12S. Such robot simulations have difficulty simulating the external state of the virtual robot 12S. Furthermore, currently widely used physical simulations have difficulty executing simulations that simulate the internal state of the virtual robot 12S.
[0020] Therefore, the control system 10 of this embodiment generates a trained model to be used when executing a task of applying a force to an object, based on data obtained by executing both the physical simulation CutSim and the robot simulation RoboSim. The data obtained while the physical simulation CutSim is being executed is virtual reaction force data that represents the reaction force when the virtual robot 12S applies a force to the virtual object BS. By using this data, it is possible to efficiently generate a trained model even when the task of applying a force to an object is irreversible.
[0021] A specific description will be given below. In the following, when there is no need to distinguish between the real world and the virtual world, the terms "robot," "object," and "knife" will be used simply.
[0022] <Framework of this embodiment> (A. System overview) Fig. 3 is a diagram for explaining the overview of the control system 10 of this embodiment (denoted as "Robot Operating System (ROS)" in Fig. 3). Fig. 3 shows a robot simulation RoboSim and a physical simulation CutSim. The robot simulation RoboSim stores data representing the position x and speed x of a virtual knife KS held by the virtual robot 12S. ・ The physical simulation CutSim sequentially outputs data representing the contact force F while the object corresponding to the virtual knife KS is cutting the virtual object BS. ext are sequentially output to the robot simulation RoboSim.
[0023] 3, the control system 10 trains a machine learning model based on data obtained while the robot simulation RoboSim and the physics simulation CutSim are being executed, thereby efficiently generating a trained model to be used when executing a task of applying force to an object. Specifically, as shown in FIG. 3, the control system 10 receives data representing the position x of the virtual knife KS output from the robot simulation RoboSim, data representing the velocity x of the virtual knife KS, and ・ and data representing the jerk x of the virtual knife KS ・・・ and the contact force F between the virtual object BS and the virtual knife KS output from the physical simulation CutSim. ext Based on the data representing the control policy, a learned model is generated using a known reinforcement learning algorithm ("RL Control Policy" in FIG. 3).
[0024] From the RL Control Policy shown in FIG. 3, the desired posture x dand the external force F generated on the end effector 16R of the real robot 12R. d In this embodiment, a Forward Dynamics Compliance Controller (FDCC) is used as a known compliance controller (denoted as "Compliance Controller" in FIG. 3). As shown in FIG. 3, the FDCC calculates the position q of the joint of the virtual robot 12S or the real robot 12R. c The FDCC is disclosed in, for example, Reference 1 below.
[0025] Reference 1: S. Scherzinger, A. Roennau, and R. Dillmann, “Forward dynamics compliance control (FDCC): A new approach to cartesian compliance for robotic manipulators,” in IEEE / RSJ International Conference on Intelligent Robots and Systems, 2017, pp. 4568-4575.
[0026] In this embodiment, the process is comprised of two phases: reflection from reality to simulation, and reflection from simulation to reality. Specifically, based on various data obtained when the real robot 12R cuts the real object BR with the real knife KR, the parameters of the robot simulation RoboSim and the parameters of the physical simulation CutSim are adjusted by a known method.
[0027] More specifically, as shown in FIG. 3, for example, the contact force F between the real object BR and the real knife KR when the real object BR is cut is ext and data representing the position x of the actual knife KR, and the velocity x of the actual knife KR. ・ Data representing the acceleration x of the actual knife KR ・・ and the data representing the jerk x of the actual knife KR. ・・・The real robot 12R acquires data representing the robot's motion and the physical simulation. Using these various data, the parameters of the robot simulation RoboSim and the parameters of the physical simulation CutSim are adjusted.
[0028] Specifically, the contact force F when cutting the real object BR is ext are collected by the real robot 12R and reflected in the virtual object BS of the physical simulation CutSim. As a result, the hardness of the real object BR and the frictional force generated when cutting the real object BR with the real knife KR are reproduced in the physical simulation CutSim.
[0029] In addition, data representing the position x of the real knife KR when the real robot 12R is operated, and the speed x of the real knife KR ・ Data representing the acceleration x of the actual knife KR ・・ and the data representing the jerk x of the actual knife KR. ・・・ is collected and reflected in the virtual robot 12S of the robot simulation RoboSim. As a result, the operable conditions of the real robot 12R are reflected in the robot simulation RoboSim, and the real robot 12R is reproduced in the robot simulation RoboSim.
[0030] Thus, if the real object BR has been cut with the real knife KR in the real world, it is not possible to cut the real object BR with the real knife KR again. Therefore, the control system 10 of this embodiment collects various data obtained when the real object BR is cut with the real knife KR and reflects the collected data in the parameters of the physical simulation CutSim and the parameters of the robot simulation RoboSim. This makes it possible to reproduce the cutting of the real object BR that has already been cut in the real world in the physical simulation CutSim and the robot simulation RoboSim, thereby enabling various data to be repeatedly obtained. Furthermore, based on the multiple pieces of data obtained in the physical simulation CutSim and the robot simulation RoboSim, a trained model used when the real robot 12R executes a task of applying force to an object can be efficiently generated.
[0031] (B. Simulation Environment) [Type of Simulator] For example, known simulators can be used for the physical simulation CutSim and the robot simulation RoboSim. For example, the known DiSECt simulator can be used for the physical simulation CutSim. Furthermore, for example, the known Gazebo simulator can be used for the robot simulation RoboSim. The DiSECt simulator and the Gazebo simulator are disclosed in the following references 2 and 3.
[0032] Reference 2: E. Heiden, M. Macklin, Y. Narang, D. Fox, A. Garg, and F. Ramos, “Disect: a differentiable simulator for parameter inference and control in robotic cutting,” Autonomous Robots, pp. 1-30, 2023. Reference 3: N. Koenig and A. Howard, “Design and use paradigms for gazebo, an open-source multi-robot simulator,” in IEEE / RSJ International Conference on Intelligent Robots and Systems, vol. 3, 2004, pp. 2149-2154.
[0033] [Calibration] As described above, in this embodiment, the parameters of the physical simulation CutSim and the parameters of the robot simulation RoboSim are adjusted based on data obtained in the real world. For example, using a known algorithm, the parameters of the physical simulation CutSim and the parameters of the robot simulation RoboSim are adjusted so as to obtain data similar to the data obtained in the real world. In this way, the real world is reproduced in the physical simulation CutSim and the robot simulation RoboSim.
[0034] [Connection Between Simulators] In this embodiment, it is assumed that the time step in the physical simulation CutSim and the time step in the robot simulation RoboSim are different. Therefore, in this embodiment, the time step in the physical simulation CutSim and the time step in the robot simulation RoboSim are synchronized. Details will be described later.
[0035] (C. Cutting Learning by Reinforcement Learning) [Markov Decision Process] In reinforcement learning, an agent acquires a desired behavior through interactions with the outside world. The task of cutting an object can be defined by the framework of a Markov decision process (MDP). When the task of cutting an object is defined by a Markov decision process, the task of cutting an object is defined by T time steps per episode. A Markov decision process is defined by a tuple (S, A, P, R, γ). S is the state space, A is the action space, and P:S × A is the transition probability function. R represents the reward function, and γ∈(0,1) represents the discount rate. At each time step t, the agent changes its current state s t Observe ∈S and take action a according to the policy π t ∈A. State s t is the transition function p(s t+1 |s t , a t ) according to s t+1 At this time, the agent takes action a t By taking the reward r t = R(s t+1 , s t , a t ) is received. t By taking t is the state s t+1 Transition to.
[0036] The final goal in reinforcement learning is to find a policy π that maximizes the reward function R(t) given by the following equation: * In the following, i represents the time step.
[0037]
[0038] In this embodiment, Soft-Actor Critic (SAC), which is an example of a reinforcement learning algorithm, is employed. The SAC algorithm is disclosed in, for example, Reference 4 below.
[0039] Reference 4: T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off policy maximum entropy deep reinforcement learning with a stochastic actor,” in International Conference on Machine Learning, 2018, pp.1861-1870.
[0040] The SAC algorithm is an algorithm based on entropy maximization. Specifically, in the SAC algorithm, the expected reward E π The policy π that maximizes * In the following equation, H is the entropy, and in the SAC algorithm, the strategy π is determined so that the entropy H is maximized. * is determined. Here, α is a ratio to a preset entropy. In addition, the SAC algorithm employs a known Off-Policy algorithm.
[0041]
[0042] [Compliance Control] As described above, the present embodiment employs a known FDCC, which is capable of performing control on Cartesian coordinates by combining impedance control, admittance control, and force control.
[0043] FIG. 4 is a diagram for explaining FDCC. d and external force F d 4 is also a diagram showing how to control the robot's end effector when F is given. The x shown in FIG. 4 is a parameter that includes the position of the robot's end effector in Cartesian coordinates and the position of the robot's joints in Cartesian coordinates. F ext is the measured contact force. d is the target posture of the robot, and F d is the external force generated on the end effector of the robot. Fc is a control command that represents the force to be applied to the end effector of the robot. -1 J T is the forward dynamics model in the virtual model, and FK(q c ) is the forward kinematics model in the virtual model. c corresponds to the position of the robot joint, and q c・・ corresponds to the acceleration of the robot's joints. Also, the net force F net includes all forces measured by the sensors and those resulting from the virtual motion. Therefore, the compliance controller shown in FIG. 4 uses a parameter K c and a parameter K representing the PD (Proportional and Differential) gain. p and a parameter K representing the PD gain d As will be described later, in this embodiment, a parameter K representing stiffness is calculated using a trained model. c and a parameter K representing the PD gain p and a parameter K representing the PD gain d In the following, the parameter K c and a parameter K representing the PD gain p and a parameter K representing the PD gain d The FDCC and the respective parameters are also referred to as control parameters. Note that the FDCC and the respective parameters are disclosed in the above-mentioned Reference 1, so please refer to Reference 1 for details.
[0044] [Reinforcement Learning Agent] Fig. 5 is a diagram for explaining the relationship between reinforcement learning and the compliance controller in this embodiment. As shown in Fig. 5, an actor, which is a reinforcement learning agent, performs a robot movement trajectory x c and the control parameter [K c , K p , K d ] to the compliance controller. The actor controls the motion trajectory x c and the control parameter [Kc , K p , K d ] is output. Note that the robot's motion trajectory x c past target posture x d By adding to the target attitude x d Therefore, the target attitude x d and motion trajectory x c The sum of these is the target attitude x d Then, the FDCC, which is a compliance controller, calculates the motion trajectory x c and the control parameter [K c , K p , K d ] and controls the joints of the robot. Feedback from the robot is the posture of the end effector (or the posture of the knife) and the forces and moments detected by the robot. The robot is equipped with sensors to detect forces and torques at n locations.
[0045] The physical quantities observed by the actor, which is an agent, are the relative position of the knife with respect to the target position, the speed of the knife, the jerk of the knife, the history of the robot's actions, and the history of the forces and torques acting on the robot at n locations. Therefore, in this embodiment, the observation data consisting of these physical quantities is called state data s t The state data s corresponding to the observation data is used as t is a vector having these physical quantities as components. Each physical quantity can be obtained from various sensors provided on the robot or from past operation data of the robot.
[0046] The actors in the SAC algorithm can be realized by, for example, a known temporal convolutional network, which is an example of a machine learning model. The temporal convolutional network can process the history of forces and torques applied to the robot at n locations, and the fully connected network included in the temporal convolutional network can store the observation history.
[0047] [Reward Function] The following formula (1) is the reward function set in this embodiment. As shown in the following formula (1), the reward function is the height x of the knife from the cutting board. cut and contact force F ext and the knife's jerk x ・・・ It is defined based on the following. 1 , w 2 , w 3 is a preset weight parameter. As shown in the following equation (1), the reward function incorporates a reward B that corresponds to the situation at the end of the episode representing a series of actions of the robot. Specifically, when the task of cutting an object with a knife is accomplished (for example, when the knife reaches the cutting board with a force within a predetermined threshold), "100" is substituted for reward B as shown in the following equation. When the knife collides with the cutting board (for example, when the knife reaches the cutting board with a force greater than a predetermined threshold), "-100" is substituted for reward B as shown in the following equation. When the knife moves outside the preset space or when it takes too long to complete the task, "-1" is substituted for reward B as shown in the following equation. As shown in the following equation, when the jerk x of the knife ・・・ The larger the value, the smaller the reward r. ・・・ corresponds to a jerky movement of the knife, so in this embodiment, a reward is set to suppress such movement.
[0048]
[0049] <Control System 10> Fig. 6 is a block diagram showing a schematic configuration of the control system 10 of this embodiment. As shown in Fig. 6, the control system 10 includes a sensor group 11, a real robot 12R, and a control device 14. The control device 14 generates a trained model for controlling the movement of the end effector 16R of the real robot 12R. The control device 14 also controls the movement of the end effector 16R of the real robot 12R using the generated trained model.
[0050] The sensor group 11 is attached to the end effector 16R of the real robot 12R and sequentially measures the above-mentioned observation data. The sensor group 11 measures the position of the real knife KR, the velocity of the real knife KR, the jerk of the real knife KR, the behavior history of the real robot 12R, and the history of forces and torques applied to the real robot 12R at n locations. For example, the forces and torques applied to the real robot 12R at n locations are so-called force sense values (e.g., three-axis torque and three-axis force). The sensor group 11 then outputs the obtained observation data to the control device 14.
[0051] 7 is a block diagram showing the hardware configuration of the control device 14 according to this embodiment. As shown in FIG. 7, the control device 14 includes a CPU (Central Processing Unit) 42, a memory 44, a storage device 46, an input / output I / F (Interface) 48, a storage medium reader 50, and a communication I / F 52. Each component is connected to each other via a bus 54 so as to be able to communicate with each other.
[0052] The storage device 46 stores a trained model generation program and a control program for executing each process described below. The CPU 42 is a central processing unit that executes various programs and controls each component. That is, the CPU 42 reads the programs from the storage device 46 and executes the programs using the memory 44 as a work area. The CPU 42 controls each component and performs various arithmetic processes in accordance with the programs stored in the storage device 46.
[0053] The memory 44 is configured with a RAM (Random Access Memory) and serves as a working area to temporarily store programs and data. The storage device 46 is configured with a ROM (Read Only Memory), an HDD (Hard Disk Drive), an SSD (Solid State Drive), etc., and stores various programs including the operating system and various data.
[0054] The input / output I / F 48 is an interface for inputting data from an external device and outputting data to an external device. Input devices for various inputs, such as a keyboard or a mouse, and output devices for various information outputs, such as a display or a printer, may also be connected. A touch panel display may be used as the output device, allowing it to function as an input device.
[0055] The storage medium reader 50 reads data stored in various storage media such as CD (Compact Disc)-ROM, DVD (Digital Versatile Disc)-ROM, Blu-ray Disc, and USB (Universal Serial Bus) memory, and writes data to the storage media.
[0056] The communication I / F 52 is an interface for communicating with other devices, and uses standards such as Ethernet (registered trademark), FDDI, and Wi-Fi (registered trademark).
[0057] Next, the functional configuration of the control device 14 will be described. As shown in Fig. 6, the control device 14 functionally includes a learning acquisition unit 18, a simulation unit 20, a generation unit 22, an acquisition unit 24, and a control unit 26. A predetermined storage area of the control device 14 also includes a data storage unit 28, a trained model storage unit 30, and a controller storage unit 32. Each functional configuration is realized by the CPU 42 reading each program stored in the storage device 46, expanding it into the memory 44, and executing it.
[0058] The data storage unit 28 stores observation data detected by the sensor group 11. The data storage unit 28 also stores control data when the end effector 16R of the real robot 12R operates. For example, the data storage unit 28 stores a history of behavior data when the end effector 16R of the real robot 12R operates.
[0059] The trained model storage unit 30 stores trained policies π generated by the process described below. * and the learned action value function Q * and is stored.
[0060] The controller storage unit 32 stores information representing the above-mentioned FDCC model. Specifically, the FDCC stored in the controller storage unit 32 is loaded with a learned policy π * The control parameter K output from c , K p , K d By substituting the above, the position q of the joint of the real robot 12R is obtained. c will be output.
[0061] First, the learning acquisition unit 18, the simulation unit 20, and the generation unit 22 acquire a learned policy π, which is a learned model for controlling the operation of the end effector 16R of the real robot 12R. * Generate.
[0062] First, the end effector 16R of the real robot 12R is made to grasp a real knife KR and execute a task of cutting a real object BR with the real knife KR. Then, the learning acquisition unit 18 acquires real reaction force data representing a real reaction force obtained while the real knife KR of the real robot 12R is cutting the real object, and stores the real reaction force data in the data storage unit 28. This real reaction force data is measured by a force sensor or the like included in the sensor group 11.
[0063] Next, the simulation unit 20 reads out the actual reaction force data stored in the data storage unit 28 and reflects the actual reaction force data in the physical simulation. Specifically, the simulation unit 20 reflects the actual reaction force data in the physical simulation by adjusting various parameters when executing the physical simulation.
[0064] 2, the simulation unit 20 executes a physical simulation CutSim that simulates a reaction force from the virtual object BS when a force is applied to the virtual object BS, and a robot simulation RoboSim that simulates the operation of the virtual robot 12S that processes the virtual object BS. The simulation unit 20 executes calculations for the physical simulation CutSim in a first period, and executes calculations for the physical simulation RoboSim in a second period. The first period and the second period represent time steps of the simulation. For example, in the first period dt 1 = 1.0 × 10 -5 [s], second period dt 2 = 1.0 × 10 -4 [s]There is.
[0065] The generation unit 22 generates a trained model that outputs behavior data of the real robot 12R corresponding to the virtual robot 12S when state data representing the environment in which the real robot 12R operates is input, based on data obtained while the physical simulation CutSim and the robot simulation RoboSim are being executed by the simulation unit 20. Note that the data obtained while the physical simulation is being executed is virtual reaction force data representing the reaction force when the virtual robot 12S applies a force to the virtual object BS. The trained model in this embodiment is based on the trained policy π * corresponds to the learned policy π * When state data is input, the state data in this embodiment is output as action data. As described above, the state data in this embodiment is the position of the real knife KR, the velocity of the real knife KR, the jerk of the real knife KR, and the action data of the real robot 12R (the motion trajectory x of the end effector 16R). c , control parameter K c , K p , K d ) and the history of forces and torques applied to n locations of the real robot 12R. c and the control parameter K c , K p , K dis.
[0066] Specifically, the generation unit 22 uses the SAC algorithm, which is an example of the reinforcement learning algorithm described above, to learn an actor corresponding to a policy in reinforcement learning and a critic corresponding to an action value function in reinforcement learning. As shown in FIG. 5, the actor is trained based on state data s t When input, behavioral data a t The motion trajectory x of the end effector 16R is c and the control parameter K c , K p , K d As shown in FIG. 5, the critic outputs the state data s corresponding to the observation data. t and state data s t When the reward r calculated by the above formula (1) according to the actor is input, the value of the action value function is output. The actor and critic are realized by a known machine learning model.
[0067] The generation unit 22 trains the actors and critics in accordance with the SAC algorithm so that the sum of the rewards r calculated by the above formula (1) becomes large, thereby generating a trained policy π corresponding to the trained actor. * and the learned action-value function Q corresponding to the learned critic * Then, the generation unit 22 generates the learned policy π * and the learned action value function Q * and are stored in the trained model storage unit 30. The reward r is set in advance depending on the type of task to be executed by the virtual robot 12S.
[0068] The generation unit 22 generates the learned policy π * When generating the first period dt 1 From the virtual reaction force data obtained from the physical simulation CutSim by the calculation, the second period dt 2 The virtual reaction force data corresponding to the second period dt 2 Based on the virtual reaction force data corresponding to *For example, the generation unit 22 generates a first period dt 1 From the virtual reaction force data for each period, the second period dt 2 The virtual reaction force data is acquired, and the acquired second period dt 2 Using the virtual reaction force data, the learned policy π * As a result, even if the calculation period of the physical simulation CutSim and the calculation period of the robot simulation RoboSim are different, the learned policy π * can be generated.
[0069] Furthermore, the generation unit 22 generates a learned policy π * When generating the first period dt 1 The virtual reaction force data is acquired at a period greater than π, and the learned policy π is calculated based on the acquired virtual reaction force data. * For example, the fluctuation of reaction force data while cutting an object with a knife may be small. In such a case, even if virtual reaction force data is acquired in a short time interval, the virtual reaction force data acquired in that time interval is almost constant, so the learned policy π * Therefore, when the fluctuation of the virtual reaction force data is within a predetermined range, the first period dt 1 a period greater than (for example, the first period dt 1 Alternatively, the virtual reaction force data may be acquired at a period (e.g., twice the period of the reference period).
[0070] Furthermore, the generation unit 22 generates a learned policy π * When generating the first period dt 1 The virtual reaction force data is acquired at a period smaller than * For example, if the reaction force data fluctuates significantly while cutting an object with a knife, the learned policy π *Therefore, when the fluctuation of the virtual reaction force data is outside the predetermined range, the first period dt 1 a period smaller than (for example, the first period dt 1 Alternatively, the virtual reaction force data may be acquired at a period (e.g., a period equal to 1 / 2 of the period).
[0071] Learned policy π * is stored in the trained model storage unit 30, the trained policy π * Therefore, the acquisition unit 24 and the control unit 26 can control the operation of the end effector 16R of the real robot 12R using the learned policy π * and the FDCC stored in the controller memory unit 32 to control the operation of the end effector 16R of the real robot 12R.
[0072] The acquisition unit 24 acquires state data s representing the environment in which the real robot 12R operates. t Get.
[0073] The control unit 26 calculates the trained policy π stored in the trained model storage unit 30. * Then, the control unit 26 reads out the status data s acquired by the acquisition unit 24. t The learned policy π * By inputting the data into the t As mentioned above, we obtain the learned policy π * From the above, the motion trajectory x of the end effector 16R is c and the control parameter K c , K p , K d and behavioral data a t is output as
[0074] The control unit 26 determines the learned policy π * The control parameter K output from c , K p , K d to the FDCC stored in the controller storage unit 32, the position q of the joint of the real robot 12R output from the FDCC c Get.
[0075] Then, the control unit 26 calculates the joint position q of the real robot 12R output from the FDCC. c and the motion trajectory x of the end effector 16R. c Specifically, the control unit 26 controls the real robot 12R based on the joint positions q c and the motion trajectory x of the end effector 16R. c The control signal is output to the real robot 12R so that the above is realized.
[0076] Next, the operation of the control system 10 according to this embodiment will be described.
[0077] First, the real robot 12R executes a task such as cutting a real object BR using a real knife KR, and the real reaction force data from this task is input to the control device 14 and stored in the data storage unit 28. Then, when the control device 14 receives a predetermined instruction signal, the CPU 42 of the control device 14 reads out the trained model generation program from the storage device 46, loads it into the memory 44, and executes it. As a result, the CPU 42 functions as each functional component of the control device 14, and the trained model generation process shown in FIG. 8 is executed.
[0078] In step S100, the simulation unit 20 acquires the actual reaction force data stored in the data storage unit .
[0079] In step S102, the simulation unit 20 sets various parameters of the physical simulation CutSim by reflecting the actual reaction force data acquired in step S100 in the physical simulation CutSim.
[0080] In step S104, the simulation unit 20 starts executing the physical simulation CutSim and the robot simulation RoboSim.
[0081] In step S106, the generation unit 22 generates state data s representing the environment in which the virtual robot 12S operates, based on data obtained while the physical simulation CutSim and the robot simulation RoboSim are being executed.t When this is input, the behavior data a of the virtual robot 12S is t A trained policy π that outputs * Specifically, the generation unit 22 generates a learned policy π corresponding to the learned actor by training the actor and the critic in accordance with the SAC algorithm so as to maximize the reward r calculated by the above formula (1). * and the learned action-value function Q corresponding to the learned critic * and generate.
[0082] In step S108, the generation unit 22 generates the learned policy π * and the learned action value function Q * and are stored in the trained model storage unit 30.
[0083] Next, when the control device 14 receives a predetermined instruction signal, the control device 14 executes the control process shown in Fig. 9. The control process in Fig. 9 is executed repeatedly.
[0084] In step S200, the acquisition unit 24 acquires state data s representing the environment in which the real robot 12R operates. t Get.
[0085] In step S202, the control unit 26 calculates the trained policy π stored in the trained model storage unit 30. * Then, the control unit 26 reads out the status data s acquired in step S200. t The learned policy π * By inputting the data into the t Get.
[0086] In step S204, the control unit 26 uses the behavior data a acquired in step S202 t The control parameter K c , K p , K d to the FDCC stored in the controller storage unit 32, the position q of the joint of the real robot 12R output from the FDCC c Get.
[0087] In step S206, the control unit 26 calculates the joint position q of the real robot 12R acquired in step S204. c and the behavior data a acquired in step S202 t The motion trajectory x of the end effector 16R c Based on this, the real robot 12R is controlled.
[0088] As described above, the control device according to this embodiment executes a physics simulation that simulates a reaction force from a virtual object when a force is applied to the virtual object, and a robot simulation that simulates the operation of a virtual robot that processes the virtual object. Based on data obtained during the execution of the physics simulation and the robot simulation, the control device generates a trained policy that outputs behavior data of the real robot when state data representing the environment in which the real robot corresponding to the virtual robot operates is input. This allows for efficient generation of a trained model used when the robot executes a task of applying a force to an object.
[0089] Furthermore, the control device according to this embodiment acquires state data representing the environment in which the real robot operates, inputs the acquired state data into the generated learned policy, acquires behavior data of the real robot, and controls the real robot based on the behavior data, thereby enabling the robot to perform a task of applying force to an object.
[0090] Furthermore, with the control device according to this embodiment, even if the processing task for an object is irreversible, it is possible to repeatedly acquire data by using physical simulation and robot simulation. This makes it possible to acquire more data on processing without actually processing the real object, and to efficiently generate a trained model.
[0091] Next, an example will be described. In this example, a simulation is performed to verify the effectiveness of the proposed method. In this simulation, an experiment was performed on a task of slicing real objects. The real objects used in the experiment were a cucumber, a tomato, a potato, and a carrot. The number of slices and the size of the slices were set as shown in the table below.
[0092]
[0093] 10A and 10B show the results of calibration performed by reflecting real-world data in a physical simulation. FIG. 10A shows the results when the real object is a cucumber, FIG. 10B shows the results when the real object is a tomato, and FIG. 10C shows the results when the real object is a potato. The horizontal axis of the graph shown in FIG. 10 represents time, and the vertical axis represents contact force. In FIG. 10, "Disect" represents the results of the physical simulation reflecting real-world data, "Gazebo" represents the results of the robot simulation, and "Ground Truth" represents real contact force data. As shown in FIG. 10, it can be seen that the contact force in "Disect" is similar to the contact force in "Ground Truth" due to the calibration performed on the physical simulation.
[0094] 11 is a diagram showing the results of the variation in contact force when a trained model is generated using only the robot simulation "Gazebo" (denoted as "Gazebo only" in FIG. 11 ) and when a trained model is generated by combining the physical simulation "Disect" and the robot simulation "Gazebo" (denoted as "Disect + Gazebo" in FIG. 11 ). As shown in FIG. 11 , "Disect + Gazebo" has smaller variation in contact force than "Gazebo only," and can slice the object more smoothly.
[0095] Figure 12 shows the time series of contact force when slicing a cucumber twice. As shown in Figure 12, "Disect + Gazebo" requires less contact force than "Gazebo only," and can slice the cucumber more smoothly with less force.
[0096] As can be seen from FIGS. 11 and 12, by using the method of this embodiment, it is possible to accurately and efficiently generate a learned policy to be used when applying a force to an object.
[0097] In the above embodiment, the task of applying a force to an object is described as a task of cutting an object. However, this is not limited to this. Any task of applying a force to an object may be used. For example, this embodiment can be applied to tasks such as bending parts, cutting parts, food processing (e.g., shaping bread), and gripping soft parts. This embodiment can also be applied to tasks that do not require irreversibility. For example, this embodiment can be applied to tasks that do not require irreversibility and require time for setting up. For example, this embodiment can be applied to tasks such as laying a tablecloth, food processing (e.g., shaping bread), and assembly that requires multiple steps (e.g., assembly tasks that require disassembly and reassembly if a step fails).
[0098] In the above embodiment, a reinforcement learning algorithm is used as the machine learning algorithm, but the present invention is not limited to this. Other machine learning algorithms (e.g., supervised learning algorithms or unsupervised learning algorithms) may be used to generate the learned data.
[0099] In the above embodiment, a part of the behavior data output from the trained model is input to the FDCC, and the joint position q cAlthough the above embodiment has been described with reference to an example in which a robot is controlled based on the FDCC, the present invention is not limited to this. As the behavior data, data of a different type from that in the above embodiment may be used, and in this case, it is also possible to control the robot according to the behavior data without using the FDCC. Furthermore, the state data is not limited to that in the above embodiment, and any data may be used as long as it represents the environment of the robot.
[0100] In the above embodiment, the reward r is defined by the formula (1), but the present invention is not limited to this. For example, the reward r can be changed as appropriate depending on the task to be performed by the robot.
[0101] In the above embodiment, the first period dt 1 and the second period dt, which is the time step of the robot simulation RoboSim 2 However, the present invention is not limited to this. For example, the first period dt 1 and the second period dt 2 may be made the same as
[0102] In the above embodiment, the control device 14 executes both the trained model generation process of Fig. 8 and the control process of Fig. 9 , but the present invention is not limited to this. For example, a trained model generation device implemented by a computer separate from the control device 14 may be provided, and the trained model generation device may execute the trained model generation process of Fig. 9 , while the control device 14 executes the control process of Fig. 10 . In this case, the trained model generation device includes at least the learning acquisition unit 18, the simulation unit 20, and the generation unit 22 described above.
[0103] Furthermore, various processors other than the CPU may execute the processes executed by loading software (programs) in the above embodiments. Examples of processors in this case include programmable logic devices (PLDs) (such as field-programmable gate arrays (FPGAs)) whose circuit configuration can be changed after manufacture, and dedicated electrical circuits, such as application-specific integrated circuits (ASICs), which are processors having circuit configurations designed specifically to execute specific processes. Each process may be executed by one of these various processors, or by a combination of two or more processors of the same or different types (e.g., multiple FPGAs, or a combination of a CPU and an FPGA). The hardware structure of these various processors is, more specifically, an electrical circuit that combines circuit elements such as semiconductor elements.
[0104] In the above embodiment, the programs are pre-stored (installed) in a storage device, but this is not limiting. The programs may be provided in a form stored in a storage medium such as a CD-ROM, DVD-ROM, Blu-ray Disc, or USB memory. The programs may also be downloaded from an external device via a network.
[0105] (Supplementary Notes) The following supplementary notes are provided regarding aspects of the present disclosure.
[0106] (Supplementary Note 1) A trained model generation device comprising: a simulation unit that executes a physical simulation that simulates a reaction force from a virtual object when a force is applied to the virtual object, and a robot simulation that simulates the operation of a virtual robot that applies a force to the virtual object, and a generation unit that generates a trained model that outputs behavior data of a real robot corresponding to the virtual robot when state data that represents an environment in which the real robot operates is input, based on data obtained while the physical simulation and the robot simulation are being executed. (Supplementary Note 2) The trained model generation device according to Supplementary Note 1, wherein the data obtained while the physical simulation is being executed is virtual reaction force data that represents a reaction force when the virtual robot applies a force to the virtual object, and the generation unit generates the trained model based on the virtual reaction force data. (Supplementary Note 3) The trained model generation device according to Supplementary Note 2, wherein the simulation unit performs calculations of the physical simulation in a first period and performs calculations of the robot simulation in a second period, and the generation unit, when generating the trained model, acquires the virtual reaction force data corresponding to the second period in the robot simulation from the virtual reaction force data obtained from the physical simulation by calculation in the first period, and generates the trained model based on the virtual reaction force data corresponding to the second period. (Supplementary Note 4) The trained model generation device according to Supplementary Note 3, wherein, when generating the trained model, the generation unit acquires the virtual reaction force data at a period longer than the second period if fluctuations in the virtual reaction force data are within a predetermined range, and generates the trained model based on the acquired virtual reaction force data. (Supplementary Note 5) The trained model generation device described in Supplementary Note 3 or Supplementary Note 4, wherein when generating the trained model, if fluctuations in the virtual reaction force data are outside a predetermined range, the generation unit acquires the virtual reaction force data at a period shorter than the second period, and generates the trained model based on the acquired virtual reaction force data.(Supplementary Note 6) The trained model generation device according to any one of Supplementary Notes 1 to 5, wherein the simulation unit acquires real reaction force data representing a real reaction force when a force is applied by a real robot corresponding to the virtual robot to a real object corresponding to the virtual object, and executes the physical simulation by reflecting the real reaction force data in the physical simulation. (Supplementary Note 7) The trained model generation device according to any one of Supplementary Notes 1 to 6, wherein the generation unit generates the trained model by reinforcement learning so as to increase a sum of rewards that are set in advance according to types of tasks to be performed by the virtual robot. (Supplementary Note 8) The trained model generation device according to any one of Supplementary Notes 1 to 7, wherein the generation unit generates the trained model by reinforcement learning so as to increase a reward according to a situation at an end of an episode representing a series of actions of the virtual robot. (Supplementary Note 9) The trained model generation device according to any one of Supplementary Notes 1 to 8, wherein the generation unit, when executing reinforcement learning according to a Soft Actor-Critic algorithm, learns an actor representing a policy in reinforcement learning and a critic representing an action value function in reinforcement learning based on data obtained while the physical simulation and the robot simulation are being executed, and generates the learned policy as the trained model. (Supplementary Note 10) A control device comprising: an acquisition unit that acquires the state data, and a control unit that acquires behavior data of the real robot by inputting the state data acquired by the acquisition unit into the trained model generated by the trained model generation device according to any one of Supplementary Notes 1 to 9, and controls the real robot based on the behavior data.(Supplementary Note 11) A trained model generation method in which a computer executes the following processes: executing a physics simulation that simulates a reaction force when a force is applied to a virtual object, and a robot simulation that simulates the operation of a virtual robot that applies a force to the virtual object, and generating a trained model that outputs behavior data of a real robot corresponding to the virtual robot when state data representing an environment in which the real robot operates is input, based on data obtained while the physics simulation and the robot simulation are being executed. (Supplementary Note 12) A trained model generation program for causing a computer to execute the following processes: executing a physics simulation that simulates a reaction force when a force is applied to a virtual object, and a robot simulation that simulates the operation of the virtual robot that applies a force to the virtual object, and generating a trained model that outputs behavior data of the real robot when state data representing an environment in which the real robot corresponding to the virtual robot operates is input, based on data obtained while the physics simulation and the robot simulation are being executed.
[0107] The disclosure of Japanese Patent Application No. 2024-007062, filed on January 19, 2024, is incorporated herein by reference in its entirety. All documents, patent applications, and technical standards mentioned herein are incorporated herein by reference to the same extent as if each individual document, patent application, and technical standard was specifically and individually indicated to be incorporated by reference.
Claims
1. A learned model generation device including: a simulation unit that executes a physical simulation for simulating a reaction force from a virtual object when a force is applied to the virtual object, and a robot simulation for simulating the operation of a virtual robot that applies a force to the virtual object; and a generation unit that generates a learned model that outputs action data of a real robot when state data representing an environment in which the real robot corresponding to the virtual robot operates is input based on data obtained while the physical simulation and the robot simulation are being executed.
2. The data obtained while the physical simulation is being executed is virtual reaction force data representing the reaction force when the virtual robot applies a force to the virtual object, and the generation unit generates the learned model based on the virtual reaction force data. The learned model generation device according to claim 1.
3. The simulation unit executes the calculation of the physical simulation in a first cycle and executes the calculation of the robot simulation in a second cycle. When generating the learned model, the generation unit acquires the virtual reaction force data corresponding to the second cycle in the robot simulation from the virtual reaction force data obtained from the physical simulation by the calculation in the first cycle, and generates the learned model based on the virtual reaction force data corresponding to the second cycle. The learned model generation device according to claim 2.
4. When generating the learned model, if the variation of the virtual reaction force data is within a predetermined range, the generation unit acquires the virtual reaction force data at a cycle larger than the second cycle, and generates the learned model based on the acquired virtual reaction force data. The learned model generation device according to claim 3.
5. When generating the learned model, if the variation of the virtual reaction force data is outside a predetermined range, the generation unit acquires the virtual reaction force data at a cycle smaller than the second cycle, and generates the learned model based on the acquired virtual reaction force data. The learned model generation device according to claim 3 or claim 4.
6. The simulation unit acquires real reaction force data representing a real reaction force when a real robot corresponding to the virtual robot applies a force to a real object corresponding to the virtual object, and executes the physical simulation by reflecting the real reaction force data in the physical simulation. The learned model generation device according to claim 1 or claim 2.
7. The generation unit generates the learned model by reinforcement learning so that the total reward preset according to the type of task executed by the virtual robot increases. The learned model generation device according to claim 1 or claim 2.
8. The generation unit generates the learned model by reinforcement learning so that the reward corresponding to the situation at the end of an episode representing a series of actions of the virtual robot increases. The learned model generation device according to claim 1 or claim 2.
9. When executing reinforcement learning according to the Soft Actor-Critic algorithm, the generation unit learns an actor representing a policy in reinforcement learning and a critic representing an action value function in reinforcement learning based on data obtained while the physical simulation and the robot simulation are being executed, and generates the learned policy as the learned model. The learned model generation device according to claim 1 or claim 2.
10. An acquisition unit that acquires the state data, and a control unit that acquires the action data of the real robot by inputting the state data acquired by the acquisition unit to the learned model generated by the learned model generation device according to claim 1 or claim 2, and controls the real robot based on the action data. A control device comprising:
11. A learned model generation method executed by a computer, which performs a physical simulation for simulating a reaction force when a force is applied to a virtual object, and a robot simulation for simulating the operation of a virtual robot that applies a force to the virtual object, and generates a learned model that outputs action data of the real robot when state data representing an environment in which the real robot corresponding to the virtual robot operates is input based on data obtained while the physical simulation and the robot simulation are being executed.
12. A learned model generation program for causing a computer to execute a process of performing a physical simulation for simulating a reaction force when a force is applied to a virtual object, and a robot simulation for simulating the operation of a virtual robot that applies a force to the virtual object, and generating a learned model that outputs action data of the real robot when state data representing an environment in which the real robot corresponding to the virtual robot operates is input based on data obtained while the physical simulation and the robot simulation are being executed.
Citation Information
Patent Citations
Master-slave type robot arm device and arm positioning / guiding method
JP1996025254A
Variational grasp generation
JP2020192676A
Device for generating learning data, method for generating learning data, and machine learning device and machine learning method using learning data
WO2023073780A1