Learned model generation apparatus, control apparatus, learned model generation method, and learned model generation program
By integrating physical and robot simulations with reinforcement learning, the system efficiently generates a learned model for robot tasks, addressing the challenge of obtaining reaction force data, enabling accurate and efficient force application.
Patent Information
- Application Number
- JP2024007062
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-19
- Publication Date
- 2025-08-01
AI Technical Summary
Existing methods struggle to efficiently generate a learned model for robot operations that apply force to objects, as obtaining reaction force data is time-consuming and difficult, whether from real-world objects or virtual simulations, leading to inefficiencies in training.
A system that combines physical and robot simulations to generate a learned model by simulating reaction forces and robot operations, using reinforcement learning algorithms like Soft-Actor Critic (SAC) to adjust simulation parameters based on real-world data, enabling efficient generation of a learned model for robot tasks.
This approach allows for the efficient generation and utilization of a learned model for robot tasks, even with irreversibility, by repeatedly acquiring data through simulations, thus improving the robot's ability to apply forces accurately and efficiently.
Smart Images

Figure 2025112678000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a learned model generation device, a control device, a learned model generation method, and a learned model generation program.
Background Art
[0002] Conventionally, a robot system that stably acquires work information according to the situation is known (see, for example, Patent Document 1). This robot system performs a physical simulation in which a virtual work is freely dropped into a virtual container, virtually images the virtual work stacked in the virtual container with a virtual imaging device to acquire image data, and acquires a learned model by machine learning using the image data as input data.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] By the way, it is difficult to train a robot to perform an operation such as applying a force to an object. For example, consider a case where a robot grasps a knife and performs a task of cutting an object using the knife. In this case, when the robot cuts the object with the knife, the robot needs to learn points such as at what angle to apply the knife to the object, how much force to apply to the knife, and at what speed to move the knife.
[0005] When the robot learns its operation taking the above points into consideration, it is necessary to obtain data on the reaction force from the object generated when applying a force to the object. For example, in this case, it is necessary to use the reaction force data to train a learned model using a machine learning algorithm, and the robot needs to control its own operation using the learned model.
[0006] However, after applying a force to the object to obtain reaction force data, the object may change, and it may be difficult to obtain reaction force data from the object again. For example, as described above, once the object is cut, it is difficult to obtain reaction force data thereafter.
[0007] In such a case, if an attempt is made to obtain reaction force data for training the robot from the real world, it is time-consuming and difficult, such as having to prepare a huge number of objects. On the other hand, even if an attempt is made to obtain virtual reaction force data using a robot simulator, it is difficult to obtain accurate reaction force data. For this reason, in the case described above, there is a problem that reaction force data when applying a force to an object cannot be obtained efficiently, and a learned model used when the robot applies a force to an object cannot be generated efficiently.
[0008] In Patent Document 1 above, when training the robot to perform an operation of gripping an object, a physical simulation is performed in which a virtual workpiece is freely dropped into a virtual container, and the scattered state of the virtual workpiece is virtually reproduced. Therefore, Patent Document 1 is a technique for virtually reproducing the scattered state of a virtual workpiece, and does not train the robot using reaction force data when the robot applies a force to an object.
[0009] The present disclosure has been made in view of the above points, and an object thereof is to efficiently generate a learned model used when a robot executes a task of applying a force to an object.
Means for Solving the Problems
[0010] To achieve the above object, a learned model generation apparatus according to the present disclosure includes a simulation unit that executes a physical simulation for simulating a reaction force when a force is applied to a virtual object and a robot simulation for simulating the operation of a virtual robot that applies a force to the virtual object, and a generation unit that generates a learned model that outputs action data of a real robot corresponding to the virtual robot when state data representing an environment in which the real robot operates is input based on data obtained while the physical simulation and the robot simulation are being executed.
[0011] Further, a learned model generation method of the present disclosure executes a physical simulation for simulating a reaction force when a force is applied to a virtual object and a robot simulation for simulating the operation of a virtual robot that applies a force to the virtual object, and generates a learned model that outputs action data of a real robot corresponding to the virtual robot when state data representing an environment in which the real robot operates is input based on data obtained while the physical simulation and the robot simulation are being executed, and is a learned model generation method in which a computer executes the process.
[0012] Further, a learned model generation program of the present disclosure executes a physical simulation for simulating a reaction force when a force is applied to a virtual object and a robot simulation for simulating the operation of a virtual robot that applies a force to the virtual object, and generates a learned model that outputs action data of a real robot corresponding to the virtual robot when state data representing an environment in which the real robot operates is input based on data obtained while the physical simulation and the robot simulation are being executed, and is a learned model generation program for causing a computer to execute the process.
Advantages of the Invention
[0013] According to the learned model generation device, learned model generation method, and learned model generation program of the present disclosure, a learned model used when a robot executes a task of applying a force to an object can be efficiently generated. Further, according to the control device of the present disclosure, a task of applying a force to an object can be executed using the efficiently generated learned model.
Brief Description of Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Embodiments for Carrying Out the Invention
[0015] Hereinafter, an example of an embodiment of the present disclosure will be described with reference to the drawings. In this embodiment, a control system equipped with a control device according to the present disclosure will be described as an example. In each drawing, the same or equivalent components and parts are given the same reference numerals. Also, the dimensions and ratios in the drawings are exaggerated for the convenience of explanation and may be different from the actual ratios.
[0016] FIG. 1 is a diagram for explaining this embodiment. As shown in FIG. 1, the control system 10 of this embodiment includes a real robot 12R and a control device 14. The real robot 12R is provided with an end effector 16R. The control system 10 efficiently generates a learned model used when the real robot 12R performs a task of applying a force to a real object BR. For example, the task of applying a force targeted by the control system 10 of this embodiment is a task having irreversibility. For example, as shown in FIG. 1, the real robot 12R of the control system 10 holds a real knife KR, and the real robot 12R operates the end effector 16R in accordance with a control signal output from the control device 14 to execute a task of cutting the real object BR using the real knife KR. The task of cutting the real object BR is a task in which once the real object BR is cut, the same location of the real object BR cannot be cut again, and it is a task having irreversibility.
[0017] FIG. 2 is a diagram for explaining the outline of this embodiment. As shown in FIG. 2, in this embodiment, a learned model used when the real robot 12R in the real world Re executes a task of cutting the real object BR using the real knife KR is generated. At this time, the control system 10 of this embodiment executes two types of simulations Sim.
[0018] The virtual object BS shown in FIG. 2 is a virtual realization of the real object BR on a computer. Similarly, the virtual robot 12S is a virtual realization of the real robot 12R on a computer.
[0019] The control system 10 of this embodiment executes a physical simulation CutSim that simulates the reaction force from the virtual object BS when applying a force to the virtual object BS, and a robot simulation RoboSim that simulates the operation of the virtual robot 12S that processes the virtual object BS.
[0020] Currently widely used robot simulations are simulations that simulate the internal state of the virtual robot 12S. Such robot simulations are difficult to simulate the external state of the virtual robot 12S. Also, currently widely used physical simulations are difficult to execute simulations that simulate the inside of the virtual robot 12S.
[0021] Therefore, the control system 10 of this embodiment generates a learned model used when executing a task of applying a force to an object based on the data obtained by executing both the physical simulation CutSim and the robot simulation RoboSim. The data obtained during the execution of the physical simulation CutSim is virtual reaction force data representing the reaction force when the virtual robot 12S applies a force to the virtual object BS. By using this data, even when the task of applying a force to an object has irreversibility, the learned model can be efficiently generated.
[0022] This will be specifically described below. In the following, when not distinguishing between the real world and the virtual world, they are simply referred to as "robot", "object", and "knife".
[0023] <The Framework of this Embodiment> (A. Overview of the System) FIG. 3 is a diagram for explaining the outline of the control system 10 of the present embodiment (denoted as "Robot Operating System (ROS)" in FIG. 3). FIG. 3 shows a robot simulation RoboSim and a physical simulation CutSim. The robot simulation RoboSim sequentially outputs data representing the position x and velocity x of the virtual knife KS grasped by the virtual robot 12S to the physical simulation CutSim. Also, the physical simulation CutSim sequentially outputs the contact force F while the object corresponding to the virtual knife KS is cutting the virtual object BS to the robot simulation RoboSim. · ext
[0024] And, as shown in FIG. 3, the control system 10 efficiently generates a learned model used when performing a task of applying a force to an object by learning a machine learning model based on the data obtained while the robot simulation RoboSim and the physical simulation CutSim are being executed. Specifically, as shown in FIG. 3, the control system 10 generates a learned model, which is a control policy (denoted as "RL Controll Policy" in FIG. 3), using a known reinforcement learning algorithm based on the data representing the position x of the virtual knife KS output from the robot simulation RoboSim, the data representing the velocity x of the virtual knife KS, the data representing the jerk x of the virtual knife KS, and the data representing the contact force F between the virtual object BS and the virtual knife KS output from the physical simulation CutSim. · ··· ext
[0025] From the RL Controll Policy shown in FIG. 3, the target posture x of the real robot 12R and the external force F generated on the end effector 16R of the real robot 12R d d is output. Also, in this embodiment, as a known compliance controller (denoted as "Compliance Controller" in FIG. 3), FDCC (Forward Dynamics Compliance Controller) is used. As shown in FIG. 3, FDCC outputs data representing the position q of the joints of the virtual robot 12S or the real robot 12R c FDCC outputs data representing the position q of the joints of the virtual robot 12S or the real robot 12R. FDCC is disclosed, for example, in Reference 1 below.
[0026] Reference 1: S. Scherzinger, A. Roennau, and R. Dillmann, “Forward dynamics compliance control (FDCC): A new approach to cartesian compliance for robotic manipulators,” in IEEE / RSJ International Conference on Intelligent Robots and Systems, 2017, pp. 4568-4575.
[0027] This embodiment consists of two phases: reflection from reality to simulation and reflection from simulation to reality. Specifically, based on various data obtained when the real robot 12R cuts the real object BR with the real knife KR, the parameters of the robot simulation RoboSim and the physical simulation CutSim are adjusted by a known method.
[0028] More specifically, as shown in FIG. 3, for example, the contact force F between the real object BR and the real knife KR when cutting the real object BR ext and the data representing the position x of the real knife KR, the velocity x of the real knife KR · data representing the acceleration x of the real knife KR ·· data representing the acceleration x of the real knife KR, and the jerk x of the real knife KR ···The real robot 12R acquires data representing the robot's motion and the physical simulation parameters of the robot simulation RoboSim and the physical simulation parameters of the physical simulation CutSim.
[0029] Specifically, the contact force F when cutting the real object BR ext are collected by the real robot 12R and reflected in the virtual object BS in the physics simulation CutSim. As a result, the hardness of the real object BR and the frictional force that occurs when cutting the real object BR with the real knife KR are reproduced in the physics simulation CutSim.
[0030] In addition, data representing the position x of the real knife KR when the real robot 12R is operated, and the velocity x of the real knife KR · Data representing the acceleration x of the reality knife KR ·· Data representing the actual knife KR's jerk x ··· The data representing the above is collected and reflected in the virtual robot 12S of the robot simulation RoboSim. As a result, the operational conditions of the real robot 12R are reflected in the robot simulation RoboSim, and the real robot 12R is reproduced in the robot simulation RoboSim.
[0031] Thus, when the real object BR is cut by the real knife KR in the real world, the real object BR cannot be cut again by the real knife KR. Therefore, the control system 10 of this embodiment collects various data obtained when the real object BR is cut by the real knife KR, and reflects the collected data on each parameter of the physical simulation CutSim and each parameter of the robot simulation RoboSim. As a result, on the physical simulation CutSim and the robot simulation RoboSim, it becomes possible to reproduce the cutting of the real object BR that has already been cut in the real world, and it becomes possible to repeatedly obtain various data. Then, based on a plurality of data obtained on the physical simulation CutSim and the robot simulation RoboSim, a learned model used when the real robot 12R executes a task of applying a force to an object can be efficiently generated.
[0032] (B. Simulation Environment) [Type of Simulator] As the physical simulation CutSim and the robot simulation RoboSim, for example, it is possible to use known simulators. For example, as the physical simulation CutSim, it is possible to use the known DiSECt simulator. Also, for example, as the robot simulation RoboSim, it is possible to use the known Gazebo simulator. The DiSECt simulator and the Gazebo simulator are disclosed in the following References 2 and 3.
[0033] Reference 2: E. Heiden, M. Macklin, Y. Narang, D. Fox, A. Garg, and F. Ramos, “Disect: a differentiable simulator for parameter inference and control in robotic cutting,” Autonomous Robots, pp. 1-30, 2023. Reference 3: N. Koenig and A. Howard, “Design and use paradigms for gazebo, an open-source multi-robot simulator,” in IEEE / RSJ International Conference on Intelligent Robots and Systems, vol. 3, 2004, pp. 2149-2154.
[0034] [Calibration] As described above, in this embodiment, based on the data obtained in the real world, the parameters of the physical simulation CutSim and the parameters of the robot simulation RoboSim are adjusted. For example, using a known algorithm, the parameters of the physical simulation CutSim and the parameters of the robot simulation RoboSim are adjusted so that data similar to the data obtained in the real world is obtained. Thereby, the real world is reproduced on the physical simulation CutSim and the robot simulation RoboSim.
[0035] [Connection between Simulators] In this embodiment, it is assumed that the time step in the physical simulation CutSim and the time step in the robot simulation RoboSim are different. Therefore, in this embodiment, the time step of the physical simulation CutSim and the time step of the robot simulation RoboSim are synchronized. Details will be described later.
[0036] (C. Cutting Learning by Reinforcement Learning) [Markov Decision Process] In reinforcement learning, an agent acquires a desired behavior through interactions with the outside world. The task of cutting an object can be defined by the framework of a Markov decision process (MDP). When the task of cutting an object is defined by a Markov decision process, the task of cutting an object is defined by T time steps per episode. A Markov decision process is defined by a tuple (S, A, P, R, γ). S is the state space, A is the action space, and P:S × A is the transition probability function. R represents the reward function, and γ∈(0,1) represents the discount rate. At each time step t, the agent changes the current state s t Observe ∈S and take action a according to policy π t ∈A. State s t is the transition function p(s t+1 |s t ,a t ) according to s t+1 At this time, the agent takes action a t Reward by taking r t =R(s t+1 ,s t ,a t ) is received. Note that if the agent takes action a t By taking the state s t is the state s t+1 Transition to.
[0037] The ultimate goal in reinforcement learning is to find a policy π that maximizes the reward function R(t) given by the following equation: * The goal is to find the i below, which represents the time step.
[0038]
number
[0039] In this embodiment, Soft-Actor Critic (SAC), which is an example of a reinforcement learning algorithm, is adopted. The SAC algorithm is disclosed in, for example, Reference 4 below.
[0040] Reference 4: T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off policy maximum entropy deep reinforcement learning with a stochastic actor,” in International Conference on Machine Learning, 2018, pp.1861-1870.
[0041] The SAC algorithm is based on entropy maximization. Specifically, in the SAC algorithm, the expected reward E is calculated as follows: π The policy π that maximizes * The H in the following equation is the entropy, and in the SAC algorithm, the strategy π is chosen so that the entropy H is maximized. * is determined. Here, α is a ratio to a preset entropy. In addition, the SAC algorithm employs a known off-policy algorithm.
[0042]
number
[0043] [Compliance Control] As described above, the present embodiment employs a known FDCC, which is capable of performing control on Cartesian coordinates by combining impedance control, admittance control, and force control.
[0044] FIG. 4 is a diagram for explaining FDCC. d and external force F dIt is also a diagram showing how to control the end effector of the robot when a certain situation is given. The x shown in FIG. 4 is a parameter that includes the position of the end effector of the robot on the Cartesian coordinates and the position of the joints of the robot on the Cartesian coordinates. F ext is the measured contact force. x d is the target posture of the robot, and F d is the external force generated on the end effector of the robot. F c is a control command representing the force applied to the end effector of the robot. Also, H -1 J T is the forward dynamics model in the virtual model, and FK(q c ) is the forward kinematics model in the virtual model. q c corresponds to the position of the joints of the robot, and q c·· corresponds to the acceleration of the joints of the robot. Also, the net force F net includes all the forces generated from the force measured by the sensor and the virtual operation. Therefore, the compliance controller shown in FIG. 4 can be represented by the parameter K c representing stiffness, the parameter K p representing PD (Proportional and Differential) gain, and the parameter K d representing PD gain. As will be described later, in this embodiment, the learned model is used to obtain the parameter K c representing stiffness, the parameter K p representing PD gain, and the parameter K d representing PD gain. Note that hereinafter, the parameter K c representing stiffness, the parameter K p representing PD gain, and the parameter K d representing PD gain are also referred to as control parameters. Note that since FDCC and each parameter are disclosed in the above-mentioned Reference 1, for details, please refer to Reference 1.
[0045] [Reinforcement learning agent] FIG. 5 is a diagram for explaining the relationship between the reinforcement learning and the compliance controller in the present embodiment. As shown in FIG. 5, an actor, which is a reinforcement learning agent, outputs, as actions, the motion trajectory x of the robot c and the control parameters [K c , K p , K d to the compliance controller. Note that the actor outputs the motion trajectory x c and the control parameters [K c , K p , K d at a control cycle of 20 Hz, which is lower than the control cycle of 500 Hz by the compliance controller. Note that the next target posture x c is calculated by adding the motion trajectory x of the robot to the past target posture x d . Therefore, the sum of the current target posture x d and the motion trajectory x d becomes the next target posture x c . Then, an FDCC, which is a compliance controller, controls the joints of the robot based on the motion trajectory x d and the control parameters [K c , K c , K p , K d . The feedback from the robot is the posture of the end effector (or the posture of the knife) and the detected force and the detected moment in the robot. Note that sensors for detecting forces and torques at n positions are installed in the robot.
[0046] The physical quantities observed by the actor, which is an agent, are the relative position of the knife with respect to the target position, the velocity of the knife, the jerk of the knife, the history of the actions of the robot, the history of the forces and torques at n positions applied to the robot. Therefore, in the present embodiment, the observation data composed of these physical quantities is used as the state data s t . Note that the state data s tis a vector having these physical quantities as components. Each physical quantity can be obtained from various sensors equipped on the robot or past operation data of the robot.
[0047] Note that the actor in the SAC algorithm can be realized by, for example, a known temporal convolutional network which is an example of a machine learning model. The temporal convolutional network can process the history of the forces and torques at n locations applied to the robot, and the fully-connected network included in the temporal convolutional network can hold the observation history.
[0048] [Reward function] The following formula (1) is the reward function set in this embodiment. As shown in the following formula (1), the reward function is based on the height x of the knife from the cutting board cut and the contact force F ext and the jerk x of the knife ··· and is defined. Here, w1, w2, and w3 are preset weight parameters. Also, as shown in the following formula (1), the reward function incorporates a reward B according to the situation at the end of an episode representing a series of operations of the robot. Specifically, when the task of cutting an object with the knife is achieved (for example, when the knife reaches the cutting board with a force within a predetermined threshold), "100" is substituted as shown in the following formula into the reward B. Also, when the knife collides with the cutting board (for example, when the knife reaches the cutting board with a force greater than a predetermined threshold), "-100" is substituted into the reward B as shown in the following formula. Also, when the knife moves outside a preset space and when it takes too much time to complete the task, "-1" is substituted into the reward B as shown in the following formula. Note that, as shown in the following formula, the greater the jerk x of the knife ··· is, the smaller the reward r becomes. The jerk x of the knife ···Since this corresponds to the clattering movement of the knife, in this embodiment, a reward is set so as to suppress such movement.
[0049]
Number
[0050] <Control system 10> FIG. 6 is a block diagram showing a schematic configuration of the control system 10 of this embodiment. As shown in FIG. 6, the control system 10 includes a sensor group 11, a real robot 12R, and a control device 14. The control device 14 generates a learned model for controlling the operation of the end effector 16R of the real robot 12R. Further, the control device 14 controls the operation of the end effector 16R of the real robot 12R using the generated learned model.
[0051] The sensor group 11 is attached to the end effector 16R of the real robot 12R and sequentially measures the above-described observation data. The sensor group 11 measures the position of the real knife KR, the speed of the real knife KR, the jerk of the real knife KR, the action history of the real robot 12R, and the history of the forces and torques applied to n locations on the real robot 12R. For example, the forces and torques applied to n locations on the real robot 12R are so-called force sense values (for example, three-axis torque and three-axis force). Then, the sensor group 11 outputs the obtained observation data to the control device 14.
[0052] FIG. 7 is a block diagram showing the hardware configuration of the control device 14 according to this embodiment. As shown in FIG. 7, the control device 14 includes a CPU (Central Processing Unit) 42, a memory 44, a storage device 46, an input / output I / F (Interface) 48, a storage medium reader 50, and a communication I / F 52. Each component is connected to be communicable with each other via a bus 54.
[0053] The memory device 46 stores a learned model generation program and a control program for executing each of the processes described later. The CPU 42 is a central processing unit that executes various programs and controls each component. That is, the CPU 42 reads a program from the memory device 46 and executes the program using the memory 44 as a work area. The CPU 42 performs control of each of the above components and various arithmetic processes according to the program stored in the memory device 46.
[0054] The memory 44 is composed of a RAM (Random Access Memory) and temporarily stores programs and data as a work area. The memory device 46 is composed of a ROM (Read Only Memory), an HDD (Hard Disk Drive), an SSD (Solid State Drive), etc., and stores various programs including an operating system and various data.
[0055] The input / output I / F 48 is an interface that performs input of data from an external device and output of data to the external device. Also, for example, an input device for performing various inputs such as a keyboard and a mouse, and an output device for outputting various information such as a display and a printer may be connected. By adopting a touch panel display as the output device, it may function as an input device.
[0056] The storage medium reader 50 reads data stored in various storage media such as a CD (Compact Disc)-ROM, a DVD (Digital Versatile Disc)-ROM, a Blu-ray disc, a USB (Universal Serial Bus) memory, etc., and writes data to the storage medium.
[0057] The communication I / F 52 is an interface for communicating with other devices, and for example, standards such as Ethernet (registered trademark), FDDI, Wi-Fi (registered trademark) are used.
[0058] Next, the functional configuration of the control device 14 will be described. As shown in FIG. 6, the control device 14 functionally includes a learning acquisition unit 18, a simulation unit 20, a generation unit 22, an acquisition unit 24, and a control unit 26. Further, a data storage unit 28, a learned model storage unit 30, and a controller storage unit 32 are provided in a predetermined storage area of the control device 14. Each functional configuration is realized by the CPU 42 reading out each program stored in the storage device 46, expanding it in the memory 44, and executing it.
[0059] The data storage unit 28 stores the observation data detected by the sensor group 11. Further, the data storage unit 28 stores the control data when the end effector 16R of the real robot 12R operates. For example, the history of the action data when the end effector 16R of the real robot 12R operates is stored.
[0060] The learned model storage unit 30 stores the learned policy π * and the learned action value function Q * which are generated by the processes described later.
[0061] The controller storage unit 32 stores information representing the above-described FDCC model. Specifically, the control parameters K * output from the learned policy π c , K p , K d are substituted into the FDCC stored in the controller storage unit 32, and the position q c of the joint of the real robot 12R is output.
[0062] First, the learning acquisition unit 18, the simulation unit 20, and the generation unit 22 generate a learned policy π * which is a learned model for controlling the operation of the end effector 16R of the real robot 12R.
[0063] First, make the end effector 16R of the real robot 12R grip the real knife KR, and execute a task of cutting the real object BR with the real knife KR. Then, the learning acquisition unit 18 acquires real reaction force data representing the real reaction force obtained while the real knife KR of the real robot 12R is cutting the real object, and stores it in the data storage unit 28. This real reaction force data is measured by a force sensor or the like included in the sensor group 11.
[0064] Next, the simulation unit 20 reads out the real reaction force data stored in the data storage unit 28 and reflects the real reaction force data in the physical simulation. Specifically, the simulation unit 20 reflects the real reaction force data in the physical simulation by adjusting various parameters when executing the physical simulation.
[0065] Then, as shown in FIG. 2 above, the simulation unit 20 executes a physical simulation CutSim that simulates the reaction force from the virtual object BS when applying a force to the virtual object BS, and a robot simulation RoboSim that simulates the operation of the virtual robot 12S that processes the virtual object BS. Note that the simulation unit 20 executes the calculation of the physical simulation CutSim in the first cycle and executes the calculation of the robot simulation RoboSim in the second cycle. The first cycle and the second cycle represent the time steps of the simulation. For example, the first cycle dt1 = 1.0×10 -5 [s], and the second cycle dt2 = 1.0×10 -4 [s].
[0066] The generation unit 22 generates a learned model that outputs the action data of the real robot 12R when state data representing the environment in which the real robot 12R corresponding to the virtual robot 12S operates is input, based on the data obtained while the physical simulation CutSim and the robot simulation RoboSim are being executed by the simulation unit 20. Note that the data obtained while the physical simulation is being executed is virtual reaction force data representing the reaction force when the virtual robot 12S applies a force to the virtual object BS. The learned model of the present embodiment corresponds to the learned policy π * The learned policy π * outputs action data when state data is input. As described above, the state data of the present embodiment includes the position of the real knife KR, the speed of the real knife KR, the jerk of the real knife KR, the action data of the real robot 12R (the motion trajectory x c of the end effector 16R, the control parameters K c , K p , K d ), and the history of the forces and torques applied to n locations on the real robot 12R. Further, the action data of the present embodiment is the motion trajectory x c of the end effector 16R and the control parameters K c , K p , K d .
[0067] Specifically, the generation unit 22 uses the SAC algorithm, which is an example of the reinforcement learning algorithm as described above, to train an actor corresponding to the policy in reinforcement learning and a critic corresponding to the action value function in reinforcement learning. As shown in FIG. 5, when the state data s t corresponding to the observation data is input, the actor outputs the action data a t as the motion trajectory x c of the end effector 16R and the control parameters K c , K p , K d . Also, as shown in FIG. 5, the critic takes the state data s t corresponding to the observation data and the state data s tWhen the reward r calculated by the above formula (1) is input according to the situation, the value of the action value function is output. Note that the actor and the critic are realized by a known machine learning model.
[0068] The generation unit 22 causes the actor and the critic to learn so that the total sum of the rewards r calculated by the above formula (1) becomes large according to the SAC algorithm, thereby obtaining the learned policy π corresponding to the learned actor * and the learned action value function Q corresponding to the learned critic * are generated. Then, the generation unit 22 stores the learned policy π * and the learned action value function Q * in the learned model storage unit 30. Note that the reward r is preset according to the type of task executed by the virtual robot 12S.
[0069] Note that when generating the learned policy π * , the generation unit 22 obtains the virtual reaction force data corresponding to the second cycle dt2 in the robot simulation RoboSim from the virtual reaction force data obtained from the physical simulation CutSim by the calculation using the first cycle dt1, and based on the virtual reaction force data corresponding to the second cycle dt2, generates the learned policy π * . For example, the generation unit 22 obtains the virtual reaction force data for the second cycle dt2 from the virtual reaction force data for each first cycle dt1 obtained by the calculation of the physical simulation CutSim, and uses the obtained virtual reaction force data for the second cycle dt2 to generate the learned policy π * . Thereby, even if the calculation cycle of the physical simulation CutSim and the calculation cycle of the robot simulation RoboSim are different, the learned policy π * can be generated.
[0070] Also, when generating the learned policy π * , if the fluctuation of the virtual reaction force data is within a predetermined range, the virtual reaction force data is obtained at a cycle larger than the first cycle dt1, and based on the obtained virtual reaction force data, the learned policy π* may be generated. For example, the variation in the reaction force data during cutting an object with a knife may be small. Even if the virtual reaction force data is acquired in a short time interval in such a case, since the virtual reaction force data acquired in that time interval is almost constant, the learned policy π * may not be useful when generating. Therefore, when the variation in the virtual reaction force data is within a predetermined range, the virtual reaction force data may be acquired at a period larger than the first period dt1 (for example, a period twice the first period dt1, etc.).
[0071] Also, when the generation unit 22 generates the learned policy π * if the variation in the virtual reaction force data is outside the predetermined range, the virtual reaction force data is acquired at a period smaller than the first period dt1, and based on the acquired virtual reaction force data, the learned policy π * may be generated. For example, when the variation in the reaction force data during cutting an object with a knife is large, it may be better to generate the learned policy π * using the detailed data at that time. Therefore, when the variation in the virtual reaction force data is outside the predetermined range, the virtual reaction force data may be acquired at a period smaller than the first period dt1 (for example, a period half of the first period dt1, etc.).
[0072] When the learned policy π * is stored in the learned model storage unit 30, it becomes possible to control the operation of the end effector 16R of the real robot 12R using the learned policy π * . Therefore, the acquisition unit 24 and the control unit 26 use the learned policy π * stored in the learned model storage unit 30 and the FDCC stored in the controller storage unit 32 to control the operation of the end effector 16R of the real robot 12R.
[0073] The acquisition unit 24 acquires the state data s t representing the environment in which the real robot 12R operates.
[0074] The control unit 26 reads out the learned policy π stored in the learned model storage unit 30. * Then, the control unit 26 inputs the state data s acquired by the acquisition unit 24 t into the learned policy π * to obtain the action data a of the real robot 12R. As described above, from the learned policy π t , the operation trajectory x of the end effector 16R * and the control parameters K c ,K c ,K p ,K d are output as the action data a t .
[0075] The control unit 26 inputs the control parameters K * output from the learned policy π c ,K p ,K d into the FDCC stored in the controller storage unit 32 to obtain the joint position q of the real robot 12R output from the FDCC. c
[0076] Then, the control unit 26 controls the real robot 12R based on the joint position q of the real robot 12R output from the FDCC c and the operation trajectory x of the end effector 16R. Specifically, the control unit 26 outputs a control signal to the real robot 12R such that the joint position q of the real robot 12R output from the FDCC c and the operation trajectory x of the end effector 16R c are realized. c
[0077] Next, the operation of the control system 10 according to the present embodiment will be described.
[0078] First, a task is executed in which the real robot 12R cuts the real object BR using the real knife KR. At this time, the real reaction force data is input to the control device 14 and stored in the data storage unit 28. Then, when the control device 14 receives a predetermined instruction signal, the CPU 42 of the control device 14 reads the learned model generation program from the storage device 46, expands it in the memory 44, and executes it. As a result, the CPU 42 functions as each functional configuration of the control device 14, and the learned model generation process shown in FIG. 8 is executed.
[0079] In step S100, the simulation unit 20 acquires the real reaction force data stored in the data storage unit 28.
[0080] In step S102, the simulation unit 20 reflects the real reaction force data acquired in step S100 in the physical simulation CutSim, thereby setting various parameters of the physical simulation CutSim.
[0081] In step S104, the simulation unit 20 starts the execution of the physical simulation CutSim and the robot simulation RoboSim.
[0082] In step S106, based on the data obtained while the physical simulation CutSim and the robot simulation RoboSim are being executed, the generation unit 22 generates the learned policy π t such that when the state data s t representing the environment in which the virtual robot 12S operates is input, the action data a * of the virtual robot 12S is output. Specifically, the generation unit 22 learns the actor and the critic according to the SAC algorithm so that the reward r calculated by the above formula (1) is maximized, thereby obtaining the learned policy π * corresponding to the learned actor and the learned action value function Q * corresponding to the learned critic.
[0083] In step S108, the generation unit 22 stores the learned policy π * and the learned action value function Q * in the learned model storage unit 30.
[0084] Next, when the control device 14 receives a predetermined instruction signal, the control device 14 executes the control process shown in FIG. 9. The control process in FIG. 9 is repeatedly executed.
[0085] In step S200, the acquisition unit 24 acquires state data s t representing the environment in which the real robot 12R operates.
[0086] In step S202, the control unit 26 reads out the learned policy π * stored in the learned model storage unit 30. Then, the control unit 26 inputs the state data s t acquired in step S200 into the learned policy π * to acquire the action data a t of the real robot 12R.
[0087] In step S204, the control unit 26 inputs the control parameter K t among the action data a c , K p , K d acquired in step S202 into the FDCC stored in the controller storage unit 32, and acquires the joint position q c of the real robot 12R output from the FDCC.
[0088] In step S206, the control unit 26 controls the real robot 12R based on the joint position q c of the real robot 12R acquired in step S204 and the operation trajectory x t of the end effector 16R among the action data a c acquired in step S202.
[0089] As described above, the control device according to the present embodiment executes physical simulation for simulating the reaction force from a virtual object when applying a force to the virtual object, and robot simulation for simulating the operation of a virtual robot that processes the virtual object. Then, based on the data obtained while the physical simulation and the robot simulation are being executed, when state data representing the environment in which the real robot corresponding to the virtual robot operates is input, the control device generates a learned policy for outputting the action data of the real robot. Thereby, a learned model used when the robot executes a task of applying a force to an object can be efficiently generated.
[0090] Further, the control device according to the present embodiment acquires state data representing the environment in which the real robot operates, inputs the acquired state data to the generated learned policy, acquires the action data of the real robot, and controls the real robot based on the action data. Thereby, the robot can execute a task of applying a force to an object.
[0091] Moreover, according to the control device according to the present embodiment, even when the task of processing an object has irreversibility, by using physical simulation and robot simulation, it is possible to repeatedly acquire data. Therefore, it is possible to acquire more data when processing without actually processing the real object, and a learned model can be efficiently generated.
Example
[0092] Next, an example will be described. In this example, a simulation for verifying the effectiveness of the proposed method is performed. In this simulation, an experiment regarding the task of slicing a real object was conducted. The real objects targeted in the experiment were cucumber, tomato, potato, and carrot. Also, the number of slices and the size of the slices were set as shown in the following table.
[0093]
Table 1
[0094] Figure 10 is a diagram showing the results when calibration is performed by reflecting real data into a physical simulation. Figure 10(A) shows the results when the real object is a cucumber, Figure 10(B) shows the results when the real object is a tomato, and Figure 10(C) shows the results when the real object is a potato. The horizontal axis of the graph shown in Figure 10 is time, and the vertical axis is the contact force. "Disect" shown in Figure 10 represents the result of the physical simulation reflected with real data, "Gazebo" represents the result of the robot simulation, and "Ground Truth" represents the real contact force data. As shown in Figure 10, it can be seen that due to the calibration being performed on the physical simulation, the contact force in "Disect" is similar to the contact force in "Ground Truth".
[0095] Figure 11 is a diagram showing the results of the variation in the contact force when a learned model is generated only with "Gazebo" which is a robot simulation (denoted as "Gazebo only" in Figure 11), and when a learned model is generated by combining the physical simulation "Disect" and the robot simulation "Gazebo" (denoted as "Disect+Gazebo" in Figure 11). As shown in Figure 11, it can be seen that "Disect+Gazebo" has less variation in the contact force than "Gazebo only", and can slice the object more smoothly.
[0096] Figure 12 is a diagram showing the time series of the contact force when a cucumber is sliced twice. As shown in Figure 12, it can be seen that "Disect+Gazebo" requires less contact force than "Gazebo only" and can slice the cucumber more smoothly with a smaller force.
[0097] As can be seen from FIGS. 11 and 12, it can be understood that by using the method of this embodiment, it is possible to accurately and efficiently generate a learned strategy used when applying a force to an object.
[0098] In the above embodiment, the case where the task of applying a force to an object is the task of cutting the object has been described as an example, but it is not limited to this. Any task that applies a force to an object may be used. For example, this embodiment can also be applied to tasks such as bending parts, cutting parts, food processing (e.g., bread molding, etc.), and gripping soft parts. Further, this embodiment can also be applied to tasks that do not have irreversibility. For example, this embodiment can also be applied to tasks that do not have irreversibility and take time to set. For example, this embodiment can also be applied to tasks such as cross-drawing of a table, food processing (e.g., bread molding, etc.), and assembly that requires multiple steps (e.g., an assembly task that must be disassembled and returned to the original state if a failure occurs during the process).
[0099] In the above embodiment, the case where a reinforcement learning algorithm is used as the machine learning algorithm has been described as an example, but it is not limited to this. Other machine learning algorithms (e.g., supervised learning algorithm or unsupervised learning algorithm) may be used to generate the learned model.
[0100] In the above embodiment, a part of the action data output from the learned model is input to the FDCC, and the case where the robot is controlled based on the joint position q c output from the FDCC has been described as an example, but it is not limited to this. Different types of data from the above embodiment may be used as the action data. In this case, it is also possible to control the robot according to the action data without using the FDCC. Also, the state data is not limited to the above embodiment, and any data that represents the environment of the robot may be used.
[0101] Also, in the above embodiment, the case where the reward r is defined by Equation (1) has been described as an example, but it is not limited thereto. For example, the reward r can be appropriately changed according to the task executed by the robot.
[0102] Also, in the above embodiment, the case where the first cycle dt1, which is the time step of the physical simulation CutSim, and the second cycle dt2, which is the time step of the robot simulation RoboSim, are different has been described as an example, but it is not limited thereto. For example, the first cycle dt1 and the second cycle dt2 may be made the same.
[0103] Also, in the above embodiment, the case where the control device 14 executes both the learned model generation process of FIG. 8 and the control process of FIG. 9 has been described as an example, but it is not limited thereto. For example, a learned model generation device realized by a computer different from the control device 14 may be prepared, the learned model generation device may execute the learned model generation process of FIG. 9, and the control device 14 may execute the control process of FIG. 10. In this case, the learned model generation device includes at least the above-described learning acquisition unit 18, simulation unit 20, and generation unit 22.
[0104] Also, in the above-described embodiment, each process executed by the CPU by loading and executing software (program) may be executed by various processors other than the CPU. Examples of the processor in this case include a PLD (Programmable Logic Device) whose circuit configuration can be changed after manufacturing, such as an FPGA (Field-Programmable Gate Array), and a dedicated electric circuit, which is a processor having a circuit configuration designed specifically for executing specific processes, such as an ASIC (Application Specific Integrated Circuit). Further, each process may be executed by one of these various processors, or may be executed by a combination of two or more processors of the same type or different types (for example, a plurality of FPGAs, a combination of a CPU and an FPGA, etc.). More specifically, the hardware structure of these various processors is an electric circuit formed by combining circuit elements such as semiconductor elements.
[0105] Also, in the above-described embodiment, an aspect in which each program is pre-stored (installed) in the storage device has been described, but the present invention is not limited to this. The program may be provided in a form stored in a storage medium such as a CD-ROM, DVD-ROM, Blu-ray Disc, USB memory, etc. Further, the program may be in a form downloaded from an external device via a network.
[0106] (Supplementary Note) Hereinafter, aspects of the present disclosure will be appended.
[0107] (Supplementary Note 1) A simulation unit that executes a physical simulation for simulating a reaction force from the virtual object when a force is applied to the virtual object, and a robot simulation for simulating the operation of a virtual robot that applies a force to the virtual object. A generation unit that generates a learned model that outputs the action data of the real robot when state data representing the environment in which the real robot corresponding to the virtual robot operates is input based on the data obtained while the physical simulation and the robot simulation are being executed; A learned model generation device including the above. (Appendix 2) The data obtained while the physical simulation is being executed is virtual reaction force data representing the reaction force when the virtual robot applies a force to the virtual object, The generation unit generates the learned model based on the virtual reaction force data. The learned model generation device according to Appendix 1. (Appendix 3) The simulation unit executes the calculation of the physical simulation in the first cycle and executes the calculation of the robot simulation in the second cycle. When generating the learned model, the generation unit acquires the virtual reaction force data corresponding to the second cycle in the robot simulation from the virtual reaction force data obtained from the physical simulation by the calculation in the first cycle, and generates the learned model based on the virtual reaction force data corresponding to the second cycle. The learned model generation device according to Appendix 2. (Appendix 4) When generating the learned model, if the variation of the virtual reaction force data is within a predetermined range, the generation unit acquires the virtual reaction force data at a cycle larger than the second cycle, and generates the learned model based on the acquired virtual reaction force data. The learned model generation device according to Appendix 3. (Appendix 5) When generating the learned model, if the variation of the virtual reaction force data is outside a predetermined range, the generation unit acquires the virtual reaction force data at a cycle smaller than the second cycle, and generates the learned model based on the acquired virtual reaction force data. The learned model generation device according to Appendix 3 or Appendix 4. (Appendix 6) The simulation unit acquires real reaction force data representing a real reaction force when a real robot corresponding to the virtual robot applies a force to a real object corresponding to the virtual object, and executes the physical simulation by reflecting the real reaction force data in the physical simulation. The learned model generation device according to any one of Appendices 1 to 5. (Appendix 7) The generation unit generates the learned model by reinforcement learning so that the total reward preset according to the type of task executed by the virtual robot increases. The learned model generation device according to any one of Appendices 1 to 6. (Appendix 8) The generation unit generates the learned model by reinforcement learning so that the reward according to the situation at the end of an episode representing a series of operations of the virtual robot increases. The learned model generation device according to any one of Appendices 1 to 7. (Appendix 9) The generation unit when performing reinforcement learning according to the Soft Actor-Critic algorithm, learns an actor representing a policy in reinforcement learning and a critic representing an action value function in reinforcement learning based on data obtained while the physical simulation and the robot simulation are being executed, and generates a learned policy as the learned model. The learned model generation device according to any one of Appendices 1 to 8. (Appendix 10) an acquisition unit that acquires the state data, for the learned model generated by the learned model generation device according to any one of Appendices 1 to 9, by inputting the state data acquired by the acquisition unit, acquires the action data of the real robot, and controls the real robot based on the action data. A control device provided with (Appendix 11) Execute a physical simulation that simulates the reaction force when a force is applied to a virtual object, and a robot simulation that simulates the operation of a virtual robot that applies a force to the virtual object, Based on the data obtained while the physical simulation and the robot simulation are being executed, generate a learned model that outputs the action data of the real robot when state data representing the environment in which the real robot corresponding to the virtual robot operates is input. A method for generating a learned model in which a computer executes processing. (Appendix 12) Execute a physical simulation that simulates the reaction force when a force is applied to a virtual object, and a robot simulation that simulates the operation of a virtual robot that applies a force to the virtual object, Based on the data obtained while the physical simulation and the robot simulation are being executed, generate a learned model that outputs the action data of the real robot when state data representing the environment in which the real robot corresponding to the virtual robot operates is input. A learned model generation program for causing a computer to execute processing.
Explanation of Signs
[0108] 10 Control system 11 Sensor group 12R Real robot 12S Virtual robot 14 Control device 16R End effector 18 Acquisition unit for learning 20 Simulation unit 22 Generation unit 24 Acquisition unit 26 Control unit 28 Data storage unit 30 Learned model storage unit 32 Controller storage unit
Claims
1. A simulation unit that executes a physical simulation for simulating a reaction force from the virtual object when a force is applied to the virtual object, and a robot simulation for simulating the operation of a virtual robot that applies a force to the virtual object, and a generation unit that generates a learned model that outputs the action data of the real robot when state data representing the environment in which the real robot corresponding to the virtual robot operates is input based on the data obtained while the physical simulation and the robot simulation are being executed. A learned model generation device including the above.
2. The data obtained while the physical simulation is being executed is virtual reaction force data representing the reaction force when the virtual robot applies a force to the virtual object. The generation unit generates the learned model based on the virtual reaction force data. The learned model generation device according to Claim 1.
3. The simulation unit executes the calculation of the physical simulation in a first cycle and executes the calculation of the robot simulation in a second cycle. When generating the learned model, the generation unit acquires the virtual reaction force data corresponding to the second cycle in the robot simulation from the virtual reaction force data obtained from the physical simulation by the calculation in the first cycle, and generates the learned model based on the virtual reaction force data corresponding to the second cycle. The learned model generation device according to Claim 2.
4. When generating the learned model, if the variation of the virtual reaction force data is within a predetermined range, the generation unit acquires the virtual reaction force data at a cycle larger than the second cycle, and generates the learned model based on the acquired virtual reaction force data. The learned model generation device according to Claim 3.
5. When generating the learned model, if the variation of the virtual reaction force data is outside a predetermined range, the generation unit acquires the virtual reaction force data at a cycle smaller than the second cycle, and generates the learned model based on the acquired virtual reaction force data. The learned model generation device according to Claim 3 or Claim 4.
6. The simulation unit Obtain real reaction force data representing the real reaction force when a real robot corresponding to the virtual robot applies a force to a real object corresponding to the virtual object. Execute the physical simulation by reflecting the real reaction force data in the physical simulation. The learned model generation device according to claim 1 or claim 2.
7. The generation unit generates the learned model by reinforcement learning so that the total reward preset according to the type of task executed by the virtual robot increases. The learned model generation device according to claim 1 or claim 2.
8. The generation unit generates the learned model by reinforcement learning so that the reward according to the situation at the end of an episode representing a series of operations of the virtual robot increases. The learned model generation device according to claim 1 or claim 2.
9. The generation unit When executing reinforcement learning according to the Soft Actor-Critic algorithm, Based on the data obtained during the execution of the physical simulation and the robot simulation, learn an actor representing a policy in reinforcement learning and a critic representing an action value function in reinforcement learning, and generate the learned policy as the learned model. The learned model generation device according to claim 1 or claim 2.
10. An acquisition unit that acquires the state data, For the learned model generated by the learned model generation device according to claim 1 or claim 2, input the state data acquired by the acquisition unit to acquire the action data of the real robot, and control the real robot based on the action data. A control unit, A control device comprising.
11. Execute a physical simulation that simulates the reaction force when a force is applied to a virtual object and a robot simulation that simulates the operation of a virtual robot that applies a force to the virtual object. Generate a learned model that outputs the action data of the real robot when state data representing the environment in which the real robot corresponding to the virtual robot operates is input based on the data obtained during the execution of the physical simulation and the robot simulation. A learned model generation method in which a computer executes a process.
12. Execute a physical simulation that simulates the reaction force when a force is applied to a virtual object, and a robot simulation that simulates the operation of a virtual robot that applies a force to the virtual object. Generate a learned model that outputs the action data of the real robot when state data representing the environment in which the real robot corresponding to the virtual robot operates is input based on the data obtained while the physical simulation and the robot simulation are being executed. A learned model generation program for causing a computer to execute the process.
Citation Information
Patent Citations
Information processing method, image processing method, robot control method, product manufacturing method, information processing apparatus, image processing apparatus, robot system, program and recording medium
JP2023168240A