Robot, motion generating device, robot control system, and motion generating method
The robot system addresses the challenge of accurate imitation movement by using state detection and optimization units to adjust control commands, ensuring precise and stable task execution.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2026-03-12
AI Technical Summary
Conventional robot systems require extensive programming and specialized knowledge, hindering their widespread adoption, and autonomous learning robots face challenges in accurately following imitation movements due to difficulties in considering control characteristics and constraints, leading to potential task failure or instability.
A robot system with a driving unit, state detection, imitation movement generation, state prediction, predicted trajectory evaluation, and control command optimization units that adjust control commands based on agreement between target and predicted trajectories to ensure accurate imitation movements.
Enables robots to accurately follow imitation movements by correcting control commands to align with environmental and operational constraints, enhancing task performance and stability.
Smart Images

Figure JP2025027806_12032026_PF_FP_ABST
Abstract
Description
Robot, motion generation device, robot control system, and motion generation method
[0001] The present invention relates to a robot, a motion generation device, a robot control system, and a motion generation method.
[0002] Conventional robot systems require extensive programming and a high level of specialized knowledge, which has hindered the introduction of robots. To address this issue, autonomous learning robots have been proposed, in which the robot itself determines its own behavior based on information from various sensors attached to the robot.
[0003] This autonomous learning robot device is able to generate flexible movements in response to a variety of environmental changes by memorizing and imitating movements mainly taught by humans or other robots. Therefore, it is expected that autonomous learning robot devices will enable robots to take over complex tasks that have traditionally been performed by humans.
[0004] Generally, an autonomous learning robot device is equipped with an imitation learner. The imitation learner stores sensor information from motion experience and adjusts parameters to generate motion. The imitation learner also has a predefined input / output relationship. The imitation learner then repeatedly learns so that the expected output value for the input value to the imitation learner is output as an imitation motion.
[0005] For example, the imitation learner stores the joint angle information of the robot when experiencing a certain movement as time-series information. Assume that joint angle information at time (t) is input to the imitation learner, and the imitation learner performs time-series learning to predict the joint angle information at the next time (t+1).
[0006] Then, by sequentially inputting the robot's joint angle information into the imitation learning device that has completed learning, the autonomous learning robot device becomes able to autonomously generate imitation movements in response to changes in the environment and its own state.
[0007] In this way, by adjusting the parameters of the imitation learner so that the autonomous learning robotic device imitates the taught motion, it becomes possible to imitate the motions of humans or other robots in a short learning time. However, when generating teaching data from humans or other robots, it is difficult to take into account the control characteristics and constraints of controlling the autonomous learning robotic device. Therefore, even if it is possible to generate motions that imitate the taught motions, the autonomous learning robotic device may not be able to adequately follow the generated motions. This may result in the autonomous learning robotic device failing to perform a target task or becoming unstable during the task.
[0008] As a means for adjusting the parameters of an imitation learner while taking into consideration the control characteristics and constraints of a robot, for example, Patent Document 1 discloses a technology such as that described in paragraph 0007. Patent Document 1 discloses a reinforcement learning method, a reinforcement learning program, and a reinforcement learning device, which, according to one embodiment, predict the state of a control target in reinforcement learning at each time point at which state measurements of the target are made, the time point being included in a period after a current action decision is made and before a next action decision is made, while the time interval at which state measurements of the target are made is different from the time interval at which an action decision is made on the target, calculate a risk level for the state of the target at each time point with respect to a constraint on the state of the target based on the predicted state of the target, identify a search range for a current action on the target based on the calculated risk level for the state of the target at each time point and the impact of the current action on the target on the state of the target at each time point, and determine a current action on the target based on the identified search range for the current action on the target (see paragraph 0007).
[0009] Japanese Patent Application Laid-Open No. 2021-033767
[0010] As in Patent Document 1, reinforcement learning makes it possible to take into account the control characteristics and constraints of a robot. However, reinforcement learning requires a high learning cost to acquire a desired behavior. In other words, reinforcement learning is a method of acquiring behavior through trial and error, so it requires the construction of an elaborate simulation environment to evaluate the behavior obtained during the parameter adjustment process, or a real environment carefully designed to create a desirable environment for the behavior. In addition, unlike imitation learning, reinforcement learning requires starting to learn a behavior from scratch without any example behaviors generated by humans or other robots. Therefore, reinforcement learning generally tends to require more time and effort than imitation learning.
[0011] Furthermore, in reinforcement learning, if control characteristics and constraints are taken into account during prior learning, it is difficult to consider control errors and disturbances that occur when the robot actually operates. The simulation and real-world environments used to learn behaviors often differ from the environments in which the robot actually performs tasks. Therefore, it is desirable to be able to correct the behavior generated by the imitation learner by taking into account the robot's control characteristics and constraints according to the situation when the robot actually operates.
[0012] The present invention has been made in view of the above background, and an object of the present invention is to enable a robot to accurately follow an imitation movement in imitation learning.
[0013] In order to solve the above-mentioned problems, the present invention includes a driving unit that drives a robot, a state detection unit that acquires information about the state of the robot, an imitation movement generation unit that generates a target imitation trajectory, which is time-series data of target values related to the movement of the robot, by imitation learning based on the state detection information acquired from the state detection unit, a state prediction unit that generates a predicted state trajectory, which is time-series data related to a predicted movement of the robot based on the movement characteristics of the driving unit, using a control command to the driving unit at a current time and the state detection information, a predicted state trajectory evaluation unit that compares the target imitation trajectory generated by the imitation movement generation unit with the predicted state trajectory generated by the state prediction unit and evaluates the degree of agreement between the target imitation trajectory and the predicted state trajectory, and a control command optimization unit that corrects the control command based on the degree of agreement evaluated by the predicted state trajectory evaluation unit and outputs the corrected control command to the driving unit. Other solutions will be described appropriately in the embodiments.
[0014] According to the present invention, in imitation learning, it is possible to make a robot accurately follow an imitation movement.
[0015] 1 is a diagram illustrating a configuration of a robot according to a first embodiment; FIG. 2 is a diagram illustrating an example of the hardware configuration of a motion generation device or a robot; FIG. 3 is a diagram illustrating an example of a task performed by a robot; FIG. 4 is a flowchart illustrating the procedure of a motion generation method performed by a control unit according to the first embodiment; FIG. 5 is a diagram illustrating an example of a target imitation trajectory that is generated; FIG. 6 is a flowchart illustrating the procedure of a motion generation method performed by a control unit according to the first embodiment; FIG. 7 is a diagram illustrating an example of an evaluation method (part 1); FIG. 8 is a diagram illustrating an example of an evaluation method (part 2); FIG. 9 is a diagram illustrating an effect of the first embodiment; FIG. 10 is a diagram illustrating an effect of the first embodiment; FIG. 11 is a diagram illustrating an effect of the second embodiment; FIG. 12 is a diagram illustrating an effect of the second embodiment; FIG. 13 is a diagram illustrating an effect of the second embodiment; FIG. 14 is a diagram illustrating an effect of the second embodiment; FIG. 15 is a diagram illustrating an effect of the second embodiment; FIG. 16 is a diagram illustrating an effect of the second embodiment; FIG. 17 is a diagram illustrating an effect of the second embodiment; FIG. 18 is a diagram illustrating an effect of the second embodiment; FIG. 19 is a diagram illustrating an effect of the second embodiment; FIG. 10 is a diagram (part 2) showing an example of a task execution environment for a robot according to the third embodiment and an evaluation method for obstacle avoidance. FIG. 11 is a diagram (part 3) showing an example of a task execution environment for a robot according to the third embodiment and an evaluation method for obstacle avoidance. FIG. 12 is a diagram (part 1) showing the effect of the method shown in the third embodiment. FIG. 13 is a diagram (part 3) showing the effect of the method shown in the third embodiment. FIG. 14 is a diagram (part 4) showing the effect of the method shown in the third embodiment. A diagram showing an example of a candidate target imitation trajectory and an evaluation method by a predicted state trajectory evaluation unit in the fourth embodiment. A diagram showing an example of a robot control system according to the present embodiment.
[0016] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. The following description shows specific examples of the contents of the present invention, and the present invention is not limited to these descriptions. Various changes and modifications are possible by those skilled in the art within the scope of the technical ideas disclosed in this specification. Furthermore, in all drawings used to explain the present invention, parts having the same function are given the same reference numerals, and repeated explanations thereof may be omitted.
[0017] In this embodiment, the term "trajectory" refers to a time series of data.
[0018] First Embodiment First, a first embodiment of the present invention will be described with reference to FIGS. 1 to 8D.
[0019] (Configuration of Robot 1) FIG. 1 is a diagram showing the configuration of a robot 1 according to the first embodiment.
[0020] The robot 1 includes a motion generating device 100 that generates information for the robot 1 to operate, a state detection unit 130, an environment detection unit 140, and a drive unit 150.
[0021] The state detection unit 130, which acquires information about the state of the robot 1, includes an angle sensor 131, a torque sensor 132, etc. In this way, the state detection unit 130 is made up of sensors that measure the internal state of the robot 1.
[0022] Furthermore, the environment detection unit 140, which acquires environmental information around the robot 1, includes a camera 141 and the like for measuring external information around the robot 1. State detection information 161 acquired by the angle sensor 131 and torque sensor 132 is used to generate a target imitation trajectory 163, which is information for the learned imitation movement of the robot 1, and to calculate a predicted state trajectory 164. Note that the state detection unit 130 is not limited to the angle sensor 131 and torque sensor 132. For example, a pressure sensor for measuring pressure information acting on a part of the robot 1, or an IMU for measuring the tilt, posture, etc. of the robot 1 may be used as an internal sensor. Note that IMU is an abbreviation for inertial measurement unit.
[0023] The images captured by the camera 141 are used to measure environmental information such as the target of the task that the robot 1 is attempting to accomplish and the surrounding environment. Note that the sensor for measuring the environmental information around the robot 1 is not limited to the camera 141. For example, LiDAR for measuring surrounding point cloud information, a microphone for measuring sound information generated in the surroundings, or the like may be used as a sensor for measuring external world information around the robot 1. Note that LiDAR is an abbreviation for Light Detection and Ranging.
[0024] The driving unit 150 that drives the robot 1 is composed of actuators 151 and the like. The actuators 151 are a plurality of rotary motors, translational motors, and the like that operate the robot 1. The actuators 151 are controlled by control commands 165 output by the motion generation device 100, and are driven by a motor driver (not shown).
[0025] In this embodiment, the configuration of the robot 1 is assumed to be an arm-type robot with multiple joints, a mobile robot with multiple wheels, or the like.
[0026] The motion generation device 100 includes an imitation motion generation unit 110 and a control unit 120. The imitation motion generation unit 110 receives state detection information 161 from the angle sensor 131 and torque sensor 132 and environment detection information 162 from the camera 141 (information on images captured by the camera 141) at each time, and generates a target imitation trajectory 163 to be imitated by the robot 1. The state detection information 161 includes angle information, torque information, etc., and the environment detection information 162 includes image information.
[0027] The imitation motion generation unit 110 includes a memory motion inference unit 111 and an imitation motion linking unit 112. The imitation motion generation unit 110 generates a target imitation trajectory 163, which is time-series data of target values 300 (see FIG. 5 ) related to a series of motions of the robot 1, by imitation learning based on state detection information 161 acquired from the state detection unit 130. Imitation learning is machine learning that generates target values 300 for the motions of the robot 1 based on teaching data 166. The teaching data 166 is data collected by a human operating a robot 1 other than the robot 1 controlled by the motion generation device 100. The imitation motion generation unit 110 can also generate the target imitation trajectory 163 using the state detection information 161 and environment detection information 162 acquired from the environment detection unit 140. In this embodiment, the target imitation trajectory 163 is generated using the state detection information 161 and the environment detection information 162. However, the target imitation trajectory 163 may be generated using only the state detection information 161.
[0028] The teaching data 166 is generated by the following procedure or the like. For example, a person operates the robot 1 via a manual controller (not shown) or the like. Then, the operation information is used as the teaching data 166. Alternatively, the robot 1 is operated so as to trace the human's movement, and the traced trajectory is used as the teaching data 166. When operating the robot 1 so as to trace the human's movement, a device for recording the human's movement or the like may be used, or the human's movement may be recorded by manually moving the hand by grasping the hand of the robot 1, for example. The teaching data 166 is composed of information on the angles of the joints of the robot 1 (angle information), information on the torque of the actuator 151 (torque information), image information captured by the camera 141, and the like.
[0029] The teaching data 166 is selected (created) appropriately depending on the nature of the task that the robot 1 is to execute.
[0030] The memory and operation inference unit 111 receives the state detection information 161 as input and calculates a target value for the next movement of the robot 1. The memory and operation inference unit 111 uses teaching data 166 to adjust learning parameters for imitation learning to calculate the target imitation trajectory 163 in advance. The teaching data 166 is not used when the memory and operation inference unit 111 generates the target imitation trajectory 163. In FIG. 1 , the arrow from the teaching data 166 to the memory and operation inference unit 111 is a dashed arrow, which indicates that the teaching data 166 is used only to adjust learning parameters for learning to calculate the target imitation trajectory 163, and is not used when generating the target imitation trajectory 163. Note that the imitation learner described above corresponds to the memory and operation inference unit 111.
[0031] The imitation movement linking unit 112 generates a target imitation trajectory 163 by linking together the plurality of target values 300 calculated by the memory movement inference unit 111 .
[0032] The control unit 120 receives the information measured by the angle sensor 131 and torque sensor 132 at each time and the target imitation trajectory 163 output by the imitation movement generation unit 110 as input, calculates a control command 165 , and outputs it to the actuator 151 .
[0033] The control unit 120 includes a state prediction unit 121 , a predicted state trajectory evaluation unit 122 , and a control command optimization unit 123 .
[0034] The state prediction unit 121 uses the current control command 165 to the drive unit 150 and the state detection information 161 to generate a predicted state trajectory 164, which is time series data regarding a series of predicted operations of the robot 1 based on the operating characteristics of the drive unit 150.
[0035] The predicted state trajectory evaluation unit 122 compares the target imitation trajectory 163 generated by the imitation action generation unit 110 with the predicted state trajectory 164 generated by the state prediction unit 121, and evaluates the degree of agreement between the target imitation trajectory 163 and the predicted state trajectory 164.
[0036] The control command optimization unit 123 corrects the control command 165 by performing mathematical optimization processing based on the degree of coincidence evaluated by the predicted state trajectory evaluation unit 122. Then, the control command optimization unit 123 outputs the corrected control command 165 to the drive unit 150.
[0037] The predicted state trajectory 164 is configured to include information on a predicted angle (predicted angle) and information on a predicted torque (predicted angle information and predicted torque information, respectively). The predicted state trajectory 164 is also configured to include a conversion formula or conversion map for converting a current value or torque to be input to the actuator 151 and a current value to be input to the actuator 151. Note that the predicted state trajectory 164 is not limited to the above content as long as it can be compared with the target simulated trajectory 163.
[0038] The target imitation trajectory 163 is generated by deep learning based on the state detection information 161 and the environment detection information 162, and serves as a target value 300 for the robot 1's motion. In contrast, the predicted state trajectory 164 indicates the predicted motion of the robot 1 when a control command 165 at the current stage is input to the actuator 151. Generally, the target imitation trajectory 163 cannot be directly converted into the control command 165. The target imitation trajectory 163 is expressed in terms of the angles and torque of the robot 1's joints, while the control command 165 is the current value input to the actuator 151. There is generally no formula for converting between angles and current values. Also, deviations between the target imitation trajectory 163 and the control command 165 can occur due to errors in the angle sensor 131 and the torque sensor 132. Furthermore, there are motions, such as sudden changes in direction, that are not problematic when operated by a human, but that the robot 1 cannot follow when operating autonomously. In other words, the motion characteristics of the drive unit 150 of the robot 1 that actually operates may not be satisfied. Therefore, it is necessary to correct the control command 165 based on the target simulated trajectory 163 and the predicted state trajectory 164 .
[0039] In the first embodiment, the control unit 120 generates a predicted state trajectory 164 based on the control command 165 and the like. Then, the control unit 120 corrects the control command 165 so that the predicted state trajectory 164 approaches the target imitation trajectory 163. The predicted state trajectory 164 approaches the target imitation trajectory 163 means that the degree of agreement between the predicted state trajectory 164 and the target imitation trajectory 163 decreases. The degree of agreement will be described later. Then, the control unit 120 generates the predicted state trajectory 164 based on the corrected control command 165 and the state detection information 161. Thereafter, the control unit 120 further corrects the control command 165 based on the generated predicted state trajectory 164 and the target imitation trajectory 163.
[0040] That is, the control unit 120 generates a tentative control command 165 and evaluates the degree of match between a predicted state trajectory 164 generated from the control command 165 and the target imitation trajectory 163. If the degree of match is large (if the degree has not reached the optimum), the control unit 120 corrects the control command 165 so that the degree of match decreases and the operating characteristics of the drive unit 150 are satisfied. By repeating this process, the control unit 120 generates a control command 165 that satisfies the operating characteristics of the drive unit 150.
[0041] The reason why the control unit 120 does not directly compare the control command 165 with the target imitation trajectory 163 is that the control command 165 and the target imitation trajectory 163 have different data contents and therefore cannot be compared. Therefore, the control unit 120 generates a predicted state trajectory 164 having a data format that allows comparison with the target imitation trajectory 163.
[0042] (Hardware Configuration) FIG. 2 is a diagram showing an example of the hardware configuration of the motion generation device 100 or the robot 1. As shown in FIG.
[0043] The motion generation device 100 is a device for realizing the functions of the imitation motion generation unit 110 and the control unit 120 shown in Fig. 1, and is, for example, a computer device. The motion generation device 100 includes a CPU 171, a ROM 172, a RAM 173, a display unit 174, an input unit 175, a communication I / F 176, and a system bus 177. Incidentally, "CPU" stands for Central Processing Unit. Furthermore, "ROM" stands for Read Only Memory. Furthermore, "RAM" stands for Random Access Memory. Finally, "I / F" stands for Interface.
[0044] The CPU 171 comprehensively controls the operations of the motion generation device 100, and processes the control program implemented in the robot 1 via a system bus 177. The ROM 172 is a non-volatile memory that stores the control program and the like required for the CPU 171 to execute processing. The control program and learning data may be stored in an external memory or a storage medium that is detachable from the motion generation device 100. The ROM 172, the external memory, and the detachable storage medium are used as examples of computer-readable storage media that store the control program executed by the motion generation device 100.
[0045] The RAM 173 is the main memory of the CPU 171 and functions as a work area, etc. That is, when the CPU 171 processes a control program, it loads the control program, etc. read from the ROM 172, into the RAM 173. Then, the CPU 171 executes the loaded control program, etc., thereby realizing the functional operations of the imitation action generator 110 and the control unit 120. In this way, the CPU 171, ROM 172, and RAM 173 work together to realize the processes executed by the imitation action generator 110 and the control unit 120.
[0046] The display unit 174 is composed of a monitor such as a liquid crystal display (LCD). LCD is an abbreviation for Liquid Crystal Display. The display unit 174 can display angle information at the joints of the robot 1 measured by the angle sensor 131, torque information applied to the joints of the robot 1 measured by the torque sensor 132, image information captured by the camera 141, and the like. The display unit 174 also allows the user to monitor the state of the robot 1 in real time. The input unit 175 is composed of a keyboard and a pointing device such as a mouse. Via the input unit 175, the user can issue commands to start or end control programs provided in the motion generation device 100.
[0047] The communication I / F 176 is an interface through which the motion generation device 100 communicates with the angle sensor 131, the torque sensor 132, the camera 141, and the actuator 151. The communication I / F 176 can be, for example, a LAN interface. Note that LAN is an abbreviation for Local Area Network. The system bus 177 communicatively connects the CPU 171, the ROM 172, the RAM 173, the display unit 174, the input unit 175, and the communication I / F 176.
[0048] As described above, the functions of each unit of the motion generation device 100 shown in Fig. 2 can be realized by the CPU 171 executing a control program. However, at least some of the functions of each unit of the motion generation device 100 shown in Fig. 2 may be configured to operate using dedicated hardware. In this case, the dedicated hardware operates under the control of the CPU 171.
[0049] Furthermore, a GPU or the like may be used instead of the CPU 171. GPU is an abbreviation for Graphic Processing Unit. Furthermore, an HDD, an SSD, or the like may be used instead of the ROM 172. HDD is an abbreviation for Hard Disk Drive, and SSD is an abbreviation for Solid State Drive.
[0050] 15, there are cases where the motion generating device 100 is provided outside the robot 1. In such a case, the hardware configuration of the robot 1 will be the same as that shown in FIG.
[0051] (Tasks Performed by Robot 1) FIG. 3 is a diagram showing an example of tasks performed by the robot 1. As shown in FIG.
[0052] In this embodiment, it is assumed that the robot 1 is holding a peg T1 and is executing a task of inserting the peg T1 into a hole T2. The camera 141 is installed so as to provide an overall view of the robot 1 inserting the peg T1 into the hole T2. The camera 141 may be fixed to the robot 1 or installed outside the robot 1, as long as it is configured to be able to communicate with the robot 1. However, if the camera 141 is installed outside the robot 1, the camera 141 is installed so as to provide an overall view of the execution of the task, including parts of the robot 1 that are closely related to the task, such as the robot's hand. Note that the actions of the robot 1, such as holding the peg T1 and moving its hand to insert the peg T1 into the hole T2, are referred to as a "series of actions."
[0053] Tasks applicable to this embodiment are not limited to those shown in Fig. 3. Tasks applicable to this embodiment include any tasks that can be performed by a human, such as the robot 1 grasping and moving an object, or the robot 1 handling or pushing a tool held in its hand.
[0054] (Movement Generation Method by Imitation Movement Generator 110) Fig. 4 is a flowchart showing the steps of the movement generation method performed by the imitation movement generator 110 in the first embodiment. Fig. 1 will be referred to as appropriate.
[0055] First, in step S101, the imitation motion linking unit 112 initializes the desired imitation trajectory 163 finally generated in the previous calculation so that it is empty.
[0056] Next, in step S102, the memory operation inference unit 111 acquires current angle information from the angle sensor 131, current torque information from the torque sensor 132, and current image information from the camera 141 (acquiring angle, torque, and image from the sensors). The angle in the angle information and the torque in the torque information are, for example, the angle and torque at the joints of the robot 1.
[0057] In step S103, the memory operation inference unit 111 predicts the target angle, target torque, and target image for one step ahead based on the input angle information, torque information, and image information. Of the target angle, target torque, and target image, the target angle and target torque become target values 300 (see FIG. 5 ). The angle and torque predicted by the memory operation inference unit 111 become the target values 300 when the robot 1 actually operates. The angle, torque, and image predicted by the memory operation inference unit 111 are referred to as the target angle, target torque, and target image, respectively. Furthermore, the information on the target angle, target torque, and target image will be referred to as target angle information, target torque information, and target image information, respectively, as appropriate.
[0058] In step S103, the plurality of target values 300 calculated by the memory operation inference unit 111, that is, the target angle, the target torque, and the target image, are linked together.
[0059] The prediction is performed by machine learning using deep learning or the like. Adjustment of the learning parameters using the teaching data 166 has been completed in advance. The target angle and target torque are target values 300, which are the angle and torque at which the robot 1 should operate one step ahead. Referring to FIG. 5, for the data point of target value 301a, the data point one step ahead is the data point of target value 301b. The target image is an image that predicts the state of the object and its surrounding environment that will change as a result of the operation of the robot 1. The memory operation inference unit 111 outputs the target angle, target torque, and target image. Incidentally, one step ahead means after a predetermined time.
[0060] Of the target angle, target torque, and target image output in step S103, the target angle and target torque become the target motion (target imitation trajectory 163) that the robot 1 should perform in the future.
[0061] In step S104 , the imitation motion linking unit 112 stores the target angle and target torque output by the memory motion inference unit 111 in the target imitation trajectory 163 .
[0062] Next, in step S105, the imitation action linking unit 112 determines whether the number of steps of the target imitation trajectory 163 generated up to this point has reached a preset number of steps. Referring to the target value 300 in Fig. 5, the number of steps is the number of data points (diamonds) constituting the target imitation trajectory 163, which is "6".
[0063] If the target imitation trajectory 163 has not yet reached the number of steps (S105→No), in step S106, the memorized operation inference unit 111 refers to the latest target angle, target torque, and target image. Then, in step S103, the memorized operation inference unit 111 predicts the target angle, target torque, and target image one step ahead based on the latest target angle, target torque, and target image referred to in step S106.
[0064] If the target imitation trajectory 163 has reached the preset number of steps (S105→Yes), in step S107, the imitation motion linking unit 112 outputs the finally calculated target imitation trajectory 163 to the control unit 120.
[0065] Steps S101 to S107 are "imitation action generation steps."
[0066] (Target Imitation Trajectory 163) FIG. 5 is a diagram showing an example of the generated target imitation trajectory 163. As shown in FIG.
[0067] In this embodiment, angle information and torque information of the joints of the robot 1, and image information captured by the camera 141 are input to the imitation action generation unit 110. The memory action inference unit 111 generates time-series data of predicted information (target angle, target torque, target image) for each of these. The time-series data of the target angle is indicated by a target value 301 (diamond), and the time-series data of the target torque is indicated by a target value 302 (diamond). Furthermore, the time-series data of the target image is indicated by reference numeral 303. Incidentally, reference numeral 311 (black circle) indicates an angle in the state detection information 161 acquired from the angle sensor 131, and reference numeral 312 (black circle) indicates a torque in the state detection information 161 acquired from the torque sensor 132. Furthermore, reference numeral 313 indicates an image in the image information acquired from the camera 141.
[0068] The time series data is configured from data points up to a preset number of steps ahead, with the time when generation of the target imitation trajectory 163 starts as the current time. As shown in Fig. 5 , of the three pieces of time series data (target angle, target torque, target image), the time series data of the target angle and target torque are output as the target imitation trajectory 163. On the other hand, when the memory operation inference unit 111 generates (learns) the target angle, target torque, and target image in step S103 of Fig. 4 , all three pieces of time series data are used.
[0069] As shown in FIG. 5 , all three time-series data are composed of data points with a common time step width. For example, as described above, the data point one step ahead of the data point for target value 301a is the data point for target value 301b. The number of steps is the number of steps for each of the target values 301 and 302 in the target imitation trajectory 163. As described above, in the example of FIG. 5 , the number of steps is six. The time step width basically matches the time step width of the teaching data 166 used during learning. However, it is also possible to generate the target imitation trajectory 163 so that the time step width of the target imitation trajectory 163 is shorter than the time step width of the teaching data 166. This makes it possible to request the robot 1 to perform a target movement at a higher speed. For example, if the time step width of the target imitation trajectory 163 is half the time step width of the teaching data 166, the movement of the robot 1 can be half the movement time of the teaching data 166. In this case, the robot 1 can perform a task in a shorter time, thereby improving efficiency. However, it becomes difficult for the robot 1 to follow the movement based on the target imitation trajectory 163. In other words, the movement of the robot 1 becomes faster, which may make it difficult to control the robot 1 to stop it, etc. Therefore, it is necessary to appropriately adjust the control unit 120, which will be described later.
[0070] (Movement Generation Method by Control Unit 120) Fig. 6 is a flowchart showing the steps of the movement generation method performed by the control unit 120 in the first embodiment. Fig. 1 will be referred to as appropriate.
[0071] First, in step S201 , the predicted state trajectory evaluation unit 122 acquires the desired simulated trajectory 163 output by the simulated action linking unit 112 .
[0072] Next, in step S202, the control command optimization unit 123 initializes the control command 165. The control command optimization unit 123 can shorten the time required for the optimization calculation in the subsequent stage by initializing the control command 165 so that the control command 165 finally determined in the previous calculation becomes the initial value. The control command optimization unit 123 may initialize the control command 165 by setting it to "0".
[0073] Subsequently, in step S203, the state predicting unit 121 acquires angle information and torque information measured by the angle sensor 131 and torque sensor 132 (acquires angle information and torque information from the sensors).
[0074] Then, in step S204, the state prediction unit 121 predicts the angle and torque of the robot 1 one step ahead using the current control command 165 and the acquired angle and torque information as initial values. To predict the angle and torque, a prediction model such as a dynamics equation that takes into account the dynamic characteristics of the robot 1, such as the moment of inertia and the moment of gravity, is used. Prediction based on the dynamics equation makes it possible to determine a control command 165 that the robot 1 can realistically follow. Using the dynamics equation, a series of predicted movements of the robot 1 are calculated based on the operating characteristics of the drive unit 150. The series of predicted movements is an operation in which the robot 1 holds a peg T1 and moves its hand to insert the peg T1 into the hole T2. The angle and torque predicted by the state prediction unit 121 are referred to as a predicted angle and a predicted torque, respectively. Furthermore, information on the predicted angle is referred to as predicted angle information, and information on the predicted torque is referred to as predicted torque information, as appropriate.
[0075] In step S205, the state prediction unit 121 stores the predicted angle and predicted torque predicted and generated in step S204 in the predicted state trajectory 164. The state prediction unit 121 also stores in the predicted state trajectory 164 a conversion formula (or a conversion map) between the predicted torque and the current value in the control command 165.
[0076] Subsequently, in step S206, the state prediction unit 121 determines whether the predicted state trajectory 164 has reached a pre-specified number of steps. The number of steps is the same as the number of steps in FIG.
[0077] If the predicted state trajectory 164 has not yet reached the number of steps (S205→No), in step S207, the state prediction unit 121 refers to the most recent predicted angle and predicted torque (S207). Then, in step S204, the state prediction unit 121 predicts the angle (predicted angle) and torque (predicted torque) for the next step based on the predicted angle and predicted torque referred to in step S207 and the current control command 165, in accordance with a prediction model such as a dynamic equation. In this way, the state prediction unit 121 first calculates a predicted state value 202 (see FIGS. 7A and 7B ), which is the predicted state detection information 161, based on the dynamic equation of the drive unit 150 and the control command 165 and the state detection information 161 as input. In the first embodiment, the predicted state value 202 is a predicted angle and a predicted torque. Furthermore, the state prediction unit 121 repeatedly calculates the next predicted state value 202 based on the calculated predicted state value 202, i.e., the predicted angle and predicted torque, and the control command 165 generated by the control command optimization unit 123. In this way, a plurality of predicted state values 202, i.e., the predicted angle and predicted torque, are generated.
[0078] Then, in step S205, the state prediction unit 121 stores the predicted angle and predicted torque, i.e., the predicted state value 202, in the predicted state trajectory 164. As a result, the generated multiple predicted state values 202 are linked together to generate the predicted state trajectory 164.
[0079] The number of steps of the predicted state trajectory 164 basically matches the number of steps of the target imitation trajectory 163. However, it is also possible to make the number of steps of the target imitation trajectory 163 smaller than the number of steps of the predicted state trajectory 164.
[0080] Note that thinning out the data points of the target imitation trajectory 163 and reducing the number of data points of the thinned target imitation trajectory 163 also has the effect of reducing the calculation cost when the memory operation inference unit 111 generates the target imitation trajectory 163. As described above, the memory operation inference unit 111 generates the target imitation trajectory 163 using deep learning or the like, and reducing the number of data points of the target imitation trajectory 163 makes it possible to reduce the calculation cost of imitation learning (deep learning).
[0081] If the predicted state trajectory 164 has reached the number of steps in step S205 (S205→Yes), in step S208, the predicted state trajectory evaluation unit 122 evaluates the generated predicted state trajectory 164. In the first embodiment, the predicted state trajectory evaluation unit 122 evaluates the predicted state trajectory 164 by calculating the degree of coincidence between the target simulated trajectory 163 and the predicted state trajectory 164 as an evaluation value. Details of the evaluation will be described later with reference to FIGS. 7A and 7B. Note that, in order to prevent the control command 165 calculated in a later stage from becoming too large, an evaluation value for evaluating the magnitude of the control command 165 may be added.
[0082] As described in step S205, in this embodiment, data on the target angle, predicted angle, target torque, and predicted torque for a predetermined number of steps (linked predicted angle and predicted torque) is used. This is because the robot 1 cannot respond to sudden changes. In other words, even if a sudden change in the target angle is taught in the target imitation trajectory 163, the actual actuator 151 may not be able to respond to such an angle change. Therefore, by using data on a predetermined number of predicted target angles, predicted angles, target torques, and predicted torques, it becomes possible to have the actuator 151 respond to a sudden change in angle, etc., from a time before that change.
[0083] Then, in step S209, the control command optimization unit 123 updates (corrects) the control command 165 in a direction that improves the evaluation value of the predicted state trajectory 164. In the first embodiment, a direction in which the degree of agreement between the target imitation trajectory 163 and the predicted state trajectory 164 increases is considered to be a good direction. The closer the target imitation trajectory 163 and the predicted state trajectory 164 are to each other, the higher the degree of agreement is. To update the control command 165, optimization calculations are performed using general mathematical optimization processes that utilize gradient methods, random sampling, and the like. At this time, the control command optimization unit 123 calculates, in real time, optimal control commands 165 that minimize the evaluation value. As a result, the control command optimization unit 123 corrects (updates) the control command 165 by performing mathematical optimization processing that solves an optimization problem that reproduces the target imitation trajectory 163 with desired (predetermined) acceleration / deceleration and force (predetermined output of the drive unit 150). That is, the control command optimization unit 123 performs mathematical optimization processing so that the degree of agreement, which is an evaluation value, decreases while satisfying the dynamics equation for the drive unit 150, thereby correcting (updating) the control command 165. Thereafter, the control command optimization unit 123 outputs the control command 165 for the robot 1.
[0084] Next, in step S210, the predicted state trajectory evaluation unit 122 determines whether the change in the evaluation value (degree of agreement) has converged and whether the evaluation value has reached an optimum value.
[0085] If the evaluation value has not yet reached the optimum value (S210→No), in step S211, the state prediction unit 121 refers to (obtains) the new control command 165 updated in step S209. Then, the processing from step S203 onwards is repeated again.
[0086] If the evaluation value has reached the optimum value (S210→Yes), the control command optimization unit 123 determines the control command 165 updated in step S209 as the optimum control command 165. Then, in step S212, the control command optimization unit 123 outputs the determined optimum control command 165 to the actuator 151.
[0087] Steps S201 to S207 and S211 are "state prediction steps." Step S208 is a "predicted trajectory evaluation step." Step S209 is a "correction step." Step S212 is an "output step."
[0088] (Evaluation Method) FIGS. 7A and 7B are diagrams showing an example of an evaluation method.
[0089] The degree of agreement between the target imitation trajectory 163 and the predicted state trajectory 164 is based on the error between the target imitation trajectory 163 and the predicted state trajectory 164, which are predicted at the same time step width (step width), for data points included in each trajectory. In Figures 7A and 7B, the data points of the target imitation trajectory 163 are indicated by target values 300 (diamonds), and the data points of the predicted state trajectory 164 are indicated by predicted state values 202 (circles). The evaluation value is the sum of the squares of the errors at each data point. The smaller the evaluation value, the better the result.
[0090] 7A and 7B, the data point (target value 300) of the target simulated trajectory 163 indicated by the target value 300 and the data point (predicted state value 202) of the predicted state trajectory 164 indicated by the predicted state value 202 are clearly closer in FIG. 7B than in FIG. 7A. Therefore, the evaluation value is smaller for the data shown in FIG. 7B than for the data shown in FIG. 7A. The predicted state trajectory evaluation unit 122 performs the evaluation in this manner.
[0091] When the number of steps of the target imitation trajectory 163 and the predicted state trajectory 164 are different, error evaluation is performed only on data points that exist on the target imitation trajectory 163 and the predicted state trajectory 164 at the same time. In other words, the number of data to be evaluated may be reduced by thinning out some of the data of the target imitation trajectory 163. Alternatively, the number of data of the target imitation trajectory 163 may be reduced when the imitation action generator 110 generates the target imitation trajectory 163.
[0092] Although only the evaluation of the angle is shown in FIGS. 7A and 7B, the same applies to the torque, and therefore a detailed description thereof will be omitted.
[0093] In this way, the target value 300 included in the target imitation trajectory 163 and the predicted state value 202 included in the predicted state trajectory 164 have time information. Then, the predicted state trajectory evaluation unit 122 evaluates the degree of coincidence between the target imitation trajectory 163 and the predicted state trajectory 164 based on the error between the (corresponding) target value 300 and the predicted state value 202 at the same time.
[0094] Based on the evaluation (degree of agreement), the control command optimization unit 123 performs a mathematical optimization process to approximate the target value 300 of the target imitation trajectory 163 and the predicted state value 202 of the predicted state trajectory 164, as shown in FIG. 7B . The mathematical optimization process is applied to correct the control command 165. In other words, the control command 165 is corrected as a result of the mathematical optimization process. The smallest evaluation value is achieved when the target value 300 of the target imitation trajectory 163 and the predicted state value 202 of the predicted state trajectory 164 agree with each other. However, due to differences in the performance of the actuators 151, the target value 300 of the target imitation trajectory 163 and the predicted state value 202 of the predicted state trajectory 164 generally do not agree with each other (although they may agree). Therefore, the control command optimization unit 123 corrects the control command 165 of the robot 1 by solving an optimization problem (using the mathematical optimization process) to reproduce the desired acceleration / deceleration and force.
[0095] (Effects of the First Embodiment) FIGS. 8A to 8D are diagrams showing the effects of the first embodiment.
[0096] 8A is a diagram showing an example of the relationship between the motion of the robot 1 based on the target imitation trajectory 163 when tracking control is performed on the target imitation trajectory 163 without taking into account the control characteristics of the robot 1 (i.e., conventional imitation learning), and the actual motion. Reference numeral 211 (diamond) indicates the time series of the motion of the robot 1 based on the target imitation trajectory 163 (e.g., the time series of the target angle). Reference numeral 212 indicates the time series of the motion of the robot 1 (e.g., the time series of the angle during the actual motion) when the robot 1 moves (actual motion) in accordance with the control command 165 based on the predicted state trajectory 164.
[0097] In FIG. 8A, the robot 1 is unable to follow the target movement trajectory, and the actual movement overshoots the target movement at the timing (reference numeral 213) when the direction of the target movement trajectory changes.
[0098] Fig. 8B shows an example of the result of task execution when tracking control is performed on the target trajectory without considering the control characteristics of the robot 1 as in Fig. 8A. As shown by reference numeral 213 in Fig. 8A, the robot 1 moves in an overshooting manner, causing the peg T1 held by the robot 1 to come into contact with the side wall of the hole T2 as shown in Fig. 8B.
[0099] 8C shows an example of the relationship between the actual movement and the movement of the robot 1 when tracking control is performed on the target imitation trajectory 163 taking into account the control characteristics of the robot 1 by the method of the first embodiment. Note that the control characteristics of the robot 1 are taken into account by using a dynamics equation in the processing of step S204 in FIG. 6. Also, in FIG. 8C, reference numerals 211 and 212 are the same as those in FIG. 8A.
[0100] As shown in Fig. 8C, according to the method described in the first embodiment, the robot 1 can accurately follow the target motion. Therefore, the robot 1 moves exactly as the target without overshooting even at the timing when the direction of the target motion changes (reference numeral 213 in Fig. 8A).
[0101] 8D shows an example of the result of task execution when tracking control is performed on the target imitation trajectory 163 taking into consideration the control characteristics of the robot 1 as shown in FIG. 8C. As shown in FIG. 8D, according to the technique shown in the first embodiment, the peg T1 held by the robot 1 can be inserted into the hole T2 without contacting the hole T2.
[0102] The motion generation device 100 in the first embodiment uses imitation learning, which allows for lower learning costs than reinforcement learning. Furthermore, because imitation learning is used, it is possible to provide a robot 1 capable of general-purpose motions without requiring specialized knowledge or complex programming. This allows the robot 1 to perform delicate and complex human motions.
[0103] In addition, in the first embodiment, the control command 165 is corrected so that the predicted state trajectory 164, which is generated taking into account the control characteristics of the robot 1, approaches the target imitation trajectory 163. Furthermore, the control command 165 can be generated in real time. That is, in the first embodiment, the control command 165 is determined by correcting the control command 165 for the imitation movement generated by the imitation movement generation unit 110 while taking into account the control characteristics and constraints of the robot 1. This makes it possible to accurately follow the imitation movement while suppressing an increase in learning cost. This improves the ability to follow the target imitation trajectory 163, enabling the robot 1 to complete the task to be executed with a high success rate. That is, according to the first embodiment, it is possible to accurately follow the imitation movement in imitation learning.
[0104] With conventional technology, it has been difficult to consider controllability and constraints when generating the teaching data 166, such as by having a human operate the robot 1. As a result, various problems have arisen when the robot 1 performs a task according to such teaching data 166. In particular, attempting to follow teaching data 166 that does not appropriately consider acceleration / deceleration or force at joints, etc., can result in a lower success rate for the task and unstable control. For example, the robot 1 may accelerate too much when performing a task, resulting in a missed movement, or may try to forcefully insert the peg T1 into the hole T2, potentially damaging the hole T2.
[0105] According to the first embodiment, by using imitation learning, imitation actions can be reproduced with low learning costs, and a control command 165 is determined so that a predicted state trajectory 164 generated in consideration of the control characteristics of the robot 1 approaches a target imitation trajectory 163, thereby achieving acceleration / deceleration and force adjustments that are desirable in terms of control. In other words, according to the first embodiment, it is possible to improve the controllability of the robot 1 while taking advantage of the low learning cost, which is an advantage of imitation learning.
[0106] In the first embodiment, a predicted state trajectory 164 is generated for a target imitation trajectory 163 generated by the imitation action generator 110, taking into consideration the control characteristics and constraints of the robot 1. Then, a control command 165 is corrected so that the generated predicted state trajectory 164 comes as close as possible to the target imitation trajectory 163. This makes it possible to explicitly consider the control characteristics and constraints of the robot 1, and the robot 1 can accurately reproduce the generated imitation action.
[0107] Furthermore, according to the first embodiment, the target imitation trajectory 163 and the dynamic equations can be selected according to the task to be applied and the robot 1 to which it is applied, so that it can be adapted to a variety of tasks independent of the application destination, enabling general-purpose application.
[0108] The imitation action generation unit 110 (memory action inference unit 111) generates a target imitation trajectory 163 by imitation learning using environment detection information 162 acquired from the environment detection unit 140 (camera 141) and state detection information 161. By using the environment detection information 162 in addition to the state detection information 161, the accuracy of the target imitation trajectory 163 can be improved.
[0109] Second Embodiment Next, a second embodiment of the present invention will be described with reference to FIGS. 9A to 10D.
[0110] (Evaluation Method) FIGS. 9A and 9B are diagrams showing an example of an evaluation method in the second embodiment.
[0111] In the second embodiment, in addition to the degree of coincidence between the target imitation trajectory 163 and the predicted state trajectory 164 shown in the first embodiment, three items are used as evaluation values: the amount of change in the state detection information 161, and the degree of satisfaction. The amount of change in the state detection information 161 is the amount of change in angle or torque in the predicted state trajectory 164, etc. The degree of satisfaction is the degree to which the predicted state trajectory 164 satisfies a preset constraint. The amount of change and the degree of satisfaction will be described later. A priority level is set by the user for each evaluation value.
[0112] The priority among the evaluation values (degree of agreement, amount of change, and degree of satisfaction) can be adjusted by a weighting coefficient that is set in advance for each evaluation value. The control command optimization unit 123 calculates a control command 165 that prioritizes and satisfies evaluation values with larger weighting coefficients. In other words, the control command optimization unit 123 performs optimization calculations for each evaluation value using a general mathematical optimization process that uses a gradient method, random sampling, or the like.
[0113] An example of evaluating the amount of change in the predicted angle in the predicted state trajectory 164 is shown in Figure 9A. Also, an example of evaluating the degree to which the predicted state trajectory 164 satisfies a preset constraint is shown in Figure 9B. In Figure 9A, predicted state values 202 indicate data points based on the predicted angle of the predicted state trajectory 164. In Figure 9B, reference numeral 203 indicates angular acceleration calculated from the predicted angle.
[0114] The change in the predicted angle in the predicted state trajectory 164 shown in Figure 9A is the change (reference numeral 221) between each data point (predicted state value 202) included in the predicted state trajectory 164 and the data points predicted in the steps before and after it. This change is the evaluation value. The evaluation value is the sum of the squares of the change between each data point and the data points before and after it. The smaller the evaluation value, the better the result.
[0115] The amount of change in the trajectory of the predicted angle shown in Figure 9A corresponds to evaluating the magnitude of the angular velocity. Furthermore, when evaluating the magnitude of the angular acceleration, the amount of change before and after a data point is evaluated for a time series of angular velocity obtained by differentiating the predicted angle. Furthermore, when taking the magnitude of the jerk into consideration, the amount of change before and after each data point is similarly evaluated for a time series of values obtained by differentiating the angular acceleration. While Figure 9A only shows the evaluation of the predicted angle, the same applies to the predicted torque, so detailed description will be omitted.
[0116] 9A , the predicted state trajectory evaluation unit 122 evaluates the magnitude of any one of the velocity, acceleration, and jerk included in the predicted state trajectory 164. In this embodiment, the velocity is the angular velocity, the acceleration is the angular acceleration, and the jerk is the derivative of the angular acceleration. Then, the control command optimization unit 123 corrects the control command 165 by mathematical optimization processing so as to reduce the degree of agreement while minimizing the magnitude of any one of the velocity, acceleration, and jerk included in the predicted state trajectory 164 as much as possible.
[0117] FIG. 9B is a diagram illustrating the degree to which the predicted state trajectory 164 satisfies the pre-set constraints.
[0118] Regarding the degree of satisfaction of the predicted state trajectory 164 with respect to the preset constraints, for example, the degree of deviation (222) from the preset maximum and minimum values for a data point (203) based on the predicted state trajectory 164 is used as the evaluation value. In the example shown in FIG. 9B , the degree of deviation from the maximum value (dashed line 231) and minimum value (dashed line 232) for each data point regarding angular acceleration is defined as a penalty. The sum of the penalties then becomes the evaluation value (degree of satisfaction). The maximum and minimum values are the preset constraints. In FIG. 9B , only one data point (203a) deviates from the maximum value, but if multiple data points deviate, the predicted state trajectory evaluation unit 122 calculates the degree of deviation for each data point and uses the sum of the respective deviation degrees as the evaluation value.
[0119] In this way, the predicted state trajectory evaluation unit 122 evaluates whether the predicted state value 202 included in the predicted state trajectory 164 satisfies the constraints of the drive unit 150 (maximum and minimum values in the example shown in Figure 9B).
[0120] The maximum and minimum values are set based on the specifications of the actuator 151, etc.
[0121] 9B shows only the evaluation of angular acceleration, but the same applies to angular velocity, jerk, and predicted torque, so detailed description will be omitted.
[0122] As described above, the user sets priorities among the degree of coincidence between the target imitation trajectory 163 and the predicted state trajectory 164, the amount of change in the predicted angle and predicted torque (see FIG. 9A), and the degree of satisfaction (see FIG. 9B). In this case, it is desirable to set the priorities (weighting coefficients) so that the degree of satisfaction is given the highest priority. This is because the degree of satisfaction is related to the specifications of the actuator 151.
[0123] The control command optimization unit 123 minimizes the evaluation value to calculate an optimal predicted state trajectory 164 that satisfies the constraints in real time. As a result, the control command optimization unit 123 corrects the control command 165 so that the degree of coincidence is reduced while satisfying the constraint conditions of the drive unit 150.
[0124] The process described with reference to FIG. 9 is the process performed in step S208 in FIG.
[0125] (Effects) FIGS. 10A to 10D are diagrams showing the effects of the second embodiment.
[0126] 10A shows an example of the relationship between the movement of the robot 1 based on the target imitation trajectory 163 and the actual movement when tracking control is performed taking into consideration only the degree of coincidence of the predicted state trajectory 164 with the target imitation trajectory 163. Note that in Fig. 10A and Fig. 10C, reference numerals 211 and 212 are the same as those in Fig. 8A.
[0127] 10A shows an example of the relationship between the motion of the robot 1 based on the target imitation trajectory 163 and the actual motion when only the technique described in the first embodiment is used. In FIG. 10A , the robot 1 follows the target motion (reference numeral 211) based on the target imitation trajectory 163. However, the robot 1 overreacts to the oscillatory motion contained in the teaching data 166 generated by a human, resulting in an oscillatory motion (reference numeral 212) (dotted line in FIG. 10A ). Such oscillatory motion occurs because the teaching data 166 contains tremors, etc., of the human being when the teaching data 166 is created.
[0128] FIG. 10B shows an example of the execution result of a task when tracking control is performed taking into consideration only the degree of coincidence of the predicted state trajectory 164 with the target simulated trajectory 163 as shown in FIG. 10A.
[0129] 10B, if the teaching data 166 includes a vibratory movement, when the robot 1 inserts the peg T1 it is holding into the hole T2, the robot 1 will vibrate as indicated by the reference numeral 321. As a result, the robot 1 will cause the peg T1 to come into contact with the side wall of the hole T2.
[0130] Figure 10C shows an example of the relationship between the movement of robot 1 based on the target imitation trajectory 163 and the actual movement when tracking control is performed taking into account not only the degree of coincidence of the predicted state trajectory 164 with the target imitation trajectory 163 but also the amount of change in the predicted angle of the predicted state trajectory 164 and constraints.
[0131] By applying the method described in the second embodiment, the robot 1 operates so as to maintain the trend of the target imitation trajectory 163 (dotted line in FIG. 10C ) while ignoring excessively oscillatory movements included in the target imitation trajectory 163. This is because the evaluation values are set so as to suppress large changes in the predicted angle and predicted torque, as shown in FIG. 9A .
[0132] Figure 10D shows an example of the results of task execution when tracking control is performed taking into account not only the degree of coincidence of the predicted state trajectory 164 with the target imitation trajectory 163 as shown in Figure 10C, but also the change in the predicted angle of the predicted state trajectory 164 and constraints.
[0133] As shown in FIG. 10D, by applying the technique shown in the second embodiment, the peg T1 held by the robot 1 can be inserted into the hole T2 without coming into contact with the hole T2.
[0134] In the second embodiment, in addition to the method described in the first embodiment, the control command 165 is determined by taking into consideration the amount of change and constraints based on the data points (predicted state values 202) of the predicted state trajectory 164. The method described in the first embodiment is a method of bringing the predicted state trajectory 164, which is generated by taking into consideration the control characteristics of the robot 1, closer to the target imitation trajectory 163. Bringing the predicted state trajectory 164 closer to the target imitation trajectory 163 means that the degree of match described above decreases. This enables the robot 1 to complete tasks with a high success rate without being affected by vibrational movements or noise unintentionally included in the teaching data 166 generated by a human. In other words, the second embodiment enables robustness against vibrations and noise unintentionally included in the teaching data 166.
[0135] Furthermore, by calculating the evaluation value (degree of satisfaction) based on the degree of deviation from the maximum and minimum values as shown in Figure 9B, it becomes possible to generate a control command 165 that takes into account the performance of the actuator 151.
[0136] Third Embodiment Next, a third embodiment of the present invention will be described with reference to FIGS. 11 to 13D.
[0137] (Configuration of Robot 1a) FIG. 11 is a diagram showing an example of the configuration of a robot 1a according to the third embodiment.
[0138] The motion generation device 100a includes an imitation motion generation unit 110, a control unit 120, and an obstacle detection unit 181 that detects an obstacle C (see FIG. 12) that exists around the robot 1a. The other configurations are the same as those in FIG. 1. The obstacle detection unit 181 identifies the obstacle C from an image (image information) measured by the camera 141, and outputs position information of the obstacle C relative to the robot 1a. The method for identifying the obstacle C is not limited to the method using the camera 141, and the obstacle C may be identified from point cloud information measured by LiDAR (not shown), for example.
[0139] 12A to 12C are diagrams showing an example of a task execution environment and an evaluation method for obstacle avoidance of the robot 1a according to the third embodiment.
[0140] 12A shows a task execution environment for the robot 1a, which is a target of the third embodiment and includes an obstacle C. In the task execution environment shown in Fig. 12A, the obstacle C, which was not present when a person generated the teaching data 166, exists between the peg T1 and the hole T2. In such a case, if the robot 1a attempts to insert the peg T1 into the hole T2 in the shortest distance, there is a risk that the robot 1a or the peg T1 held by the robot 1a will come into contact with the obstacle C along the way, as shown in Fig. 12A.
[0141] 12B and 12C show examples of evaluation functions for evaluating the possibility of contact between the predicted state trajectory 164 and the obstacle C in the evaluation (calculation of the evaluation value) of the predicted state trajectory 164 performed by the predicted state trajectory evaluation unit 122. The evaluation function shown in Fig. 12C determines an evaluation value based on the reference positions R1 to R4 shown in Fig. 12B and the distance to the obstacle C.
[0142] The predicted state trajectory evaluation unit 122 calculates the distance to the obstacle C for each of the plurality of reference positions R1 to R4 set on the robot 1a or peg T1, which are calculated from each of the data points included in the predicted state trajectory 164. At this time, the predicted state trajectory evaluation unit 122 determines that the possibility of contact increases the closer the distance between the obstacle C and the closest reference position, which is the reference position R1 to R4 that is closest to the obstacle C, as shown in Fig. 12C. The greater the distance between the closest reference position and the obstacle C and the lower the possibility of contact, the better the evaluation result.
[0143] Step S209 in FIG. 6 is performed by performing an optimization calculation using the evaluation value based on the distances between the reference positions R1 to R4 and the obstacle C and the evaluation value based on the degree of match shown in the first embodiment. Note that the user sets a priority between the degree of match between the target imitation trajectory 163 and the predicted state trajectory 164 and the evaluation value based on the distances between the reference positions R1 to R4 and the obstacle C. In this case, it is desirable that the evaluation value based on the distances between the reference positions R1 to R4 and the obstacle C is given priority (has a larger weighting coefficient) over the degree of match between the target imitation trajectory 163 and the predicted state trajectory 164. This is because avoiding contact with the obstacle C is given priority over matching between the target imitation trajectory 163 and the predicted state trajectory 164.
[0144] The control command optimization unit 123 corrects the control command 165 by performing optimization calculations for each evaluation value using a general mathematical optimization method that uses a gradient method, random sampling, etc. As a result, the control command optimization unit corrects the control command 165 through mathematical optimization processing so that the predicted state trajectory 164 reduces the degree of coincidence with the obstacle C while avoiding contact with the obstacle C.
[0145] The processing shown in FIGS. 12A to 12C is the processing performed in step S209 in FIG.
[0146] (Effects) FIGS. 13A to 13D are diagrams showing the effects of the technique shown in the third embodiment.
[0147] 13A shows an example of the relationship between the actual movement and the movement of the robot 1a based on the target imitation trajectory 163 when tracking control is performed taking into consideration only the degree of coincidence of the predicted state trajectory 164 with the target imitation trajectory 163 as an evaluation value. Note that in FIGS. 13A and 13C, the reference numerals 211 and 212 are the same as those in FIG. 8A.
[0148] 13A shows an example of the relationship between the motion of the robot 1a based on the target imitation trajectory 163 and the actual motion when only the method shown in the first embodiment is used. As shown in FIG. 13A, the robot 1a satisfactorily follows the target motion without overshooting.
[0149] 13B shows an example of the execution result of a task when tracking control is performed taking into consideration only the degree of coincidence of the predicted state trajectory 164 with the target imitation trajectory 163 as shown in FIG. 13A. Because obstacle C did not exist when the human created the teaching data 166, the presence of obstacle C is not taken into consideration in the target imitation trajectory 163. Therefore, when the robot 1a operates to follow the target imitation trajectory 163, a part of the robot 1a comes into contact with obstacle C, as shown in FIG. 13B.
[0150] 13C shows an example of the relationship between the actual movement and the movement of the robot 1a based on the target imitation trajectory 163 when tracking control is performed taking into consideration not only the degree of agreement shown in the first embodiment as an evaluation value but also the possibility of contact between the robot 1a and the obstacle C. Here, it is assumed that the weighting coefficient of the evaluation value is adjusted so that a lowering of the possibility of contact with the obstacle C takes priority over the degree of agreement with the target imitation trajectory 163. In other words, the control command optimization unit 123 performs optimization calculations so that the priority (weighting coefficient) of the evaluation value based on the distance between the reference positions R1 to R4 and the obstacle C is higher than the evaluation value based on the degree of agreement shown in the first embodiment.
[0151] When such processing is performed, priority is given to avoiding the obstacle C even if the target imitation trajectory 163 is ignored to some extent. As a result, after avoiding the obstacle C, the robot 1a operates so as to come as close as possible to the target imitation trajectory 163.
[0152] Fig. 13D shows an example of the result of task execution when the robot 1a operates as shown in Fig. 13C. As shown in Fig. 13D, the robot 1a operates to avoid the obstacle C, and then operates to move the peg T1 it is holding to approach the hole T2 in order to perform the original task.
[0153] In the third embodiment, in addition to the method described in the first embodiment, a control command 165 is determined so that the robot 1a avoids contact with the obstacle C. The method described in the first embodiment is a method in which a predicted state trajectory 164 generated in consideration of the control characteristics of the robot 1a is brought closer to the target mimic trajectory 163. This enables the robot 1a to complete the task to be performed without contacting the obstacle C that did not exist when the teaching data 166 was generated. Therefore, the environmental conditions in which the robot 1a can be used are relaxed.
[0154] Fourth Embodiment Next, a fourth embodiment of the present invention will be described with reference to FIG.
[0155] FIG. 14 is a diagram showing an example of candidates for the target simulated trajectory 163 and an evaluation method of the predicted state trajectory evaluation unit 122 in the fourth embodiment.
[0156] In the fourth embodiment, the imitation movement generator 110 outputs a plurality of target imitation trajectories 163 .
[0157] This is possible by designing deep learning so that multiple target imitation trajectories 163 are output in deep learning that generates multiple target imitation trajectories 163. The multiple target imitation trajectories 163 are output, for example, according to accuracy. Accuracy is the rate of agreement of the target imitation trajectory 163 with the teaching data 166. In other words, multiple target imitation trajectories 163 with different rates of agreement with the teaching data 166 are intentionally output. The multiple target imitation trajectories 163 are referred to as target imitation trajectories 162a to 162c.
[0158] Specifically, the imitation movement generator 110 ranks the multiple target imitation trajectories 162a to 162c that it predicts and outputs in descending order of accuracy. The imitation movement generator 110 then outputs a specified number of target imitation trajectories 162a to 162c in descending order of accuracy. In the example shown in Fig. 14, three target imitation trajectories 162a to 162c are output, so the imitation movement generator 110 outputs the three target imitation trajectories 162a to 162c in descending order of angle.
[0159] The predicted state trajectory evaluation unit 122 then evaluates the degree of agreement between the predicted state trajectory 164 and each of the target imitation trajectories 162a to 162c (calculates an evaluation value). The degree of agreement between the predicted state trajectory 164 and the target imitation trajectories 162a to 162c is based on the error between the target imitation trajectories 162a to 162c predicted at the same time step (data point) and the data point (predicted state value 202) of the predicted state trajectory 164 for each data point (target value 300) included in each of the target imitation trajectories 162a to 162c. The sum of the squared sums of the errors for each data point is defined as the evaluation value. As a result, the smaller the evaluation value, the better the result.
[0160] Then, the control command optimization unit 123 updates (corrects) the control commands 165 by optimization calculation using a mathematical optimization method based on such evaluation values. In this way, the control command optimization unit 123 corrects the control commands 165 based on the priorities set in advance for the plurality of target imitation trajectories 163, respectively, and the degrees of coincidence between the plurality of target imitation trajectories 163 and the predicted state trajectory 164.
[0161] Here, the accuracy of the target imitation trajectories 162 a to 162 c may be used as a weighting factor, and the predicted state trajectory evaluation unit 122 may prioritize evaluation of the error between the target imitation trajectories 162 a to 162 c and the highly accurate target imitation trajectories 162 a to 162 c. If the target imitation trajectories 162 a to 162 c have the same priority, the average value of each data point on the target imitation trajectories 162 a to 162 c becomes the control command 165.
[0162] The process described with reference to FIG. 14 is the process performed in step S208 in FIG.
[0163] (Effects) In the fourth embodiment, the control command 165 is determined so that the predicted state trajectory 164, which is generated in consideration of the control characteristics of the robot 1, matches the multiple target imitation trajectories 162a to 162c. This makes it possible to prevent excessive tracking of target actions far in the future from the current time, which generally has low accuracy. This makes it possible to robustly perform a task even when an incorrect imitation action is predicted (with a large prediction error), contributing to an improvement in the success rate of the task.
[0164] In both the generation of the target imitation trajectory 163 and the generation of the predicted state trajectory 164, predicted values are used to calculate predicted values for the next step. Here, the predicted values are predicted angles and torques (i.e., predicted state values 202). Therefore, it is expected that the errors of the respective data points will become larger as we move from the present, when the generation of the target imitation trajectory 163 and the predicted state trajectory 164 starts, into the future. In other words, it is expected that the errors of the data points of the target imitation trajectory 163 and the predicted state trajectory 164 relative to the true data points (true angles and torques) will become larger. The true angles and torques are the actual angles and torques of the robot 1 at a given time.
[0165] If the control command 165 is updated using only one target imitation trajectory 163, there is a risk that the control command 165 will be generated and include an error in the target imitation trajectory 163 that is being used. In the fourth embodiment, a plurality of target imitation trajectories 162a to 162c are used, and the control command 165 is updated so that the errors in the respective target imitation trajectories 162a to 162c are averaged (the errors are alleviated). This makes it possible to obtain the above-described effects.
[0166] <Robot Control System Z> FIG. 15 is a diagram showing an example of a robot control system Z according to this embodiment.
[0167] In the robot control system Z shown in Fig. 15, a motion generating device 100 is provided outside the robot 1 so as to be able to communicate with it. The configuration of the motion generating device 100 is the same as that shown in Fig. 1, and the configuration of the robot 1 is the same as that shown in Fig. 1 except that the motion generating device 100 is not provided inside the robot 1.
[0168] Then, angle information acquired by an angle sensor 131 provided in the robot 1 and torque information acquired by a torque sensor 132 are input to the memory operation inference unit 111 and the state prediction unit 121. Also, an image captured by a camera 141 is input to the memory operation inference unit 111. Furthermore, a control command 165 generated by the control command optimization unit 123 is input to an actuator 151 of the robot 1 and also to the state prediction unit 121.
[0169] In the first to fourth embodiments, the motion generating device 100 is configured to be mounted inside the robot 1. However, if the motion generating device 100 is configured to be able to communicate with the angle sensor 131, torque sensor 132, camera 141, and actuator 151, the motion generating device 100 can be installed outside the robot 1 as shown in FIG.
[0170] As described above, when the motion generation device 100 is installed outside the robot 1 as shown in Fig. 15, the hardware configuration is as shown in Fig. 2. The motion generation device 100 shown in Fig. 15 may also be provided with an obstacle detection unit 181 as shown in Fig. 1.
[0171] The robot 1 shown in this embodiment is applicable to industrial robot arms, medical robot arms, visual inspection robot systems, and the like.
[0172] The present invention is not limited to the above-described embodiments and includes various modifications. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those having all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, and it is also possible to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace part of the configuration of each embodiment with other configurations.
[0173] Furthermore, some or all of the above-described components, functions, memory operation inference unit 111 to obstacle detection unit 181, ROM 172, RAM 173, etc. may be implemented in hardware, for example, by designing them as integrated circuits. Furthermore, as shown in FIG. 2 , the above-described components, functions, etc. may be implemented in software by a processor, such as CPU 171, interpreting and executing programs that implement the respective functions. Information, such as programs, tables, and files, that implement the respective functions may be stored on a hard disk or in a memory, a recording device such as an SSD, or a recording medium such as an IC (Integrated Circuit) card, an SD (Secure Digital) card, or a DVD (Digital Versatile Disc). Furthermore, in each embodiment, control lines and information lines are shown that are considered necessary for explanation, and not all control lines and information lines are necessarily shown in the actual product. In reality, it is safe to assume that almost all components are interconnected.
[0174] 1, 1a Robot 100, 100a Motion generation device 110 Imitation motion generation unit 111 Memory motion inference unit 112 Imitation motion linking unit 121 State prediction unit 122 Predicted state trajectory evaluation unit 123 Control command optimization unit 130 State detection unit 131 Angle sensor (state detection unit) 132 Torque sensor (state detection unit) 140 Environment detection unit 141 Camera (environment detection unit) 150 Driving unit 151 Actuator (driving unit) 161 State detection information 162, 162a to 162c Environment detection information 164 Predicted state trajectory 165 Control command 202 Predicted state value 300, 301, 301, 301a, 301b,302 Target value C Obstacle Z Robot control system S101 Initialize target imitation trajectory (imitation motion generation step) S102 Acquire angle, torque, and image from sensor (imitation motion generation step) S103 Learning motion inference unit predicts target angle, target torque, and target image for one step ahead (imitation motion generation step) S104 Imitation motion linking unit stores predicted angle and predicted torque in target imitation trajectory (imitation motion generation step) S105 Determine whether the number of steps in the target imitation trajectory has reached a specified number of steps (imitation motion generation step) S106 Refer to the latest target angle, target torque, and target image (imitation motion generation step) S107 Imitation motion linking unit outputs target imitation trajectory (imitation motion generation step) S201 Acquire target imitation trajectory (state prediction step) S202 Initialize control command (state prediction step) S203 Acquire angle information and torque information from sensor (state prediction step) S204 The state prediction unit predicts the angle and torque of the robot one step ahead (state prediction step). S205 The predicted angle and predicted torque are stored in the predicted state trajectory (state prediction step). S206 It is determined whether the number of steps in the predicted state trajectory has reached a specified number of steps (state prediction step). S207 The latest predicted angle and predicted torque are referenced (state prediction step). S208 The predicted state trajectory evaluation unit evaluates the predicted state trajectory (predicted trajectory evaluation step). S209 The control command optimization unit updates the current control command (correction step). S211 The latest control command is obtained (state prediction step). S212 The control command optimization unit outputs the optimal control command (output step).
Claims
1. A robot comprising: a driving unit that drives a robot; a state detection unit that acquires information regarding the state of the robot; an imitation movement generation unit that generates a target imitation trajectory, which is time series data of target values related to the movement of the robot, by imitation learning based on the state detection information acquired from the state detection unit; a state prediction unit that uses a control command to the driving unit at the current time and the state detection information to generate a predicted state trajectory, which is time series data related to a predicted movement of the robot based on the movement characteristics of the driving unit; a predicted state trajectory evaluation unit that compares the target imitation trajectory generated by the imitation movement generation unit with the predicted state trajectory generated by the state prediction unit and evaluates the degree of agreement between the target imitation trajectory and the predicted state trajectory; and a control command optimization unit that corrects the control command based on the degree of agreement evaluated by the predicted state trajectory evaluation unit, and outputs the corrected control command to the driving unit.
2. The robot described in claim 1, wherein the imitation action generation unit comprises a memory action inference unit and an imitation action connection unit, wherein the memory action inference unit uses the state detection information as input to calculate a target value for the robot's next action, the imitation action connection unit generates the target imitation trajectory by connecting the multiple target values calculated by the memory action inference unit, and the state prediction unit uses the control command and the state detection information as input to calculate a predicted state value, which is the state detection information to be predicted, based on the dynamics equation of the drive unit, and further calculates the next predicted state value based on the calculated predicted state value and the control command generated by the control command optimization unit, thereby generating multiple predicted state values, and connecting the generated multiple predicted state values to generate a predicted state trajectory.
3. The robot described in claim 1, further comprising an environment detection unit that acquires environmental information around the robot, and wherein the imitation motion generation unit generates the target imitation trajectory using the state detection information and the environmental detection information acquired from the environment detection unit.
4. The robot described in claim 1, characterized in that the target value included in the target imitation trajectory and the predicted state value included in the predicted state trajectory have time information, and the predicted state trajectory evaluation unit evaluates the degree of agreement between the target imitation trajectory and the predicted state trajectory based on the error between the target value and the predicted state value at the same time.
5. The robot according to claim 1, characterized in that the control command optimization unit corrects the control commands by performing a mathematical optimization process that solves an optimization problem that reproduces the target imitation trajectory with desired acceleration / deceleration and force.
6. The robot described in claim 5, characterized in that the predicted state trajectory evaluation unit evaluates the magnitude of any one of the velocity, acceleration, and jerk included in the predicted state trajectory, and the control command optimization unit corrects the control command by the mathematical optimization process so as to reduce the degree of agreement while minimizing the magnitude of any one of the velocity, acceleration, and jerk included in the predicted state trajectory as much as possible.
7. The robot described in claim 1, characterized in that the predicted state trajectory evaluation unit evaluates whether the predicted state value included in the predicted state trajectory satisfies the constraint conditions of the drive unit, and the control command optimization unit corrects the control command so as to approach the target imitation trajectory while satisfying the constraint conditions of the drive unit.
8. The robot described in claim 5, characterized in that the control command optimization unit corrects the control commands through the mathematical optimization process so that the predicted state trajectory avoids contact with an obstacle while reducing the degree of coincidence.
9. The robot described in claim 1, characterized in that the imitation motion generation unit outputs a plurality of target imitation trajectories, and the control command optimization unit corrects the control command based on priorities set in advance for each of the plurality of target imitation trajectories and the degree of coincidence between the plurality of target imitation trajectories and the predicted state trajectory.
10. An action generation device comprising: an imitation action generation unit that generates a target imitation trajectory, which is time series data of target values related to the action of the robot, by imitation learning based on state detection information, which is information acquired from a state detection unit and is information related to the state of the robot; a state prediction unit that generates a predicted state trajectory, which is time series data related to the predicted action of the robot based on the action characteristics of the drive unit, using a control command to a drive unit that drives the robot at the current time and the state detection information; a predicted state trajectory evaluation unit that compares the target imitation trajectory generated by the imitation action generation unit with the predicted state trajectory generated by the state prediction unit and evaluates the degree of agreement between the target imitation trajectory and the predicted state trajectory; and a control command optimization unit that corrects the control command based on the degree of agreement evaluated by the predicted state trajectory evaluation unit and outputs the corrected control command to the drive unit.
11. A robot control system comprising: a robot; and an action generation device, wherein the robot comprises: a drive unit that drives the robot; a state detection unit that acquires information regarding the state of the robot; and the action generation device described in claim 10.
12. A motion generation method in which a motion generation device that generates information for a robot to operate executes the following steps: an imitation motion generation step of generating a target imitation trajectory, which is time series data of target values related to the robot's motion, by imitation learning based on state detection information, which is information acquired from a state detection unit and is information related to the state of the robot; a state prediction step of generating a predicted state trajectory, which is time series data related to the predicted motion of the robot based on the motion characteristics of the drive unit, using a control command to a drive unit that drives the robot at the current time and the state detection information; a predicted trajectory evaluation step of comparing the target imitation trajectory generated in the imitation motion generation step with the predicted state trajectory generated in the state prediction step, and evaluating the degree of agreement between the target imitation trajectory and the predicted state trajectory; a correction step of correcting the control command based on the degree of agreement evaluated in the predicted trajectory evaluation step; and an output step of outputting the corrected control command to the drive unit.
Citation Information
Patent Citations
Driving control device, driving control method, and program
JP2022034227A
Track formation device, track formation method and track formation program
JP2022066086A