Robot control device, robot operation learning method, and robot operation learning device
The robot control device uses simulation-based learning models to enhance annotation efficiency and feature detection, enabling robots to perform manipulation tasks effectively across varying environments and objects.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2026-04-09
AI Technical Summary
Existing robot control systems face challenges in efficiently performing manipulation tasks under varying environmental conditions due to labor-intensive manual annotation and inaccurate automatic annotation, particularly when fluorescent markers are hidden or lighting conditions vary, leading to inefficiencies and increased costs.
A robot control device that includes a coordinate calculation unit, first and second learning model generation units, and a robot control unit, which utilizes simulation environments to generate learning models that estimate object and robot positions, allowing for automatic annotation and detection of feature regions even in challenging conditions.
Enables robots to perform manipulation tasks efficiently under diverse environmental conditions by increasing the number of automatically annotated processes and detecting hidden features, facilitating learning and expanding the range of environmental conditions and object types handled.
Smart Images

Figure JP2025025339_09042026_PF_FP_ABST
Abstract
Description
Robot control device, robot motion learning method, and robot motion learning device
[0001] This disclosure relates to a robot control device, a robot motion learning method, and a robot motion learning device.
[0002] To perform manipulation tasks under different environmental conditions, a robot control system has been developed that uses a learning model to extract task-related two-dimensional positions from raw images, and then controls the robot based on these extracted two-dimensional positions.
[0003] A training method for a learning model that extracts two-dimensional positions involves collecting a training dataset consisting of images and ground truth position data. Typically, the collected training dataset is used to approximate the estimated position output by the learning model with the ground truth position data in the training dataset using machine learning techniques. Ground truth positions can be obtained by manually annotating positions on the image, or by applying image processing techniques that automatically detect these positions based on recognized shapes and colors.
[0004] Patent Document 1 discloses an inference device comprising: a trained model trained using a second training image of a training target captured under second lighting conditions and a first training image of the training target illuminated under first lighting conditions in which a marking unit provided on the training target is recognizable under the first lighting conditions as learning data; an estimation unit that uses the trained model to estimate information of the marking unit from the second image captured by an imaging unit; and a robot control unit that creates an action plan for the robot according to the information of the marking unit estimated by the estimation unit and controls the robot's movements based on the action plan.
[0005] Japanese Patent Publication No. 2021-047839
[0006] Manual annotation is time-consuming and laborious. On the other hand, automatic annotation can be inaccurate if there are variations in background color, lighting conditions, etc., or if relevant image features are hidden from view.
[0007] When using the robot motion control device described in Patent Document 1, preparatory work is required to attach fluorescent markers, etc., to the training object used for photography under the second lighting conditions, and a process of irradiating with ultraviolet light to make the fluorescent markers emit light is required to create a trained model, which is time-consuming and costly. In addition, depending on the circumstances in which the training object is placed, the fluorescent markers, etc., attached to the training object may be hidden, making it impossible to detect the position of the fluorescent markers, etc.
[0008] The purpose of this disclosure is to increase the number of processes that can be automatically annotated, thereby saving labor, and to detect feature areas of an image that are not detectable during automatic annotation.
[0009] The robot control device of this disclosure includes: a coordinate calculation unit that calculates coordinates on a simulation image from data of an object and a robot manipulator in a simulation environment; a first learning model generation unit that generates a first learning model including data of a first estimated position of the object and the robot manipulator in the simulation environment using first teaching data including coordinates on a simulation image; a second learning model generation unit that generates a second learning model including data of a second estimated position of the object and the robot manipulator in real space based on second teaching data generated by predetermined processing using data of the object and the robot manipulator in real space, the first learning model and the weights of the first learning model; and a robot control unit that controls the operation of the robot manipulator based on the data of the second estimated position.
[0010] According to this disclosure, the number of processes that can be automatically annotated is increased, reducing labor, and even if the feature regions of an image necessary for processing cannot be detected during automatic annotation, those feature regions can be detected.
[0011] This is an overall configuration diagram showing the robot control device according to Example 1. This is a functional block diagram showing the simulation environment generated by the robot control device according to Example 1. This is a configuration diagram showing the hardware of the robot control device according to Example 1. This is a flowchart showing the learning method of the first learning model 4 in Figure 2A. This is a perspective view showing the coordinates on the simulation image generated by the robot control device according to Example 1, the coordinate system necessary for calculating those coordinates, camera parameters, etc. This is a functional block diagram showing the configuration for operating a robot arm in a real environment using the first learning model 4 in Figure 2A. This is a flowchart showing the learning of the second learning model of the robot control device according to Example 1 and the generation of robot motion in a real environment. This is a flowchart showing the learning of robot motion and the generation of robot motion in a real environment by the robot control device according to Example 1. This is a diagram showing the variations of three-dimensional position obtained from the simulator of the robot control device according to Example 1. This is a flowchart showing the camera posture optimization algorithm of the robot control device according to Example 1. This is a functional block diagram showing the simulation environment generated by the robot control device according to Example 2. This is a diagram showing an example of an image 24 of a T-shirt 1001 obtained from the camera 10 in the simulation environment 900 of Figure 10. This is a diagram showing another example of an image 24 of a T-shirt 1001 obtained from the camera 10 in the simulation environment 900 of Figure 10. This figure shows an example of an image 24 that shows the calculation units 1002, 1003, 1004, and 1005 shown on the display unit 26 of Figure 10 together with the estimated positions 1102, 1103, 1104, and 1105.
[0012] The robot control device, robot motion learning method, and robot motion learning device disclosed herein enable robots to efficiently perform manipulation tasks under diverse environmental conditions. Furthermore, they enable the creation of simple datasets and the training of learning models that estimate the parts to be focused on when controlling the robot's motion.
[0013] Furthermore, this disclosure makes it possible to facilitate the learning of manipulation tasks. In addition, it is possible to expand the environmental conditions in which the robot is installed and the types of objects that the robot handles.
[0014] The following describes an example using drawings.
[0015] Figure 1 is an overall configuration diagram showing the robot control device according to this embodiment.
[0016] As shown in this figure, the robot control device 100 includes a first model learning device 151, a second model learning device 152, and a robot control unit 153.
[0017] The first model learning device 151 receives data such as images, the position of the robot end-effector, the position of the object being manipulated, and the camera orientation from the simulation environment 155, and creates a first learning model. This data is also transmitted from the simulation environment 155 to the robot control unit 153. The first model learning device 151 transmits the created first learning model to the second model learning device 152. The second model learning device 152 creates a second learning model using the transmitted first learning model. The second model learning device 152 transmits the second learning model to the robot control unit 153.
[0018] The robot control unit 153 can send motion commands to the robot 154, which is the target of control, using the second learning model and data such as images, the position of the robot end-effector, the position of the object being manipulated, and the camera orientation. The robot control unit 153 can also send virtual motion commands to the simulation environment 155.
[0019] Figure 2A is a functional block diagram showing the simulation environment generated by the robot control device of this embodiment.
[0020] In this figure, the first model learning device 1 includes a coordinate calculation unit 5, a database 6, a first model learning unit 7 (first learning model generation unit), and a display unit 26.
[0021] Simulation environment 2 is obtained as a result of a simulation that mimics the real environment, and is constructed in which objects and other things (real objects) from the real environment are represented as existences on a coordinate system in the virtual space. Simulation environment 2 includes a robot arm 8, an object to be manipulated 9 (also called "object"), and a camera 10 in the virtual space. Here, a predetermined part of the object to be manipulated 9 will be called the object position 20.
[0022] The control device 3 includes a trajectory generation unit 13, a robot input unit 14, a switch 15 for switching the robot control mode to either manual mode 16 or automatic mode 17, and a robot control unit 18. The position and shape of objects in the simulation environment 2 are specified in the coordinates of the virtual space.
[0023] The coordinate calculation unit 5 is configured to receive data on the robot end-effector position 19, the object position 20, and the camera orientation 21. The coordinate calculation unit 5 is configured to use this data to calculate the coordinates 22 on the simulation image and transmit this data to the database 6 and the display unit 26. Here, the robot end-effector position 19 is the part of the robot manipulator that handles the object. Specifically, it may be a predetermined part between the two fingers of the robot hand of the robot manipulator, or it may be the fingertips of the fingers of the robot hand. It may also be a predetermined part of at least two sides of the robot hand.
[0024] Furthermore, the object position 20 may be the part corresponding to the center of gravity of the object's three-dimensional shape.
[0025] Database 6 stores the image 24 data transmitted from the simulation environment 2.
[0026] The first model learning unit 7 receives the data stored in the database 6 as teaching data (first teaching data), learns from it, and creates a trained model. The data of the trained model includes the estimated position 23 (first estimated position) of objects, etc., and is designated as the first trained model 4. The image data 24 is also designated as the first trained model 4.
[0027] The display unit 26 receives data of the estimated position 23 and the image 24, and displays the estimated position 23, the image 24, etc. for the user.
[0028] The trajectory generation unit 13 receives data on the estimated position 23 and the robot end-effector position 19. In automatic mode 17, the robot control unit 18 receives data on the estimated position 23 and the robot end-effector position 19 from the trajectory generation unit 13 and controls the operation of the robot arm 8 in the simulation environment 2.
[0029] On the other hand, in manual mode 16, the robot control unit 18 controls the movement of the robot arm 8 in the simulation environment 2 based on signals from the robot input unit 14 operated by the user, rather than data from the trajectory generation unit 13.
[0030] In summary, the first model learning device 1 generates the first learning model 4. The control device 3 generates motion commands to move the robot arm 8 in the simulation environment 2 based on the first learning model 4, data obtained from the simulation environment 2, etc.
[0031] Figure 2B is a configuration diagram showing the hardware of the robot control device according to this embodiment.
[0032] In this figure, the robot control device 102 is implemented using a computer and includes a display 121, an input device 122, a CPU 123 (Central Processing Unit), a RAM 124 (Random Access Memory), a ROM 125 (Read Only Memory), a communication device 126, a media reader 127, and an auxiliary storage device 128. The display 121 corresponds to the display unit 26 in Figure 1A, etc. The input device 122 corresponds to the robot input unit 14 in Figure 1A, etc.
[0033] First, the CPU 123 is a so-called processor and is a unit that performs various calculations. The CPU 123 performs various processes by executing robot control programs and the like that loaded from the auxiliary storage device 128 into the RAM 124.
[0034] Here, the robot control program is, for example, an application program executable on the program of an OS (Operating System). In this embodiment, the robot control program is composed of a plurality of modules for each function, but may also be realized by a plurality of independent programs for each function.
[0035] Further, the robot control program may be installed in the auxiliary storage device 128 from a portable storage medium via the media reading device 127. That is, the robot control program can be stored in the storage medium. The CPU 123 may execute processing according to a program other than the robot control program. In this case, it is desirable to store this program in the auxiliary storage device 128.
[0036] Further, the RAM 124 is a memory that stores programs such as the robot control program executed by the CPU 123 and data necessary for program execution. The ROM 125 is a memory that stores programs necessary for starting the robot control device 102, the OS, and the like. The communication device 126 is connected to sensors such as a robot arm and a camera in the actual environment via a communication path. The communication path can be realized by a network such as a LAN (Local Area Network) or the Internet, and can be wired or wireless.
[0037] The media reading device 127 is a device that reads information from a portable storage medium such as a flash memory or a CD-ROM.
[0038] The auxiliary storage device 128 can be realized by, for example, an HDD (Hard Disk Drive) or the like, and is a device that stores data and programs for executing various processes. The auxiliary storage device 128 may also be realized by an SSD (Solid State Drive) using a flash memory or the like.
[0039] The RAM 124, the ROM 125, and the auxiliary storage device 128 correspond to the database 6 in FIG. 1A. The auxiliary storage device 128 stores the robot control program, various information, data, and the like.
[0040] Figure 3 is a flowchart showing the learning method of the first learning model 4 in Figure 2A. In the following explanation, the components of Figure 2A will be indicated with reference numerals.
[0041] First, the coordinate calculation unit 5 acquires data on the robot end-effector position 19 and the object position 20 (step S201). Similarly, the coordinate calculation unit 5 acquires data on the camera posture 21 (step S202). Here, this data includes the positions and shapes of the robot arm 8, the object to be operated 9, and the camera 10 in the virtual space obtained as the simulation environment 2. Then, the coordinate calculation unit 5 uses this data to calculate the coordinates 22 on the simulation image (step S203).
[0042] Next, the data of the image 24 from the simulation environment 2 is sent to the database 6 (step S204). Then, the coordinate calculation unit 5 sends the coordinates 22 on the simulation image to the database 6 and stores them together with the data of the image 24 (step S205).
[0043] Next, it is determined whether the manipulation task has been completed (step S206). If the manipulation task has been completed, the process proceeds to step S208.
[0044] On the other hand, if the manipulation task is not completed, the process proceeds to step S207. In this case, the user manually moves the robot arm 8 to a predetermined position using the robot input unit 14 (step S207). In this case, the process returns to step S201, and the processes from step S201 to step S207 are repeated. The processes from step S201 to step S207 are repeated until it is determined in step S206 that the manipulation task is completed. For the manual operation in step S207, the user switches the switch 15 to manual mode 16 and operates the robot input unit 14. The directional information generated by this manual operation is input to the robot control unit 18 and transmitted to the robot arm 8 as an operation command. This performs the movement process of the robot arm 8 in the simulation environment 2. The repeated processing from step S201 to step S207, which involves manual operation, is performed to acquire predetermined data.
[0045] If the manipulation task is completed, step S208 determines whether the number of acquired data points has reached a predetermined value. If the number of acquired data points has not reached the predetermined value, the simulation environment 2 is reset or modified (step S209), and the process returns to step S201.
[0046] The repetitive processing from step S201 to step S209 is repeated until the number of acquired data points reaches a predetermined value.
[0047] The processing from steps S201 to S209 makes it easy to create a training dataset for the first learning model 4, which extracts important positions from images for the robot arm 8 to accomplish the manipulation task.
[0048] In step S208, if the number of acquired data points reaches a predetermined value, the first model learning unit 7 uses the training dataset stored in the database 6 to train the first learning model 4. The first model learning unit 7 uses a machine learning algorithm to approximate the estimated position 23 of the first learning model 4 to the coordinates 22 (ground truth position data) on the simulation image in the training dataset when an input image 24 from the same training dataset is given (step S210). In other words, in step S210, the learning is performed to minimize the error between the coordinates of the estimated position 23 of the first learning model 4 and the coordinates 22 in the training dataset. This learning can be performed using a machine learning algorithm (such as backpropagation).
[0049] This paper summarizes the model training method and the method for creating the training dataset for the first trained model, as well as the training method for the first trained model to estimate important locations from images for the manipulation task.
[0050] Once the first learning model 4 has completed its training, the image 24 extracted from the simulation environment 2 is input to the first learning model 4 in order to automatically move the robot arm 8 in the simulation environment 2 and perform the manipulation task. The first learning model 4 estimates the relevant position (estimated position 23) in the image plane. The estimated position 23 is input to the trajectory generation unit 13. The robot end-effector position 19 is also input to the trajectory generation unit 13. The trajectory generation unit 13 generates a trajectory so that the target robot end-effector position is located on both sides of the object to be manipulated 9, and inputs this trajectory to the robot control unit 18 by passing it through the switch 15 of automatic mode 17. The robot control unit 18 outputs an operation command to the robot arm 8 and controls the robot so that the target robot end-effector position of the robot arm 8 moves along the generated trajectory. By repeating this process, the manipulation task can be performed.
[0051] To allow the user to qualitatively verify the validity of the first learning model 4, the display unit 26 may display the coordinates 22 on the simulation image. In this case, the coordinates 22 can be marked with a circle (○) and the estimated position 23 with a triangle (△), and the distance between the circle and the triangle can be checked. In this case, it is desirable that the circle and the triangle overlap. The results from the display unit 26 are shown in Figure 12.
[0052] Furthermore, once the first learning model 4 is trained, the object's position is no longer necessary. In real-world environments, it is difficult to directly determine the object's position as in the simulation, so the estimated position 23 in the image plane is necessary for trajectory generation.
[0053] The trajectory generation unit 13 is responsible for generating a trajectory based on the current robot end-effector position 19 and the estimated position 23 in the image plane. The method for generating the trajectory from these two input pieces of information involves calculating the actual position of the object in the three-dimensional environment from the estimated two-dimensional position 23 through a predetermined transformation.
[0054] If the conversion can be done directly using analytical formulas or tables, the resulting three-dimensional position and the current robot end-effector position 19 can be input into a motion planning algorithm such as a probabilistic roadmap (PRM) or a fast search random tree (RRT) to generate a trajectory. The robot arm 8 can then execute this trajectory according to motion commands from the robot control unit 18 and accomplish the manipulation task.
[0055] However, if a direct conversion between the two-dimensional estimated position in the image plane and the actual position of the object in the three-dimensional environment is not possible, an alternative approach may be to train a machine learning model that generates a trajectory that accomplishes the manipulation task of the robot arm 8 based on the two-dimensional estimated position and the robot end-effector position 19, which are estimated from both the robot end-effector position and the object position. For example, imitation learning can be applied as a machine learning algorithm.
[0056] Next, we will explain, using Figure 4, how to calculate the coordinates 22 on the simulation image based on the three-dimensional positions such as the robot end-effector position 19 and object position 20 obtained from the simulation environment 2.
[0057] Figure 4 is a perspective view showing the coordinates on a simulation image generated by the robot control device according to this embodiment, the coordinate system necessary for calculating those coordinates, camera parameters, etc.
[0058] In this figure, the robot arm 8 and the object to be manipulated 9, which have three-dimensional coordinates, are shown in the simulation environment 2 of Figure 2A. The coordinate systems used here are the world coordinate system W, which represents the three-dimensional coordinates of the robot arm 8 and the object to be manipulated 9, and the camera coordinate system C, which has the camera's focus point Fc as its origin.
[0059] In the world coordinate system W, the coordinates P of the robot arm 8's end-effector position 19 are 19 w and the coordinates P of the object position 20, which is the representative position of the object 9 to be operated on. 20 w These are each set. Below, the coordinates of these positions are summarized as P w It is called that.
[0060] On the other hand, the robot arm 8 and the object to be manipulated 9 in the simulation environment 2 are projected onto the image plane 301. The focus point Fc of the camera is located at a position of the focal length f from the image plane 301 along the optical axis 302.
[0061] To obtain the coordinates 22 (FIG. 2A) on the simulation image in the image plane 301, first, the position P w needs to be converted into the position P in the camera coordinate system C using the external camera matrix T W C C Here, "conversion" means changing the position P C into a coordinate system within the camera coordinate system. This conversion can be calculated by the following formula (1). In the formula, r ij (i = 1, 2, 3, j = 1, 2, 3) is a rotation matrix that rotates the coordinates of P w from the world coordinate system W to the camera coordinate system C. t k (k = 1, 2, 3) can be used not only for the conversion of P w between the same coordinate systems but also for translation (parallel movement).
[0062]
[0063] After obtaining the position P C in the camera coordinate system C by the above formula (1), these positions are projected onto the image plane 301 to obtain the coordinates 22 on the simulation image, that is, the coordinates (u, v). This can be obtained using the following formula (2). In the formula, ρ u , ρ v represent the pixel width and height, respectively. These are measured in mm / pixel. Furthermore, c x , c y typically indicate the principal point located at the center of the image.
[0064]
[0065] Therefore, the coordinates 22 in the image plane 301 can be acquired in a form corresponding to the robot end position 19 and the object position 20. These coordinates 22 can be used as the correct data for training the first learning model 4.
[0066] Next, a method for applying the first learning model 4 obtained in the first model learning device 1 shown in Figure 2A to the real world (real space) will be explained using Figures 5 and 6.
[0067] The first trained model 4, which was trained in the simulation environment 2 shown in Figure 2A, may be usable in the real world without adjustment, but only if the robot arm 8 and the object 9 used in the simulation environment 2 are similar to those in the real world. In this case, it is assumed that the simulation parameters, such as image conditions, are very similar to the parameters of the real world. However, in many cases this is not the case, so a different approach is needed. By first measuring the parameter distribution of the real world and randomizing the simulation parameters within that distribution, it is possible to train a robust model that can effectively respond to various changes in the real world.
[0068] Figure 5 is a functional block diagram illustrating this scenario. Specifically, Figure 5 is a functional block diagram showing a configuration in which a robot arm or the like in a real environment is operated using the first learning model 4 of Figure 2A.
[0069] Figure 5 shows the second model learning device 404 and control device 410 that constitute the robot control device, as well as the real environment 401 (real space) that is the target of control. The second model learning device 404 includes a parameter adjustment unit 403, a simulator 422, a second model learning unit 405 (second learning model generation unit), and a display unit 26.
[0070] The parameter adjustment unit 403 receives data from the real environment 401 and calculates the range for randomizing the simulation parameters in that data. The randomization range is determined according to the range of change of the components of the work. Examples of components include the work entity such as a robot, the work object such as an object to be grasped, and the work environment in which the robot works.
[0071] More specifically, for example, in the case of a work object, the range of randomization is calculated according to the magnitude of the change between the position, orientation, and type of the work object in the actual work and the learned position, orientation, and type of the work object.
[0072] While the location, orientation, and type of the object being worked on are given as examples, at least one of these may suffice.
[0073] Furthermore, in the case of a work environment, the range of randomization is calculated based on the degree of change between the brightness, saturation, and background pattern of the actual work environment and the brightness, saturation, and background pattern of the learned work environment.
[0074] This range is then sent to the simulator 422. Specifically, the parameter adjustment unit 403 acquires data of the robot end-effector position 19 from the robot arm 8 in the real environment 401, and image 24 (real-space image) data from the camera 10 in the real environment 401. Here, the robot arm 8 may directly calculate the robot end-effector position 19 etc. using kinematics with respect to the joint angles obtained from the motor encoders of each joint.
[0075] The simulator 422 generates a training dataset (second teaching data). The training dataset, along with the pre-trained first training model 4 and its weights, is sent to the second model learning unit 405. Here, the weights of the first training model 4 are also parameters within the first training model 4. The second model learning unit 405 creates a trained model. The data of the trained model becomes the second training model 402. The second model learning unit 405 transmits the estimated position 23 (second estimated position) and image 24 of the robot arm 8, etc., to the display unit 26. The display unit 26 displays the estimated position 23, image 24, etc., on the screen. Here, the estimated position 23 is estimated by the second model learning device 404 based on the image 24 of the real environment 401.
[0076] Thus, the simulator 422 can estimate the coordinates even when a predetermined part of the image of the robot arm 8 or the object to be manipulated 9 is hidden in real space.
[0077] The control device 410 includes a trajectory generation unit 13 and a robot control unit 18.
[0078] The trajectory generation unit 13 acquires the estimated position 23 and generates a trajectory for the robot arm 8, etc. The robot control unit 18 acquires the data of that trajectory and issues operation commands for the robot arm 8, etc.
[0079] Figure 6 is a flowchart showing the learning process of the second learning model of the robot control device according to this embodiment and the generation of robot motion in a real environment.
[0080] In this figure, the second model learning unit 405 uses the first learning model 4 and the teaching data created by the simulator 422 using data from the real environment 401 to create a second learning model 402 so that it can estimate the position of the object to be operated 9 in the real environment 401 (step S501). Specifically, the second learning model 402 is created so that it can estimate the position of the object to be operated 9 in the real image plane included in the data of the real environment 401.
[0081] Next, the second model learning device 404 estimates the estimated position 23 using the image 24 of the real environment 401 based on the second learning model 402. Then, the trajectory generation unit 13 generates a trajectory based on the estimated position 23 (step S502).
[0082] Then, the robot control unit 18 acquires trajectory data from the trajectory generation unit 13 and transmits an operation command to the robot arm 8, etc. (step S503).
[0083] Using the robot control device 100 shown in Figure 1, a first learning model 4 is generated from the simulation environment 2, and a second learning model 402 is generated from the real environment 401. This allows the estimated position 23 of the robot arm 8, etc., to be estimated from the image 24 of the real environment 401. Furthermore, the estimated position 23 can be used to accomplish a manipulation task in the real environment 401. Additionally, by acquiring a learning dataset while changing the simulation conditions and training the first learning model 4 and the second learning model 402, the manipulation task in the real environment 401 can be executed more robustly.
[0084] In this embodiment, a manipulation task of the robot arm 8 is described, but the robot control device according to this disclosure is not limited to this and can be applied to various manipulation tasks. In particular, important positions can be selected in advance and subsequent position changes can be automatically randomized, making it easy to create a dataset.
[0085] In this embodiment, the robot end-effector position 19 and the object position 20 were selected as simulation positions because they are important positions related to whether or not the object can be grasped. By setting the robot end-effector position 19 and the object position 20 as simulation positions, it is possible to learn to predict positions that are important for the task, thereby improving learning performance and shortening learning time. However, since the appropriate positions differ depending on the manipulation task, the learning method of this disclosure requires a human to select the positions to be used for each manipulation task.
[0086] Figure 7 is a flowchart that summarizes the learning of robot motion and the generation of robot motion in a real environment by the robot control device according to this embodiment.
[0087] In this figure, steps S601 to S603 summarize the contents of the flowchart in Figure 3 in three steps, and steps S604 to S606 are the same as steps S501 to S503 in Figure 6.
[0088] Specifically, first, the coordinate calculation unit 5 acquires the image 24 and the position of the object to be manipulated 9 from the simulation environment 2 (step S601). Next, the coordinate calculation unit 5 calculates the coordinates of a predetermined part on the image (coordinates 22 on the simulation image) from the position of the object to be manipulated 9 (step S602). Next, the first model learning unit 7 learns the first learning model 4 so that it can estimate the calculation position (coordinates 22 on the simulation image) from the image 24 (Figure 2A) in the simulation environment 2 (step S603).
[0089] Next, the second model learning unit 405 learns the second learning model 402 so that the first learning model 4 can estimate the calculation position from the image 24 of the real environment 401 (step S604).
[0090] Next, the second model learning device 404 estimates the position 23 from the image 24 of the real environment 401 based on the second learning model 402. Then, the trajectory generation unit 13 generates a trajectory based on the estimated position 23 (step S605).
[0091] Then, the robot control unit 18 acquires trajectory data from the trajectory generation unit 13, sends an operation command to the robot (object to be operated 9), and executes the task (step S606).
[0092] Regarding the robot motion learning portion of the robot control device in this embodiment, it consists only of learning a learning model that can obtain the position of the image plane important for manipulation from an image, but it can also include a motion generation model that generates the robot's trajectory from the estimated position of the image plane and the robot's end-effector position.
[0093] Figure 8 shows the variations in three-dimensional position obtained from the simulator of the robot control device according to this embodiment.
[0094] The three-dimensional position shown in this figure can be used as input for calculating the two-dimensional position in the image plane.
[0095] This figure shows variations in the robot end-effector position, the position of the object being manipulated, and the camera orientation.
[0096] The robot end-effector position 19 in Figure 2A provides positional information, but when combined with positions 701 and 702, it can also provide information about the robot end-effector's angle. Furthermore, positions 703 and 704 also provide information about the state of the robot's fingers (open or closed). Similarly, the object position 20 in Figure 2A provides positional information, but when combined with positions 705 and 706, it can also provide information about the angle of the object being manipulated 9.
[0097] Furthermore, Figure 8 shows the change in the posture 716 of the camera 10. During the movement of the robot arm 8 (Figure 2A) to perform the manipulation task, the two-dimensional calculated positions calculated from, for example, positions 701, 702, 703, 704, 705, and 706 overlap, which can reduce the effectiveness of the two-dimensional calculated positions and lower the success rate of the manipulation task. In such cases, the effect of changing the posture 716 of the camera 10 can be confirmed by changing the posture 716 of the camera 10, recalculating the coordinates 22 on the simulation image (Figure 2A), and then using these to re-execute the manipulation task with the robot. By repeating these steps using an optimization algorithm, the overall task performance can be improved.
[0098] Figure 9 is a flowchart showing the camera orientation optimization algorithm for the robot control device according to this embodiment.
[0099] In this diagram, first, the average angle between the end-effector of the robot arm 8 and the object being manipulated is calculated (step S801). Then, the camera angle is adjusted so that it is perpendicular to this average angle (step S802). Next, the camera position is adjusted to a distance where both the end-effector of the robot arm 8 and the object can be seen (step S803). Then, the manipulation task is executed (step S804). The average angle between the robot arm end-effector and the object being manipulated is calculated again (step S805).
[0100] Next, the new camera pose is determined (step S806). Then, the difference between the camera pose before and after the pose change is calculated (step S807). If the difference between the camera poses is greater than a preset threshold (step S808), the process returns to step S804.
[0101] The data acquisition method, learning method, trajectory generation method, and robot control method of this embodiment enable the robot to efficiently perform manipulation tasks under diverse environmental conditions.
[0102] Example 2 will be explained using Figures 10 to 12.
[0103] Figure 10 is a functional block diagram showing the simulation environment generated by the robot control device of this embodiment.
[0104] This figure shows the first model learning device 1, the simulation environment 900, and the control device 3.
[0105] The differences between this embodiment and Embodiment 1 are two points: the object being manipulated in the simulation environment 900 and the position of the camera 10.
[0106] In Example 2, the object being manipulated is a T-shirt 901 that deforms during the manipulation. The camera 10 in the simulation environment 900 is positioned above the T-shirt 901 in the simulation environment 900. The T-shirt 901 has gripping positions 902, 903, 904, and 905. In this example, only the gripping positions 902, 903, 904, and 905 of the T-shirt 901 are converted to coordinates 22 on the simulation image in the image plane 301 (Figure 4). This is because it is difficult to observe the end effector of the robot arm because the camera 10 is above.
[0107] Therefore, the trajectory generation unit 13 needs to estimate the gripping positions 902, 903, 904, and 905 from the estimated position 23 estimated by the first learning model 4. Subsequently, the trajectory generation unit 13 generates a gripping motion using a trajectory generation algorithm based on these estimated object positions. Aside from these differences, this embodiment follows the same learning method as Embodiment 1 and enables the manipulation of objects.
[0108] Figure 11A shows an example of an image 24 of the T-shirt 1001 obtained from the camera 10 in the simulation environment 900 of Figure 10.
[0109] In Figure 11A, the T-shirt 1001 is laid flat without any folds and is in an unfolded state. The two-dimensional calculation areas 1002, 1003, 1004, and 1005 corresponding to the gripping positions 902, 903, 904, and 905 of the T-shirt 901 in Figure 10 are shown.
[0110] Figure 11B shows another example of an image 24 of the T-shirt 1001 obtained from the camera 10 in the simulation environment 900 of Figure 10.
[0111] In Figure 11B, the T-shirt 1001 is folded at the left shoulder portion in the figure.
[0112] As shown in Figure 11A, when the T-shirt 1001 is stretched, the calculation areas 1002, 1003, 1004, and 1005 can be observed without interference. However, as shown in Figure 11B, when one sleeve of the T-shirt 1001 is partially folded so that it is located behind the main part of the T-shirt, the usual approach to automatically detect the position to be grasped cannot estimate the calculation areas 1002 and 1003 due to the blind spot. As a result, the manipulation task cannot be performed.
[0113] However, as in this embodiment, since the gripping positions 902 and 903 can be obtained from the simulation environment 2, the calculation units 1002 and 1003 can be calculated. The information obtained by the calculation units 1002 and 1003 can be used to generate motions that reveal the gripping positions 902 and 903, making manipulation tasks such as gripping and folding a T-shirt easier.
[0114] Figure 12 is an image 24 showing the calculation units 1002, 1003, 1004, and 1005 shown on the display unit 26 of Figure 10 together with the estimated positions 1102, 1103, 1104, and 1105.
[0115] In Figure 12, the estimated positions 1102, 1103, and 1105 are close to the calculation areas 1002, 1003, and 1005. Therefore, it can be concluded that the estimated positions 1102, 1103, and 1105, which are the estimation results for the calculation areas 1002, 1003, and 1005 in the first learning model 4 in Figure 10, are valid.
[0116] On the other hand, the estimated position 1104 is far from the calculation unit 1004. Therefore, it can be concluded that the estimated position 1104, which is the estimation result of the calculation unit 1004 in the first learning model 4 in Figure 10, is not valid.
[0117] To numerically determine whether the first learning model 4 is valid, for example, the mean of the distance between the estimated positions 1102, 1103, 1104, and 1105 and the corresponding computation units 1002, 1003, 1004, and 1005 can be calculated, and the standard deviation calculated at this time can be used.
[0118] According to this embodiment, even if a predetermined calculation area is in the camera's blind spot, it is possible to calculate the estimated position corresponding to that calculation area and perform the manipulation task. Such estimation is particularly effective when dealing with flexible and easily deformable objects (easily deformable objects).
[0119] The embodiments of this disclosure will be described below in summary.
[0120] The robot control device comprises: a coordinate calculation unit that calculates coordinates on a simulation image from data of an object and a robot manipulator in a simulation environment; a first learning model generation unit that generates a first learning model including data of a first estimated position of the object and the robot manipulator in the simulation environment using first teaching data including coordinates on the simulation image; a second learning model generation unit that generates a second learning model including data of a second estimated position of the object and the robot manipulator in real space based on second teaching data generated by predetermined processing using data of the object and the robot manipulator in real space, the first learning model and the weights of the first learning model; and a robot control unit that controls the movement of the robot manipulator based on the data of the second estimated position.
[0121] The first estimated position included in the first learning model, specifically the position estimated for the robot manipulator, is the part of the robot manipulator that handles objects.
[0122] The first estimated position included in the first learning model, specifically the position estimated for the robot manipulator, is a predetermined location between the two fingers of the robot hand of the robot manipulator.
[0123] Of the first estimated positions included in the first learning model, the position estimated for the robot manipulator is the fingertips of the robot hand of the robot manipulator.
[0124] The first estimated position included in the first learning model, which is estimated for the robot manipulator, is a predetermined part of at least two sides of the robot hand of the robot manipulator.
[0125] The first estimated position included in the first learning model, specifically the position estimated for an object, corresponds to the center of gravity of the object's three-dimensional shape.
[0126] The object is easily deformable.
[0127] The prescribed process calculates the range within which the simulation parameters in the data for the object and robot manipulator in real space are randomized, and generates second teaching data within this range.
[0128] The first learning model generation unit learns to minimize the error between the coordinates of the first estimated position of the first learning model and the coordinates on the simulation image of the first training data.
[0129] The robot motion learning method involves a coordinate calculation unit calculating coordinates on a simulation image from data of an object and a robot manipulator in a simulation environment; a first learning model generation unit generating a first learning model that includes data of the first estimated positions of the object and the robot manipulator in the simulation environment using first teaching data that includes coordinates on the simulation image; and a second learning model generation unit generating a second learning model that includes data of the second estimated positions of the object and the robot manipulator in real space, based on second teaching data generated by predetermined processing using data of the object and the robot manipulator in real space, the first learning model, and the weights of the first learning model.
[0130] The robot motion learning device comprises: a coordinate calculation unit that calculates coordinates on a simulation image from data of an object and a robot manipulator in a simulation environment; a first learning model generation unit that generates a first learning model including data of a first estimated position of the object and the robot manipulator in the simulation environment using first teaching data including coordinates on the simulation image; and a second learning model generation unit that generates a second learning model including data of a second estimated position of the object and the robot manipulator in real space based on second teaching data generated by predetermined processing using data of the object and the robot manipulator in real space, the first learning model, and the weights of the first learning model.
[0131] 1, 151: First model learning device, 2, 155, 900: Simulation environment, 3: Control device, 4: First learning model, 5: Coordinate calculation unit, 6: Database, 7: First model learning unit, 8: Robot arm, 9: Object to be operated, 10: Camera, 13: Trajectory generation unit, 14: Robot input unit, 15: Switch, 16: Manual mode, 17: Automatic mode, 18: Robot control unit, 19: Robot end-effector position, 20: Object position, 21: Camera orientation, 22: Coordinates, 23: Estimated position, 24: Image, 26: Display unit, 100: Robot control device, 102: Robot control device, 121: Display, 122: Input device, 123: CPU, 124: RAM ,125: ROM, 126: Communication device, 127: Media reader, 128: Auxiliary storage device, 152, 404: Second model learning device, 153: Robot control unit, 154: Robot, 301: Image plane, 302: Optical axis, 401: Real environment, 402: Second learning model, 403: Parameter adjustment unit, 405: Second model learning unit, 410: Control device, 422: Simulator, 701, 702, 703, 704, 705, 706: Position, 716: Posture change, 901, 1001: T-shirt, 902, 903, 904, 905: Gripping position, 1002, 1003, 1004, 1005: Calculation unit, 1102, 1103, 1104, 1105: Estimated position.
Claims
1. A robot control device comprising: a coordinate calculation unit that calculates coordinates on a simulation image from data of an object and a robot manipulator in a simulation environment; a first learning model generation unit that generates a first learning model including data of a first estimated position of the object and the robot manipulator in the simulation environment using first teaching data including the coordinates on the simulation image; a second learning model generation unit that generates a second learning model including data of a second estimated position of the object and the robot manipulator in the real space based on second teaching data generated by predetermined processing using data of the object and the robot manipulator in real space, the first learning model and the weights of the first learning model; and a robot control unit that controls the operation of the robot manipulator based on the second learning model.
2. The robot control device according to claim 1, wherein the position estimated for the robot manipulator among the first estimated positions included in the first learning model is the part of the robot manipulator that handles the object.
3. The robot control device according to claim 1, wherein the position estimated for the robot manipulator among the first estimated positions included in the first learning model is a predetermined part between the two fingers of the robot hand of the robot manipulator.
4. The robot control device according to claim 1, wherein the position estimated for the robot manipulator among the first estimated positions included in the first learning model is the fingertip of the robot hand of the robot manipulator.
5. The robot control device according to claim 1, wherein the position estimated for the robot manipulator among the first estimated positions included in the first learning model is a predetermined portion of at least two side portions of the robot hand of the robot manipulator.
6. The robot control device according to claim 1, wherein the position estimated for the object among the first estimated positions included in the first learning model is a part corresponding to the center of gravity of the three-dimensional shape of the object.
7. The robot control device according to claim 1, wherein the object is easily deformable.
8. The robot control device according to claim 1, wherein the predetermined processing involves calculating a range for randomizing the simulation parameters in the data of the object and the robot manipulator in the real space, and generating the second teaching data within this range.
9. The robot control device according to claim 1, wherein the first learning model generation unit learns to minimize the error between the coordinates of the first estimated position of the first learning model and the coordinates of the first teaching data on the simulation image.
10. A robot motion learning method comprising: a coordinate calculation unit calculates coordinates on a simulation image from data of an object and a robot manipulator in a simulation environment; a first learning model generation unit generates a first learning model including data of the first estimated positions of the object and the robot manipulator in the simulation environment using first teaching data including the coordinates on the simulation image; and a second learning model generation unit generates a second learning model including data of the second estimated positions of the object and the robot manipulator in real space based on second teaching data generated by predetermined processing using data of the object and the robot manipulator in real space, the first learning model, and the weights of the first learning model.
11. A robot motion learning device comprising: a coordinate calculation unit that calculates coordinates on a simulation image from data of an object and a robot manipulator in a simulation environment; a first learning model generation unit that generates a first learning model including data of a first estimated position of the object and the robot manipulator in the simulation environment using first teaching data including the coordinates on the simulation image; and a second learning model generation unit that generates a second learning model including data of a second estimated position of the object and the robot manipulator in the real space based on second teaching data generated by predetermined processing using data of the object and the robot manipulator in real space, the first learning model and the weights of the first learning model.
Citation Information
Patent Citations
Robot control system and robot control method
JP2017094482A
Learning device, control device, learning method, and learning program
JP2020057161A
Control device, control method and control program
JP2021035714A
Generation method for dataset for machine learning, learning model generation method, learning data generation device, inference device, robotic operation controlling device, trained model, and robot
JP2021047839A
Calibration system, information processing system, robot control system, calibration method, information processing method, robot control method, calibration program, information processing program, calibration device, information processing device, and robot control device
JP2021160037A