Control device, learning device, and training data generation device
The control device employs a machine learning-based inference model to assess gripping quality and feasibility, addressing the challenge of high-quality gripping by a robot hand and achieving robust and precise gripping operations.
Patent Information
- Application Number
- JP2023205154
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-17
AI Technical Summary
Existing technologies face challenges in estimating whether a gripping device can perform high-quality gripping, as they may not accurately determine suitable gripping positions.
A control device that utilizes an inference model generated by machine learning to calculate a score indicating the quality of gripping and a mask value indicating whether gripping can be executed, allowing for high-quality gripping by a robot hand.
Enables high-quality gripping by accurately determining suitable gripping positions and executing gripping operations with improved robustness and precision.
Smart Images

Figure 2025090123000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a control device, a learning device, and a training data generation device.
Background Art
[0002] Patent Document 1 discloses an estimation device. The estimation device according to Patent Document 1 inputs information about a target object into a neural network model that outputs information about the position and orientation of a gripping device capable of gripping the target object, and estimates information about the position and orientation at which the gripping device can grip the target object. Non-Patent Document 1 discloses a neural network that takes a TSDF (Truncated Signed Distance Function) volume as input and outputs a grasping success probability and a grasping direction for each voxel.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Non-Patent Documents
[0004]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] In the technology according to Patent Document 1, there is a possibility that it is not possible to estimate whether the position of the grippable gripping device is a position suitable for gripping. In other words, in the technology according to Patent Document 1, there is a possibility that it is not possible to estimate whether the gripping device can actually perform high-quality gripping at the position of the grippable gripping device.
[0006] The present disclosure provides a control device, a learning device, and a training data generation device that enable high-quality gripping using a robot hand.
Means for Solving the Problems
[0007] The control device according to the present disclosure uses an inference model generated in advance by machine learning to calculate, for each position in the three-dimensional space, a score indicating the quality of gripping when the robot hand grips an object with respect to the position of the reference point of the robot hand in the three-dimensional space, and a mask value indicating whether gripping can be executed. An estimation unit that estimates, and a hand control unit that controls the robot hand based on the score for positions where the mask value indicates that gripping can be executed.
Advantages of the Invention
[0008] According to the present disclosure, it is possible to provide a control device, a learning device, and a training data generation device that enable high-quality gripping using a robot hand.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Embodiments for Carrying Out the Invention
[0010] Hereinafter, this embodiment will be described with reference to the drawings. However, the present invention is not limited to the following embodiments. Also, for clarity of explanation, the following description and drawings are simplified as appropriate.
[0011] (Embodiment 1) FIG. 1 is a diagram showing the configuration of a gripping system 1 according to Embodiment 1. The gripping system 1 includes one or more detection devices 2, a robot hand 10, and a control system 100. The control system 100 is connected to the detection device 2 and the robot hand 10 via a wireless or wired communication network.
[0012] The gripping system 1 estimates a score and a mask value for each position in the three-dimensional space with respect to the position of the reference point of the robot hand 10 in the three-dimensional space using an inference model generated in advance by machine learning. The gripping system 1 executes gripping using the estimated score and mask value. Here, the inference model can be learned by, for example, a neural network. Also, the score indicates the quality of gripping when the robot hand 10 grips an object. The better the quality of gripping, the more firmly the robot hand 10 can grip the object. Also, the mask value indicates whether the robot hand 10 can execute gripping. Thereby, the gripping system 1 according to Embodiment 1 can achieve high-quality gripping using the robot hand.
[0013] The detection device 2 detects each position in the three-dimensional space where the object is placed. In other words, the detection device 2 detects (captures, measures) the position where the object placed in the three-dimensional space exists. The detection device 2 detects whether an object exists at each position in the three-dimensional space. The detection device 2 is, for example, a three-dimensional camera such as an RGB-D camera or a stereo camera, a depth camera, or LiDAR (Light Detection And Ranging), but is not limited thereto. Also, in the present embodiment, the position in the three-dimensional space is represented by voxels, but is not limited to this.
[0014] The robot hand 10 is configured to grip an object placed in the three-dimensional space. The operation of the robot hand 10 is controlled by the control system 100. That is, the robot hand 10 grips the object by the control of the control system 100. The robot hand 10 can be an end effector provided at the tip of a robot arm (not shown). The robot hand 10 is, for example, an underactuated hand, but is not limited thereto.
[0015] FIG. 2 is a diagram illustrating the robot hand 10. The robot hand 10 includes a hand body 12, two finger portions 14, a plurality of links 16, and a plurality of joint portions 18. The finger portions 14 are connected to the hand body 12 via the plurality of links 16 and joint portions 18. When one or more of the joint portions 18 are driven, the finger portions 14 operate. Here, when the robot hand 10 is an underactuated hand, not all of the joint portions 18 are driven, but some of the plurality of joint portions 18 are driven. A driving device such as a motor is incorporated in the drivable joint portions 18. As will be described later, the robot hand 10 illustrated in FIG. 2 can perform precision grasp and power grasp. Also, the robot hand 10 can take a gripping posture with six degrees of freedom in the three-dimensional space.
[0016] Here, a reference point Pr is set for the robot hand 10. The reference point Pr is also called the TCP (Tool Center Point). The reference point Pr is the origin of the hand coordinate system (x, y, z). The +x direction is the direction in which the robot hand 10 approaches the object. Also, the y direction is the direction along which the finger part 14 moves (opens and closes). Further, the z direction, which is perpendicular to the xy plane, is the direction of the normal vector of the plane where the finger part 14 moves. Note that the reference point Pr can be arbitrarily determined. In the example of FIG. 2, the reference point Pr is provided near the center in the y direction on the front surface 12a of the hand body 12, but it is not limited to this. Note that the position of the reference point Pr can be varied to perform learning of the inference model described later, and the position of the reference point Pr can be adjusted so that the accuracy of the output of the inference model improves (robust grasping can actually be performed at a position with a high score in the real space). The reference point Pr may be a position outside the hand body 12 or a position inside the hand body 12.
[0017] FIG. 3 is a diagram for explaining the grasping operation of the robot hand 10 according to the first embodiment. (a) of FIG. 3 shows a state in which the robot hand 10 is grasping the object W by precise grasping. (b) of FIG. 3 shows a state in which the robot hand 10 is grasping the object W by inclusive grasping. As shown in (a), in the state of precise grasping, the robot hand 10 grasps the object W at two points by sandwiching it with the two finger parts 14 at a position slightly separated from the hand body 12. As shown in (b), in the state of inclusive grasping, the robot hand 10 grasps the object W in a state where three points, namely, the two finger parts 14 and the front surface 12a of the hand body 12, are in contact at a position close to the hand body 12. Here, compared with inclusive grasping, in precise grasping, the object is more likely to fall off from the finger part 14 due to an external force (such as gravity) applied to the object. Therefore, generally, inclusive grasping can grasp the object W more robustly than precise grasping. Therefore, the quality of grasping can be higher in inclusive grasping than in precise grasping.
[0018] The control system 100 is a computer such as a server, for example. The control system 100 can be realized by, for example, cloud computing. Also, the control system 100 can be realized by a plurality of computers. In this case, the plurality of components of the control system 100 described later may be realized by physically different computers, respectively.
[0019] The control system 100 has, as a main hardware configuration, a control unit 102, a storage unit 104, a communication unit 106, and an interface unit 108 (IF; Interface). The control unit 102, the storage unit 104, the communication unit 106, and the interface unit 108 are interconnected via a data bus or the like. Note that when the control system 100 is realized by a plurality of computers, each of the plurality of computers may have the hardware configuration shown in FIG. 1.
[0020] The control unit 102 is a processor such as a CPU (Central Processing Unit), for example. The control unit 102 has a function as an arithmetic unit that performs control processing, arithmetic processing, and the like. Note that the control unit 102 may have a plurality of processors. The storage unit 104 is a storage device such as a memory or a hard disk, for example. The storage unit 104 is, for example, a ROM (Read Only Memory) or a RAM (Random Access Memory). The storage unit 104 has a function of storing a control program, an arithmetic program, and the like executed by the control unit 102. That is, the storage unit 104 (memory) stores one or more instructions. Also, the storage unit 104 has a function of temporarily storing processing data and the like. The storage unit 104 may include a database. Also, the storage unit 104 may have a plurality of memories.
[0021] The communication unit 106 performs processes necessary for communicating with other devices via a network. The communication unit 106 may include a communication port, a router, a firewall, etc. The interface unit 108 is, for example, a user interface (UI). The interface unit 108 has an input device such as a keyboard, a touch panel, or a mouse, and an output device such as a display or a speaker. The interface unit 108 may be configured such that the input device and the output device are integrated, for example, like a touch panel. The interface unit 108 receives an operation of inputting data by a user and outputs information to the user.
[0022] The control system 100 includes a training data generation device 120, a learning device 130, and a control device 140. The training data generation device 120, the learning device 130, and the control device 140 may be physically separate devices. In this case, each of the training data generation device 120, the learning device 130, and the control device 140 has the above-described hardware configuration. Also, two or more of the training data generation device 120, the learning device 130, and the control device 140 may be physically the same device. For example, the function of the training data generation device 120 may be incorporated into the learning device 130.
[0023] The training data generation device 120 includes, as components, a simulation execution unit 122 and a training data generation unit 124. The learning device 130 includes, as components, a training data acquisition unit 132 and a learning unit 134. The control device 140 includes, as components, a position acquisition unit 142, an estimation unit 144, and a hand control unit 146.
[0024] Each of the above-described components can be realized, for example, by causing a program to be executed under the control of the control unit 102. More specifically, each component can be realized by the control unit 102 executing a program (instruction) stored in the storage unit 104. Further, by recording a necessary program on an arbitrary non-volatile recording medium and installing it as needed, each component may be realized. Also, each component is not limited to being realized by software based on a program, and may be realized by any combination of hardware, firmware, and software, etc. Further, each component may be realized using a user-programmable integrated circuit such as, for example, an FPGA (field-programmable gate array) or a microcomputer. In this case, a program composed of the above-described components may be realized using this integrated circuit.
[0025] FIG. 4 is a flowchart showing the processing executed by the gripping system 1 according to the first embodiment. The processing in S120 shows the training data generation method by the training data generation device 120. The processing in S130 shows the learning method by the learning device 130. The processing in S142 to S146 shows the control method by the control device 140.
[0026] The training data generation device 120 generates training data by physical simulation (step S120). Specifically, the simulation execution unit 122 executes a physical simulation so that the robot hand 10 grips an object in a virtual three-dimensional space and the object is gripped in the virtual three-dimensional space where the object is arranged. The simulation execution unit 122 may execute, for example, the physical simulation according to Non-Patent Document 1, but is not limited thereto.
[0027] The training data generation unit 124 generates training data used to generate an inference model by performing machine learning using a physical simulation to be executed. Specifically, the training data generation unit 124 generates training data indicating a score and a mask value for each position in the three-dimensional space realized by the physical simulation when the reference point Pr of the robot hand 10 exists. In the first stage, the training data generation unit 124 uses the physical simulation to acquire various gripping postures for various objects and calculates the raw data of the score for each gripping posture. Also, in the second stage, the training data generation unit 124 calculates the score for each gripping posture for each object in the three-dimensional space where the object is arranged using the physical simulation.
[0028] Specifically, in the first stage, the training data generation unit 124 acquires a plurality of shapes represented by the three-dimensional meshes of various objects. The training data generation unit 124 samples the surface of the three-dimensional mesh for each object and acquires a set of pairs of opposing points. The training data generation unit 124 determines the precision gripping posture of the robot hand 10 such that each pair of opposing points becomes the contact points of the two finger parts 14 respectively. The training data generation unit 124 determines a plurality of precision gripping postures by rotating the robot hand 10 about the line connecting the pair of contact points as an axis. The training data generation unit 124 determines the enveloping gripping posture by bringing the robot hand 10 closer to the object by a certain distance for each of the determined precision gripping postures. Here, the accuracy of the gripping posture determined as described above may not be good. Therefore, the training data generation unit 124 corrects the gripping posture by physical simulation. Specifically, in the physical simulation, the training data generation unit 124 arranges the robot hand 10 with the finger parts 14 opened in a floating state of the object so as to match the determined gripping posture. The training data generation unit 124 closes the finger parts 14 of the robot hand 10 from that state and performs the simulation until the object stops, and sets the posture when the object stops as the gripping posture. Thereby, the accuracy of the gripping posture is increased.
[0029] Also, in the first stage, the training data generation unit 124 adds an external force to the object in a certain direction in the virtual space while the robot hand 10 holds the object, and increases the magnitude of the external force. Then, the training data generation unit 124 uses the magnitude of the external force when the object detaches from the robot hand 10 as the original data for the score. The training data generation unit 124 adds the external force in the same way for each of the plurality of directions, and obtains the original data for the score for each of the plurality of directions. In this way, in the first stage, the training data generation unit 124, through physical simulation, obtains the maximum external force f that can maintain the grip for each of the plurality of directions i when an external force is applied to the gripped object in a plurality of directions. i as the original data for the score. The training data generation unit 124 performs the above processing for each of the plurality of gripping postures for each individual object.
[0030] In the second stage, the training data generation unit 124 obtains the score by mapping the original data corresponding to each direction obtained in the first stage to the gravity direction in the virtual space in the virtual space where the objects are randomly arranged, which is realized by physical simulation. Specifically, the training data generation unit 124 arranges a plurality of objects obtained in the first stage in a three-dimensional virtual space by physical simulation to generate a plurality of scenes. The training data generation unit 124 selects a gripping posture in which the object is gripped in the posture of the object arranged in the generated scene from the gripping postures obtained in the first stage and projects it onto the generated scene.
[0031] FIG. 5 is a diagram illustrating the physical simulation according to the first embodiment. Various objects W are arranged in the three-dimensional virtual space realized by the physical simulation. And in that virtual space, the robot hand 10 holds the arranged object W in various gripping postures. At this time, for each gripping posture, the position of the reference point Pr and the posture of the robot hand 10 in the scene (virtual space) are respectively determined. Note that gripping postures in which the robot hand 10 interferes with other objects or the floor can be excluded.
[0032] Also, the training data generation unit 124 calculates a one-dimensional score for each gripping posture projected onto the scene. Specifically, the training data generation unit 124 maps the maximum value of the external force in each of the plurality of directions obtained in the first stage (the raw data of the score) for each gripping posture to the gravity direction of the object being gripped, thereby calculating the score corresponding to each gripping posture. Therefore, the score corresponds to the maximum value of the external force that can maintain the grip when the robot hand 10 executes the grip against the external force applied to the object. Here, let the index for each direction of the external force applied to the object in the scene be i, and the unit vector indicating direction i in the scene be n i be. Also, let the maximum value of the external force applied in direction i (the raw data of the score) be f i be. Also, let the unit vector indicating the gravity direction in the scene be g. In this case, the score corresponding to that gripping posture is represented by Equation 1 below. Also, ε is a sufficiently small positive value. Therefore, in Equation 1, when n i ·g is 0 or less, that is, for the maximum value f of the external force in direction i in the horizontal direction and in the direction upward from the horizontal direction in the scene i is excluded.
Equation
[0033] FIG. 6 is a diagram for explaining the score according to Embodiment 1. FIG. 6 shows a state in which the robot hand 10 is gripping the object W in a scene (virtual space). Here, let the maximum value of the external force in the direction along the x direction (unit vector n2) of the robot hand 10 shown in FIG. 2 be f2 [N], and the maximum value of the external force in the y direction (unit vector n1) orthogonal to that direction be f1 [N]. Here, f1 > f2. In this case, the external force in the direction of gravity that generates n2f2 (the component force in the direction of n2) is f2 / (n2·g) [N], and the external force in the direction of gravity that generates n1f1 (the component force in the direction of n1) is f1 / (n1·g) [N]. In this case, when the external force in the direction of gravity increases and reaches f2 / (n2·g) [N], the component force in the n2 direction becomes f2 / (n2·g)*(g·n2) = f2 [N], and the object W will fall off in the n2 direction. Therefore, as shown in Equation 1, the score is preferably the minimum value of f i / (n i ·g) for i.
[0034] The training data generation unit 124 generates training data using, for each of the plurality of gripping postures in each scene (virtual space) as described above, the position of the reference point Pr in that gripping posture, the score corresponding to that gripping posture, and the posture data indicating that gripping posture. Here, the posture data includes the unit vector in the direction in which the robot hand 10 approaches the object when realizing the corresponding gripping posture (the direction when the x direction in FIG. 2 is projected onto the scene), and the normal vector of the plane in which the finger portion 14 moves (the plane when the xy plane in FIG. 2 is projected onto the scene). The posture data can be uniquely determined from the corresponding gripping posture.
[0035] The training data generation unit 124 uses the position data (TSDF volume) at each position in the three-dimensional space where the object is placed as the input data in the training data. Specifically, the training data generation unit 124 generates a TSDF volume for each voxel from a depth image obtained by photographing (rendering) a scene where an object is placed on a virtual space by physical simulation from a predetermined direction, and uses it as the input data in the training data. The TSDF volume indicates the distance from each voxel in the virtual space to the nearest object.
[0036] Also, the training data generation unit 124 uses the score, mask value, and pose data when the reference point Pr of the robot hand 10 exists at each position of the input data as the output data in the training data. Specifically, the training data generation unit 124 generates the mask value, score, and pose data corresponding to each voxel in the virtual space as the output data in the training data. In the output data, the mask value indicates a value (e.g., "1") indicating "true" when an object can be grasped (i.e., there is a grasping pose) when the reference point Pr exists at each position (voxel) in the virtual space. On the other hand, the mask value indicates a value (e.g., "0") indicating "false" when an object cannot be grasped (i.e., there is no grasping pose) when the reference point Pr exists at that position. Also, in the output data, the score is the value of the score shown in Equation 1 corresponding to the grasping pose when the reference point Pr exists at each position (voxel) in the virtual space. Also, in the output data, the pose data corresponds to the grasping pose when the reference point Pr exists at each position (voxel) in the virtual space.
[0037] The learning device 130 generates an inference model by machine learning (step S130). Specifically, the training data acquisition unit 132 acquires the training data generated by the training data generation device 120. The learning unit 134 learns the inference model so as to input the input data in the training data and output the output data in the training data by executing machine learning. Thereby, the learning unit 134 generates a learned inference model. The inference model can be realized by a neural network such as an FCN (Fully Convolutional Network) described in Non-Patent Document 1, for example, but is not limited thereto.
[0038] Note that the input of the neural network is, for example, voxel data (TSDF volume) of a scene that may include a plurality of objects with a dimension of 40×40×40. In this case, the output of the neural network is a score with a dimension of 40×40×40, a mask value with a dimension of 40×40×40, and pose data with a dimension of 40×40×40×3 (vector dimension)×2 (number of vectors). As described above, for each of the plurality of voxels, a score, a mask value, and pose data are output. Here, the score, the mask value, and the pose data may be output, for example, by the multi-head method of the neural network. That is, three heads output a score, a mask value, and pose data, respectively. The input-side network of the multi-head can be learned together without distinguishing the score, the mask value, and the pose data. On the other hand, the three heads can be learned independently. Note that since the score is a continuous value indicating the degree of grasping quality, while the mask value is a discrete value indicating the presence or absence of a grasping pose, the score and the mask value have different properties. Therefore, as described above, it is effective to separately learn a plurality of heads that output a score and a mask value, respectively.
[0039] The control device 140 controls the robot hand 10 so as to grip an object arranged in the three-dimensional space (S142 to S146). The position acquisition unit 142 acquires position data in the three-dimensional space (step S142). Specifically, the position acquisition unit 142 detects each position (voxel) of the three-dimensional space detected by the detection device 2. Then, the position acquisition unit 142 acquires the TSDF volume, which is the position data, for each voxel of the three-dimensional space.
[0040] The estimation unit 144 estimates, for each position, the score, the mask value, and the posture data with respect to the reference point Pr of the robot hand 10 using the inference model (step S144). Specifically, the estimation unit 144 uses an inference model generated in advance by machine learning to estimate the score, the mask value, and the posture data with respect to the position of the reference point Pr of the robot hand 10 in the three-dimensional space for each position in the three-dimensional space. More specifically, the estimation unit 144 inputs the TSDF volume of each voxel acquired by the position acquisition unit 142 into the inference model. Thereby, the estimation unit 144 acquires the score, the mask value, and the posture data for each voxel output from the inference model.
[0041] The hand control unit 146 controls the robot hand 10 based on the estimated score and posture data for positions where the estimated mask value indicates that gripping is executable (step S146). For example, the hand control unit 146 may control to position the reference point Pr at each voxel where the mask value indicates "true: grippable" in the order according to the score. In this case, the hand control unit 146 may reproduce the posture of the robot hand 10 using the rotation matrix obtained according to the posture data of the voxel, by placing the reference point Pr at the voxel in order from the ones with higher scores among the voxels where the mask value indicates "true: grippable". Thereby, the hand control unit 146 causes the robot hand 10 to execute gripping. Note that it is not necessary to try gripping by positioning the reference point Pr in order from the position with the highest score. For example, even for a position with a high score, if the user knows in advance that it is difficult to grip the object (for example, a position where the robot hand 10 has difficulty approaching), it may be excluded from the gripping trial (or the priority of the trial may be lowered). On the other hand, for a position where the user knows in advance that it is easy to grip the object (for example, a position where the robot hand 10 can approach from above the object) and that has a high score, the reference point Pr may be positioned to try gripping.
[0042] Note that the rotation matrix R is represented by R = [e x e y e z . Here, e x , e y , e z are the basis vectors of the rotation matrix R corresponding to the x-direction, y-direction, and z-direction of the hand coordinate system, respectively, and are represented by, for example, the following equation 2.
Equation
[0043] Here, ~ e xcorresponds to the unit vector in the direction in which the robot hand 10 approaches the object in the posture data output from the inference model. Also, ~ e z corresponds to the normal vector of the plane in which the finger part 14 moves in the posture data output from the inference model. Note that once the posture data ( ~ e x and ~ e z ) is determined, the posture of the robot hand 10 is determined from the above formula 2. Therefore, the position and posture of the robot hand 10 are uniquely determined once the position of the reference point Pr and the posture data ( ~ e x and ~ e z ) at that position are determined. Also, ~ e x (the direction in which the robot hand 10 approaches the object) output from the inference model is determined by the relative positional relationship between the reference point Pr and the object, and thus can be estimated accurately. Also, ~ e z (the normal vector of the plane in which the finger part 14 moves) output from the inference model may depend on the curvature of the object, and thus can be estimated accurately. Therefore, by expressing the rotation matrix as in the above formula 2, the gripping performance can be improved.
[0044] Note that the output from the inference model does not show data indicating whether the robot hand 10 performs precision grasping or enveloping grasping when there is a reference point Pr in each voxel. However, based on the positional relationship between the reference point Pr and the object to be grasped, it can be uniquely determined whether the robot hand 10 performs precision grasping or enveloping grasping on the object. That is, when the distance between the reference point Pr and the object is considerably far to the extent that the robot hand 10 cannot reach the object, there is no grasping posture. In that case, for the voxel where the reference point Pr is located, the mask value output from the inference model indicates "false". Also, when the distance between the reference point Pr and the object is slightly close to the extent that the robot hand 10 can grasp the object near the tip of the finger part 14, the robot hand 10 can perform precision grasping. Further, when the distance between the reference point Pr and the object is considerably close to the extent that the robot hand 10 can grasp the object by contacting the hand body 12, the robot hand 10 can perform enveloping grasping. Note that when the robot hand 10 grasps the same object, the score for enveloping grasping can be higher than that for precision grasping. Therefore, as long as the mask value indicates "true" for a position (voxel) close to the object, it is possible that enveloping grasping is attempted by positioning the reference point Pr at that position. Therefore, by the inference model outputting the score and the mask value, it becomes possible to cause the robot hand 10 to perform not only precision grasping but also enveloping grasping.
[0045] As described above, the grasping system 1 according to the first embodiment is configured to estimate a score and a mask value using an inference model learned by machine learning and control the robot hand 10. Therefore, as described above, the grasping system 1 according to the first embodiment can position the robot hand 10 so that the robot hand 10 can perform enveloping grasping. Therefore, high-quality grasping using the robot hand 10 can be realized.
[0046] (Modification example) Note that the present invention is not limited to the above-described embodiments, and can be appropriately modified without departing from the gist thereof. For example, in the above-described embodiments, the position in the three-dimensional space is represented by the TSDF volume (voxel), but it is not limited thereto. The position may be represented by point cloud data. However, since the TSDF volume (voxel) corresponds to the pixel in the two-dimensional image, it has the advantage of being easy to handle by the neural network. Also, while the point cloud data is the position on the object, the TSDF volume (voxel) is not limited to the position on the object. Therefore, by representing the position by the TSDF volume (voxel), gripping can be executed even if the reference point Pr is away from the object.
[0047] When the above-described program is loaded into a computer, it includes a group of instructions (or software code) for causing the computer to perform one or more functions described in the embodiments. The program may be stored in a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, the computer-readable medium or the tangible storage medium includes random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technologies, CD-ROM, digital versatile disk (DVD), Blu-ray (registered trademark) disk or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices. The program may be transmitted on a transitory computer-readable medium or a communication medium. By way of example and not limitation, the transitory computer-readable medium or the communication medium includes electrical, optical, acoustic, or other forms of propagated signals.
Explanation of Reference Numerals
[0048] 1... Gripping system, 2... Detection device, 10... Robot hand, 12... Hand body, 14... Finger part, 100... Control system, 120... Training data generation device, 122... Simulation execution part, 124... Training data generation part, 130... Learning device, 132... Training data acquisition part, 134... Learning part, 140... Control device, 142... Position acquisition part, 144... Estimation part, 146... Hand control part
Claims
1. Using an inference model generated in advance by machine learning, for each position in the three-dimensional space, a score indicating the quality of grasping when the robot hand grasps an object with respect to the position of the reference point of the robot hand, and a mask value indicating whether grasping can be executed are estimated by an estimation unit, A hand control unit that controls the robot hand based on the score for positions where the mask value indicates that grasping can be executed, A control device having the above.
2. The score corresponds to the maximum value of the external force that can maintain the grasp when the robot hand executes the grasp against the external force applied to the object, The control device according to claim 1.
3. The estimation unit further estimates posture data indicating the grasping posture of the robot hand when grasping the object by the inference model, The hand control unit controls the posture of the robot hand based on the estimated posture data, The control device according to claim 1.
4. A training data acquisition unit that acquires training data with position data indicating each position in the three-dimensional space where the object is placed as input data, and a score indicating the quality of grasping when the robot hand grasps the object and a mask value indicating whether grasping can be executed as output data for each position where there is a reference point of the robot hand, A learning unit that generates an inference model that outputs the output data when the input data is input by executing machine learning, A learning device having the above.
5. A simulation execution unit that executes a physical simulation so that an object is grasped in the three-dimensional space where the object is placed, Training data used to generate an inference model by performing machine learning, the training data generation unit generating training data indicating a score indicating the quality of grasping when the robot hand grasps the object and a mask value indicating whether grasping can be performed when there is a reference point of the robot hand at each position in the three-dimensional space realized by the physical simulation. A training data generation device having
Citation Information
Patent Citations
Information processing device, workpiece recognition device, and workpiece removal device
JP6758540B1
Robot control device and robot control method
JP7337285B2
Learning device, learning method, learning model, detection device, and holding system
JP2019164836A