Efficient Data Generation for Grasp Learning by General Grippers
A method using a database, iterative optimization, and neural network training generates efficient and robust grasps for robots, addressing computational inefficiencies and real-world grasping challenges in cluttered environments.
Patent Information
- Application Number
- JP2021136160
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-10
- Filing Date
- 2021-08-24
- Publication Date
- 2025-06-25
- Estimated Expiration
- 2041-08-24
AI Technical Summary
Existing learning-based grasping detection methods for robots are computationally expensive and require extensive empirical trials, often failing to provide optimal grasping solutions in real-world scenarios, especially when dealing with stacked parts and collision avoidance.
A method involving a database of solid models, random initialization, iterative optimization, and physical environment simulation to generate high-quality grasps, using a neural network trained on simulated grasping data to identify optimal grasps from 3D camera images.
Generates diverse and robust grasps efficiently, reducing computational effort and enabling real-time grasping in cluttered environments, with simulation speeds 10-100 times faster than actual trials, and capable of handling various gripper and object combinations.
Smart Images

Figure 0007698514000005 
Figure 0007698514000006 
Figure 0007698514000007
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to a method for generating a good gripping posture for gripping parts with a robot, and in particular to a random initialization where a random object and a gripper are selected from a large-scale object database, an iterative optimization where hundreds of grippings are calculated for each part in surface contact with the gripper, and a physical environment simulation where the gripping of each part is applied to a pile of objects simulated in a bin. The present disclosure relates to a method for robot gripping learning including these aspects.
Background Art
[0002] It is well known to widely perform operations of manufacturing, assembling, and material movement using industrial robots. One such application is a pick-and-place operation where a robot takes out individual parts from a bin and places each part on a conveyor or in a transport container. In an example of this application, molded or machined parts are dropped into a bin and arranged in random positions and orientations, and the robot picks up each part and places it on a conveyor in a predefined orientation (pose), and the conveyor transports the parts for packaging or further processing. In another example, in a warehouse that processes e-commerce orders, it is necessary to reliably handle items of various dimensions and shapes. Depending on how much the parts in the bin are stacked, a finger-type glass gripper or a suction-type gripper can be used as a robot tool. A vision system (one or more cameras) is typically used to identify the position and orientation of individual parts in the bin.
[0003] Conventional grasping generation methods manually teach known 3D shapes or picking points of objects. In these methods, a significant amount of time is required for trial-and-error designs to identify the optimal grasping pose, and manually designed heuristics may not function for unknown objects or occlusions. Due to the difficulty of using heuristic grasping teaching, learning-based grasping detection methods that can adapt to unknown objects have become popular.
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, existing learning-based grasping detection methods also have drawbacks. One of the known learning-based approaches uses a mathematically rigorous grasp quality to search for grasping candidates, but this process is performed before feeding those candidates into a convolutional neural network (CNN) classifier. However, this method usually has a high computational cost, and due to the simplifications involved in optimization, the solution may not be optimal in the real world. For realistic grasping, other methods use empirical trials to collect data, but this method usually requires tens of thousands of robot hours with complex force control, and the entire process needs to be repeated when the gripper is changed.
[0005] In light of the above situation, a high-quality grasping candidate generation and computationally efficient robot grasping learning technology without manual teaching is desired. This technology provides a grasping scenario applicable to real-world situations including stacked parts and collision avoidance between the robot arm and the bin side.
Means for Solving the Problems
[0006] According to the teachings of the present disclosure, a grasping generation technique for picking up parts with a robot is presented. For all objects and grippers to be evaluated, a database of solid models or surface models is provided. A gripper is selected and a random initialization is performed. Here, a plurality of random objects are selected from the object database and a plurality of poses are randomly started. Next, an iterative optimization calculation is performed, where hundreds of grasps are calculated for each part having a surface in contact with the gripper and sampling is performed for grasp diversity and global optimization. Finally, a physical environment simulation is performed, where the grasp of each part is mapped to a pile of objects simulated in a bin scenario. Next, a neural network for grasp learning in actual robot operations is trained using the grasp points and approach directions from the physical environment simulation. The simulation results are associated with the depth image data of the camera to identify high-quality grasps.
[0007] Additional features of the disclosed method will become apparent from the following description and the claims, in conjunction with the accompanying drawings.
Brief Description of the Drawings
[0008]
Figure 1
[0009]
Figure 2
[0010]
Figure 3
[0011]
Figure 4
DETAILED DESCRIPTION OF THE INVENTION
[0012] The following description of embodiments of the present disclosure directed to optimization-based grasp generation techniques is illustrative in nature and is not intended to limit the disclosed techniques or their applications or uses.
[0013] It is well known to use industrial robots to pick parts from a source and place them at a destination. In one common application, the parts are supplied by bins containing a random pile of parts that have just been cast or molded, in which case the parts need to be transferred from the bin to a conveyor or transport container. It has always been difficult to teach a robot to recognize and grasp individual parts within a bin filled with parts in real time.
[0014] To improve the speed and reliability of the part picking operation by a robot, it is known to pre-compute specific gripper grasping actions for grasping a particular part in various poses. This pre-computation of the grasping action is known as grasp generation, and the pre-computed (generated) grasps are used to make decisions in real time during the part picking operation by the robot.
[0015] Conventional grasping generation methods manually teach picking points of the known 3D shape of an object. In these methods, a significant amount of time is required for a trial-and-error (heuristic) design to identify the optimal grasping pose, and the manually designed heuristics may not function for unknown objects or occlusions. Since it is difficult to use trial-and-error grasping teaching, learning-based grasping detection methods capable of adapting to unknown objects have become popular.
[0016] However, existing learning-based grasping generation methods also have drawbacks. In one well-known learning-based technique, mathematically rigorous grasping quality is used to search for grasping candidates, but this process is performed before feeding those candidates to a CNN classifier. However, this is computationally expensive, and due to the simplifications used for optimization, the solution may not be optimal in the real world. In other methods, empirical trials are used to collect data for generating realistic grasps, but this method usually requires tens of thousands of robot hours involving complex force control, and the entire process needs to be repeated when the gripper is changed. Furthermore, some existing grasping generation methods are limited to identifiable types of grasping poses, for example, being limited only to the direct top-down approach direction.
[0017] The technology described by this disclosure is automatically applicable to any combination of gripper and part / object designs, generating a large number of realistic grasps with minimal computational effort in simulation, and further simulating the complexity of grasping individual parts from a pile of multiple parts mixed in a bin, as is often seen in real-world robotic part-picking operations. To enhance the robustness of the grasp, a mathematically rigorous grasp quality is used, and the contact is modeled as a surface. A specially designed solver is used to efficiently solve the optimization. Finally, the generated grasps are tested and refined in a physical environment simulation step to account for interference between the gripper and the parts that occurs in a cluttered environment. The grasps generated and evaluated in this way are used in actual robotic part-picking operations to identify the target object and the grasp pose from 3D camera images.
[0018] FIG. 1 is an exemplary flowchart diagram 100 of a grasp generation process that calculates optimized grasps for individual parts / objects according to an embodiment of the present disclosure and applies the calculated grasps to the objects in a simulated pile. In block 110, a database 112 of three-dimensional (3D) object models is provided. The object model database 112 can include hundreds of different objects for which grasps are to be generated, and each object is provided with 3D solid data or surface data (typically from a CAD model). The object model database 112 can also include data from widely available shared source repositories such as ShapeNet®.
[0019] In box 110, a gripper database 114 is also provided, which includes both the 3D geometric data and joint data of each gripper. For example, one particular gripper has three mechanical fingers, and each finger has two knuckle joints and two finger segments. The 3D geometry of this gripper is provided in a specific configuration, such as when all the knuckle joints are fully open, and the pivot geometry of the joints is also provided. The gripper database 114 can include various types of grippers, such as two-finger and three-finger articulated grippers, parallel-jaw grippers, full-human-hand-style grippers, underconstraint-actuated grippers, suction-cup-style grippers (single or multiple cups), and so on.
[0020] In the random initialization box 120, a group of objects (e.g., 10 - 30 objects) is randomly selected from the object model database 112, and a gripper for each object is also selected from the gripper database 114. For example, a rabbit 122 (which could be a molded plastic toy) and a corresponding three-finger gripper 124 are illustrated, where the rabbit 122 is just one of the objects in the object database 112, and the gripper 124 is one of the grippers included in the gripper database 114.
[0021] As another example, a teapot-shaped object 126 is illustrated with the same gripper 124. For clarity, each object selected in the random initialization box 120 is analyzed individually with the selected gripper to generate a number of robust grasps for the object, which will be described in detail below. Merely to facilitate the operation, a number (e.g., 10 - 30) of objects may be selected (randomly or by the user), in which case, in box 120, the grasp generation calculation is automatically performed for all the selected objects. In a preferred embodiment, the same gripper (gripper 124 in this example) is used for all the objects selected in box 120. The reason for this is that in later analysis, it may involve stacking a large number or all different objects in a bin and using a robot equipped with a gripper to pick up the objects one by one.
[0022] In box 130, an iterative optimization calculation is performed for each object / gripper pair, and a number of robust grasps are generated and stored. In one embodiment, the iterative optimization routine is configured to calculate 1000 grasps for each object and the selected gripper. Of course, more or fewer than 1000 grasps can also be calculated. The iterative optimization calculation models the surface contact between the object and the gripper while preventing collisions and penetrations. This calculation uses a solver specially designed for efficiency to quickly calculate each grasp. The initial conditions (the pose of the gripper with respect to the object) are changed to provide a variety of combinations of robust grasps. The iterative optimization calculation in box 130 will be described in detail below.
[0023] In box 140, the group of objects selected by the random initialization box 120 is used for simulating the physical environment (i.e., simulating the parts that randomly fall into the bin and pile up), and the related grasping from box 130 is mapped to the physical environment. The simulated pile of objects in the bin in box 140 may include many different types of objects (e.g., all the objects selected by box 120), or may include bins filled with all objects of a single type. The simulations are run in parallel to test the performance of each grasp in a cluttered environment. The physical environment simulation in box 140 will be further described below.
[0024] In box 150, for forming the grasping database, the point cloud, the grasping pose, and the success rate from the physical environment simulation in box 140 are recorded. The point cloud depth image 152 shows the pile of objects from box 140 as seen from a specific viewpoint. In a preferred embodiment, the depth image 152 is as seen from the approach direction calculated for the best grasp. From the image 152, calculations in box 140 determined several grasping candidates that could be executed by the robot gripper. Each of the grasping candidates is represented by a grasping pose and a point map 154. The point map 154 shows the points that can be used as grasping targets, together with the approach angle defined by the viewpoint of the image 152, and uses the gripper width and gripper angle calculated in box 140. Thus, the data stored in box 150 includes the depth map from the desired approach angle, the point map 154 showing the grasping coordinates x / y / z including the best grasp, the gripper rotation angle and gripper width, and the grasping success rate from the physical environment simulation. The points in the point map 154 are ranked from the perspective of grasping quality and will result in successfully grasping an object from the pile of objects in the bin.
[0025] The grasps generated and evaluated by the method shown in the flowchart of FIG. 1 (FIG. 100) and stored in the box 150 are later used to train a grasp learning system for identifying the best grasp pose from the 3D camera images of the actual bins filled with parts.
[0026] FIG. 2 is an exemplary flowchart diagram 200 of the steps included in the iterative optimization box 130 of the grasp generation process of FIG. 1 according to an embodiment of the present disclosure. The iterative optimization calculation of the flowchart diagram 200 is performed for each object / gripper pair to generate and store many robust grasps.
[0027] In box 210, the surface of the gripper and the surface of the object are discretized into points. The points on the gripper surface (the palm surface and the inner surfaces of the finger segments) are designated as p i and each of the points p i has a normal vector n i p . The points on the outer surface of the object are designated as q i and each of the points q i has a normal vector n i q .
[0028] In box 220, based on the current pose of the gripper with respect to the object (the overall gripper position and the individual finger joint positions will be described later), point contact pairs and collision pairs are calculated. This calculation starts by identifying, using the nearest neighbor method, the points on the gripper surface that match the closest points on the object surface. After filtering out pairs of points with a distance exceeding a threshold, the remaining pairs of points (p i , q i ) define the contact surface S f on the (gripper) and the contact surface S o on the (object).
[0029] In 222, the cross-section of the object and one finger of the gripper are for the corresponding pairs of points (p i , q i) and are shown together with their respective surface normal vectors. At the position shown in 222 (the initial position or other iterations in the optimization process), there is interference between the object and the outer segments of the gripper fingers. This interference results in a penalty being imposed using a constraint function in the optimization calculation to move the gripper away from the object to eliminate the interference, as will be described later.
[0030] In box 230, the grasping search problem is modeled as an optimization and one iteration calculation is performed. Surface contact and exact mathematical properties are used in the modeling to calculate a stable grasp. A penalty is also imposed in the optimization to avoid penetration in the case of a collision between the gripper and the object, as described above. The optimization formulation shown in box 230 is reproduced below as equations (1a)-(1f) and will be explained in the following paragraphs.
Equation
[0031] The optimization formulation includes an objective function (Equation 1a) defined to maximize the grasping quality Q, where the grasping quality Q is related to the contact surfaces S f , S o and the object geometry O. The grasping quality Q can be defined in any suitable way and is calculated from the force contribution of all contact points to properties of the object such as mass and center of gravity. A stable grasp means that small movements of the object within the gripper are quickly stopped by frictional forces and the grip is not lost.
[0032] Equations (1a)-(1f) from the optimization formulation include many s functions. The constraint function (Equation 1b) indicates that S o is a subset of the object surface δ0 transformed by the initial pose T ini,o of the object, and that S f is a subset of the hand / gripper surface δF transformed by the pose and joint positions (T,q) of the gripper. The constraint function (Equation 1c) indicates the contact surface S of the objecto and the contact surface S of the hand / gripper f indicates that they are the same.
[0033] The constraint function (Equation 1d) indicates that during the interaction between the gripper and the object, the contact force f remains within the friction cone FC. The friction cone FC is characterized by imposing a penalty on the deviation of the finger force from the center line of the friction cone. Here, the friction cone, as is well known in the art, is a cone such that the force exerted by one surface on the other surface must be placed within the friction cone when both surfaces are stationary as determined by the coefficient of static friction. The constraint function (Equation 1e) indicates that the transformed hand surface (T(δF; T, q)) must not penetrate the environment ε, i.e., the distance must be zero or greater.
[0034] The constraint function (Equation 1f) indicates that the joint position q remains within the space constrained by qmin and qmax. Here, the joint position boundaries qmin and qmax are known and applied to the selected gripper. For example, the finger joints of the illustrated three-finger gripper are constrained to have an angular position between qmin = 0° (straight finger extension) and approximately qmax ≈ 140° - 180° (maximum inward flexion; each joint of the gripper has a specific qmax value within this range).
[0035] The optimization formulation of Equations (1a)-(1f) remains a non-convex problem due to the contact surface and non-linear kinematics in S f ⊂ (δF; T, q). To solve the kinematic non-linearity, the search is changed from the hand / gripper configuration T, q to the increments δT, δq of the hand / gripper configuration. Specifically, T = δT + T0, q = δq + q0, where T0, q0 represent the current hand / gripper configuration. In the present disclosure, δT = (R, t) is referred to as the transformation.
[0036] To solve the non-linearity introduced by surface contact and solve Equations (1a)-(1f) by a gradient-based method, the hand / gripper surface δF and the object surface δ0 are each point sets JPEG0007698514000002.jpg9119 and discretize to JPEG0007698514000003.jpg10116. Here, JPEG0007698514000004.jpg8117 represents the position of the point and the normal vector. The discretization of the point is described above with respect to boxes 210 and 220.
[0037] S f and S o Using the nearest neighbor point matching approach described above to define, for the translational Jacobian matrix at point p i By using it and describing the point distance in the surface normal direction of the object, the contact proximity penalty can be formulated. With this point - plane distance, the points on the gripper can exist on the surface of the object and slide on the surface. Also, the sensitivity of the algorithm to incomplete point cloud data is reduced.
[0038] The collision constraint (Equation 1e) is formulated to receive a penalty and impose a penalty on collision only for the currently penetrating points. Since the collision approximation introduces a differential form with respect to δT, δq, the computational efficiency is greatly improved. However, due to the lack of preview, when the hand moves, the approximate penalty becomes discontinuous, so the optimized δT, δq may exhibit a zigzag behavior. To reduce the influence of the zigzag caused by the approximation, the hand surface can be inflated to preview possible collisions.
[0039] Returning here to Figure 2, the above discussion explains the optimization formulation in box 230, which includes a convexification simplification that can make the optimization calculation to be executed very fast. In box 240, the pose of the gripper with respect to the object (both the pose T of the base / palm of the gripper and the joint position q) is updated, and the steps of box 220 (calculating the contact point pairs and collision point pairs) and the steps of box 230 (solving the iteration of the optimization problem) are updated and iterated until the optimization formulation converges to a predefined threshold.
[0040] The steps shown in the flowchart 200 shown in FIG. 2 calculate a single grasp of an object (such as a rabbit 122) by a gripper (such as 124). That is, when the optimization calculation converges, a single grasp with appropriate quality and no interference / penetration is provided. In the cross-section shown at 242, by relaxing the angle of the gripper finger joint, the interference between the object and the gripper finger is eliminated, and it can be seen that the point pairs (p i ,q i ) are just in contact.
[0041] As described above, it is desirable to calculate many different grasps for each object / gripper pair. In one embodiment, the iterative optimization routine is configured to calculate 1000 grasps for each object using the selected gripper.
[0042] To obtain different converged grasps, it is also desirable to sample from different initial grasp poses, and the resulting grasps are comprehensive. In other words, the initial conditions (the pose of the gripper relative to the object) vary randomly for each of the 1000 calculated grasps. The reason for this is that the optimization formulation of equations (1a)-(1f) converges to a local optimum. To grasp various parts of the head and body (such as the rabbit 122) from various directions (front, back, up, down, etc.), the initial conditions need to reflect various approach directions and target grasp positions.
[0043] Even if the initial conditions change to provide a diversity of grasping poses, considering a large number (e.g., 500 - 1000) of grasps calculated in the iterative optimization steps, a large number of very similar grasps will necessarily exist. For example, it is easily foreseeable that many similar grasps will be calculated for the head of the rabbit 122 from the front. Therefore, after 500 - 1000 grasps are calculated for an object, those grasps are grouped by similar poses and an average is calculated. In one embodiment of the grasp generation method, only the average grasp is stored. In other words, for example, all of the grasps of the head of the rabbit 122 from the front are averaged into a single stored grasp. The same is true for other approach directions and grasp positions. In this way, the 500 - 1000 calculated grasps can be reduced to a number in the range of 20 - 50 stored grasps, and each of the stored grasps is significantly different from the other stored grasps.
[0044] Figure 3 is an illustrated flowchart diagram 300 of the steps included in the physical environment simulation box 140 of the grasp generation process of FIG. 1, according to one embodiment of the present disclosure. In box 310, a set of optimized grasps for an individual object (for a particular gripper) is provided. The optimized grasps are calculated and provided by the iterative optimization process of box 130 (FIG. 1) and the flowchart diagram 200 of FIG. 2. As described above, the optimized grasps are preferably stored (20 - 50) grasps representing the diversity of the grasp positions / angles, and each of the stored grasps is significantly different from each other.
[0045] In box 320, a simulated pile of objects is provided. The objects may all be of the same type (illustrated as bin 322), or may include many different types (such as those provided in the random initialization step of box 120 in FIG. 1 and illustrated as bin 324). The grasping provided in box 310 must include all types (if multiple) of objects within the simulated pile using the same gripper. In box 320, the pile of objects is simulated using an actual dynamic simulation of the objects, and the objects of the simulation fall into the bin in a random orientation, land on the bottom of the bin, collide with other objects already landed, and roll down the side of the formed pile until reaching an equilibrium position, etc. The dynamic simulation of the pile of objects includes the actual part shape, bin dimensions / shape, and surface - to - surface contact simulation for high realism.
[0046] After the simulated pile of objects in the bin is provided in box 320, the grasping (recorded optimized grasping) provided in box 310 is mapped to the simulated pile of objects. This step is shown in box 330. Since the simulated pile of objects includes the known positions of the individual objects, the optimized grasping can be mapped to the simulated pile of objects by using the pose and ID (identity) of the objects. Thereby, a simulated grasping of the objects is obtained, including a 3D depth map of the pile of objects, the ID of the selected object, the corresponding approach angle, three - dimensional grasping points, gripper angle, and gripper width (or finger joint positions).
[0047] The exposed surface of the pile of simulated objects is modeled as a 3D point cloud or depth map. Since the pile of simulated objects includes the known positions of individual objects, the 3D depth map can be calculated from any suitable viewing angle (such as within 30° from the vertical, an angle suitable for robot grasping). Next, by analyzing the 3D depth map from each simulated grasping viewing angle, a correlation can be found between the exposed parts of the objects in the 3D depth map and the simulated grasping corresponding to one of the memorized optimized graspings.
[0048] The provision of the pile of simulated objects (box 320) can be repeated many times for a given set of optimized graspings. Each pile of simulated objects uses different random streams of objects and postures during dropping. Therefore, all the simulated piles are different and provide different viewpoints for grasping decisions. More specifically, for any random pile of objects, the grasping approach direction can be randomly selected, and the grasping close to that approach direction can be tested in the simulation. The individual object grasping simulations (box 330) can be repeated for each pile of simulated objects until all the objects are grasped (in the simulation) and removed from the bin.
[0049] By repeating the steps of boxes 320 and 330, each grasping can be simulated under different conditions, which include objects intertwined with each other, objects that are partially exposed but stuck in a certain position by other objects in the pile, and the sides / corners of the bin. Furthermore, the grasping simulation may incorporate variations and uncertainties, which include uncertainties in the pose of the object, sensing uncertainties, friction uncertainties, and various environments (bin, object). By performing grasping trials under these various situations, the robustness of each grasping under uncertainties, variations, and interferences can be simulated and recorded.
[0050] Returning to FIG. 1, the success rates from the physical environment simulation of box 150, the point cloud, the grasping pose, and box 140 (and as described above with respect to FIG. 3) are recorded to form a grasping database. The point cloud depth image 152 shows a pile of objects of box 140 as viewed from a particular viewpoint. From the image 152, calculations at box 140 determined several grasping candidates that could be adopted by the robot's gripper. Each of the grasping candidates is represented by a region within the point map 154, the points included in the point map 154 being available as grasping targets, and the approach angle being defined by the viewpoint of the image 152 using the gripper angle and gripper width calculated at box 140. The points within the point map 154 are ranked with respect to grasping quality, resulting in the object being successfully grasped from the pile of objects within the bin.
[0051] The grasps generated and evaluated by the method shown in the flowchart of FIG. 1 in FIG. 100 and stored in box 150 are later used as training samples for the neural network system. This neural network system is used in an actual robot part picking operation to identify a target grasping pose from a 3D camera image of an actual bin filled with parts.
[0052] FIG. 4 shows a block diagram of a robot part picking system that uses a neural network system for grasping calculations, and this neural network system is trained using the grasps generated by the process of FIGS. 1-3 according to an embodiment of the present disclosure. A robot 400 having a gripper 402 operates within a work space, where the robot 400 moves parts or objects from a first location (bin) to a second location (conveyor). The gripper 402 is the gripper identified in box 120 of FIG. 1.
[0053] The movement of the robot 400 is controlled by a controller 410, which typically communicates with the robot 400 via a cable 412. As is well known in the art, the controller 410 sends joint movement instructions to the robot 400 and receives joint position data from the encoders of the joints of the robot 400. The controller 410 also sends instructions (rotation angle and width, as well as instructions regarding grip / ungrip) for controlling the operation of the gripper 402.
[0054] A computer 420 communicates with the controller 410. The computer 420 includes a processor and memory / storage configured to calculate high-quality grasps for bins filled with objects in real time in one of two ways. In a preferred embodiment, the computer 420 runs a neural network system pre-trained for grasp learning using a grasp database from the box 150. The neural network system then calculates grasps in real time based on live image data. In other embodiments, the computer 420 calculates grasps during live robot operation directly from the grasp database of the box 150, including point clouds, grasp poses, and success rates from physical environment simulations. The computer 420 may be the same computer that performs all of the grasp generation calculations described above with respect to FIGS. 1-3.
[0055] A pair of 3D cameras 430 and 432 communicate with the computer 420 and provide images of the work space. In particular, the cameras 430 / 432 provide images of the object 440 within the bin 450. Images (including depth data) from the cameras 430 / 432 provide point cloud data that defines the position and orientation of the object 440 within the bin 450. Since there are two 3D cameras 430 and 432 with different viewpoints, a 3D depth map of the object 440 within the bin 450 can be calculated or projected from any suitable viewpoint.
[0056] The task of the robot 400 is to pick out one of the objects 440 from the bin 450 and transfer the object to the conveyor 460. In the illustrated example, individual parts 442 gripped by the gripper 402 of the robot 400 are selected and moved to the conveyor 460 along the path 480. For each part picking operation, the computer 420 receives an image of the object 440 in the bin 450 from the cameras 430 / 432. The computer 420 calculates a depth map of the pile of objects 440 in the bin 450 from the camera images. Since the camera images are provided from two different viewpoints, the depth map of the pile of objects 440 can be calculated from different viewpoints.
[0057] In a preferred embodiment, the computer 420 includes a neural network system trained for grasping learning. The neural network system is trained using supervised learning using data from the grasping database (point clouds, grasping poses, and success rates from physical environment simulations) of the box 150. The methods of FIGS. 1-3 of the present disclosure provide the amount and variety of data required for complete training of the neural network system. This includes various grasps for each object and various random piles of objects with the degrees of freedom (DOF) of the target grasp and grasping success rates. All data in the grasping database from the box 150 can be used to automatically train the neural network system. Next, the neural network system of the computer 420 is executed in inference mode during live robot operation, and the neural network system calculates a high-quality grasp based on the object pile image data from the cameras 430 / 432.
[0058] In other embodiments, computer 420 directly identifies the grasp during live robot operation based on the grasp database from box 150. In this embodiment, computer 420 knows in advance the type of object 440 contained in bin 450. This information is included in the grasp database from box 150 (along with the point cloud, grasp pose, and success rate from the physical environment simulation). When a depth map containing an object (such as object 442) at a high-quality grasp position according to a pre-generated grasp database is found, computer 420 sends the individual object grasp data to controller 410, which then commands robot 400 to grasp and move the object.
[0059] In any of the above embodiments, the grasp data provided by computer 420 to controller 410 includes the 3D coordinates of the grasp target point, the approach angle that gripper 402 should follow, and the rotation angle and width of the gripper (or the position of all finger joints). Controller 410 can use the grasp data to calculate a robot motion command, which causes gripper 402 to grasp an object (such as object 442) and move the object along a collision-free path (path 480) to a target location.
[0060] The target location may be a transport container where objects are placed in individual compartments, instead of conveyor 460, or another surface or device where the objects are further processed in a later operation.
[0061] After the object 442 is moved to the conveyor 460, since the pile of objects 440 changes, new image data is provided by the cameras 430 / 432. Next, the computer 420 must identify a new target grasp based on the new image data and a pre-generated grasp database. The new target grasp must be identified very quickly by the computer 420 because the identification of the grasp and the calculation of the path must be performed in real time at the same speed as the robot 400 can move one of the objects 440 and pick up the next object. The generation of a database of quality grasps including corresponding depth image data, grasp points, approach angles, and gripper configurations for each grasp enables pre-training of the neural network system and allows real-time calculations to be performed very quickly and efficiently during actual robot operation. The disclosed method easily and automatically generates a database of grasps for many objects and corresponding grippers.
[0062] The above-described grasp generation technique has several advantages over existing methods. The disclosed method generates high-quality full-DOF grasps. This method generates reasonable grasps with surface contact, so the generated grasps are more robust to uncertainty and disturbances. Furthermore, the disclosed optimization formulation and customized iterative solver are very efficient and calculate grasps in a time range from 0.06 seconds (gripper with one joint) to 0.5 seconds (multi-fingered hand with 22 joints). Regarding physical simulation, it is 10 - 100 times faster than actual trials, can test grasp trials within 0.01 - 0.05 seconds, and generates 1 million grasps in 10 hours.
[0063] Furthermore, the disclosed method generates diverse grasping data that includes various variations and interferences. The generation pipeline simulates the grasping performance under variations (variations in the shape of the object, uncertainty in pose) and interferences (entanglement, jamming, corners). Therefore, subsequent learning algorithms can learn robust grasping strategies based on this grasping data. The optimization framework is effective for suction grippers, conventional finger-type grippers, customized grippers, multi-fingered hands, and micro-deformable soft grippers. It is also effective for both under-actuated hands and fully actuated hands. Finally, the disclosed method is mathematically stable, easily solvable, and optimizes the exact grasping quality to generate proper grasps. Despite the exact quality and all the constraints, the algorithm can be solved iteratively with basic linear algebra.
[0064] Throughout the foregoing description, various computers and controllers have been described and implied. It should be understood that the software applications and modules of these computers and controllers are executed on one or more computing devices having a processor and a memory module. In particular, this includes a processor within a robot controller 410 that controls a robot 400 that performs object grasping, and a processor within a computer 420 that performs grasping generation calculations and identifies an object to be grasped in real-time operation.
[0065] Although multiple exemplary aspects and embodiments of optimization-based grasping generation techniques have been described, those skilled in the art will recognize their modifications, permutations, additions, and sub-combinations. Therefore, the appended patent claims and claims should be construed to include all such modifications, permutations, additions, and sub-combinations that are within their true spirit and scope. [Configuration 1] A method for generating a gripping database used in a robot, comprising: preparing an object database including 3D shape data of a plurality of objects, and a gripper database including 3D shape data and operating parameters of one or more grippers; using a computer having a processor and a memory to perform an initialization including selection of a gripper from the gripper database and selection of one or more objects from the object database; repeatedly performing iterative optimization to calculate, for each selected object, a plurality of quality grippings by the selected gripper that each achieve a predetermined quality criterion; performing a physical environment simulation to generate the gripping database, wherein the physical environment simulation includes repeatedly simulating a random pile of objects, repeatedly identifying a gripping pose for applying to the pile of objects using one of the quality grippings, and outputting data for each successfully simulated gripping to the gripping database; A method as described above. [Configuration 2] The method according to Configuration 1, further comprising: using the gripping database by the computer during live robot operation to identify a target object from a pile of objects, analyzing depth images from a plurality of 3D cameras, calculating one or more synthetic depth maps from the depth images, identifying, from the gripping database, an object having a quality gripping corresponding to one of the synthetic depth maps as the target object, and providing individual object gripping data to a robot controller for instructing the robot having the gripper to grip and move the target object. [Configuration 3] The method according to Configuration 1, further comprising: using the gripping database by the computer to train a neural network gripping learning system in supervised learning, and providing individual object gripping data to a robot controller for instructing the robot having the gripper to grip and move an object, wherein the trained neural network gripping learning system is used in inference mode to calculate gripping commands during live robot operation. [Configuration 4] The gripping command is the method according to Configuration 3, including the 3D coordinates of the gripping target point, the approach angle that the gripper should follow, and the rotation angle and width of the gripper or the positions of all finger joints. [Configuration 5] The 3D shape data of the plurality of objects is the method according to Configuration 1, including the solid model, surface model, or surface point data of each object. [Configuration 6] The operating parameters of the gripper are the method according to Configuration 1, including, for a finger-type gripper, the position of the finger joint, the pivot direction, and the flexion angle range of the finger joint. [Configuration 7] Repeatedly executing iterative optimization includes randomly changing the initial pose of the gripper with respect to the object, according to the method of Configuration 1. [Configuration 8] Repeatedly executing iterative optimization includes discretizing the surfaces of the gripper and the object into points, calculating contact point pairs and collision pairs based on the current pose of the gripper with respect to the object, repeatedly calculating the optimization model, updating the pose of the gripper, and repeatedly calculating the optimization model until convergence to achieve the quality criteria, according to the method of Configuration 1. [Configuration 9] The optimization model includes an objective function defined to maximize the quality criteria and a constraint function that defines the contact surfaces of the gripper and the object as subsets of the surface transformed and discretized by the pose of the gripper with respect to the object. The contact surface of the gripper and the contact surface of the object are equal to each other, and the contact surface of the gripper and the contact surface of the object have friction defined by the friction cone model, a penalty for the gripper penetrating into the object, and a gripper joint angle within a limited angle, according to the method of Configuration 8. [Configuration 10] The optimization model is convexified and simplified before obtaining a solution, and the convexification and simplification include changing the constraint function from the gripper configuration to an increment of the gripper configuration, defining the distance from each gripper surface point along the object surface perpendicular to the point of the nearest object, and considering only pairs of points within a threshold distance from each other, according to the method of Configuration 9. [Configuration 11] For each selected object, the plurality of quality grasps include at least 500 grasps, the at least 500 grasps are grouped by similar grasp poses, the average grasp of each group is calculated, and only the average grasp is used in the physical environment simulation, the method according to Configuration 1. [Configuration 12] The data of each successfully simulated grasp output to the grasp database includes the 3D depth map of the pile of the object, the gripper approach angle of the robot, the three-dimensional grasp point, the gripper angle, and the gripper width or knuckle position, the method according to Configuration 1. [Configuration 13] Executing the physical environment simulation includes mathematically simulating the dropping of a plurality of objects into a bin, the simulation including gravity, contact between the objects, and the walls of the bin, and the resulting pile of objects in the bin is used in the physical environment simulation, the method according to Configuration 1. A method including. [Configuration 14] Executing the physical environment simulation includes calculating one or more synthetic depth maps of the simulated pile of objects, identifying quality grasps in the synthetic depth maps, and simulating grasps using the gripper approach angle of the robot, the three-dimensional grasp point, the gripper angle, and the gripper width or knuckle position, the method according to Configuration 1. [Configuration 15] A method for generating a grasp database used in a robot, preparing an object database including 3D shape data of a plurality of objects, and a gripper database including 3D shape data and operating parameters of one or more grippers; using a computer having a processor and a memory to perform an initialization including selecting a gripper from the gripper database and selecting a plurality of objects from the object database; performing iterative optimization to calculate, for each selected object, a plurality of quality grasps by the selected gripper, grouping the quality grasps by similar grasp poses to calculate the average grasp of each group, and saving only the average grasp; Repeatedly simulating a random pile of objects, repeatedly identifying a gripping pose for applying to the pile of objects using one of the average grippings, and outputting data for each successfully simulated gripping to the gripping database, performing a physical environment simulation including the above to generate the gripping database, Training a neural network gripping learning system used in inference mode to calculate gripping instructions during live robot operation using the gripping database, and providing the gripping instructions to a robot controller that instructs the robot having the gripper to grip and move an object. A method including the above. [Configuration 16] In a system for a robot that grips an object, the system includes: A computer having a processor and a memory, Generating a gripping database, Providing an object database including 3D shape data of a plurality of objects, and a gripper database including 3D shape data and operating parameters of one or more grippers, Performing an initialization including selecting a gripper from the gripper database and selecting a plurality of objects from the object database, Performing iterative optimization to calculate a plurality of quality grippings by the selected gripper for each selected object, Repeatedly simulating a random pile of objects, repeatedly identifying a gripping for applying to the pile of objects using one of the quality grippings, and outputting data for each successfully simulated gripping to the gripping database, performing a physical environment simulation including the above to generate the gripping database, A computer configured to train a neural network gripping learning system using the gripping database in supervised learning, During live robot operation, providing a plurality of 3D cameras that provide a depth image of a pile of objects to the computer that calculates gripping instructions using the neural network gripping learning system in inference mode after training, A robot controller that communicates with the computer and receives the gripping instructions from the computer, A robot including the gripper that grips and moves an object based on an instruction from the computer. A system having... [Configuration 17] The gripping instruction includes the 3D coordinates of the gripping target point, the approach angle that the gripper should follow, and the rotation angle and width of the gripper or the positions of all finger joints. The system according to Configuration 16. [Configuration 18] Repeatedly executing iterative optimization includes discretizing the surfaces of the gripper and the object into points, calculating contact point pairs and collision pairs based on the current pose of the gripper with respect to the object, repeatedly calculating the optimization model, updating the pose of the gripper, and repeatedly calculating the optimization model until it converges to achieve the quality criteria. The system according to Configuration 16. [Configuration 19] The optimization model includes an objective function defined to maximize the quality criteria and a constraint function that defines the contact surface of the gripper and the contact surface of the object as a subset of the discretized surface transformed by the pose of the gripper with respect to the object. The contact surface of the gripper and the contact surface of the object are equal to each other. The contact surface of the gripper and the contact surface of the object have friction defined by the friction cone model, a penalty for the gripper penetrating into the object, and a gripper joint angle within a limited angle. The system according to Configuration 18. [Configuration 20] Executing the physical environment simulation includes mathematically simulating the dropping of multiple objects into a bin. The simulation includes gravity, contact between the objects, and the walls of the bin. The resulting pile of objects in the bin is used in the physical environment simulation. The system according to Configuration 16.
Claims
1. A method for generating a gripping database for use in a robot, comprising: providing an object database including 3D shape data of a plurality of objects, and a gripper database including 3D shape data and operating parameters of one or more grippers; performing an initialization including selection of a gripper from the gripper database and selection of one or more objects from the object database using a computer having a processor and a memory; repeatedly performing iterative optimization to calculate, for each selected object, a plurality of gripping poses by the selected gripper each achieving a predetermined gripping quality; performing a physical environment simulation to generate the gripping database, wherein the physical environment simulation includes repeatedly simulating a random pile of objects, repeatedly identifying a gripping pose for application to the pile of objects using one of the gripping poses, and outputting data for each simulated gripping to the gripping database; a method.
2. The method of claim 1, further comprising using the gripping database by the computer during live robot operation to identify a target object from a pile of objects, analyzing depth images from a plurality of 3D cameras, calculating one or more composite depth maps from the depth images, identifying, from the gripping database, an object having a gripping pose corresponding to one of the composite depth maps as the target object, and providing individual object gripping data to a robot controller for instructing the robot having the gripper to grip and move the target object.
3. The method of claim 1, further comprising training a neural network gripping learning system with supervised learning using the gripping database by the computer, and providing individual object gripping data to a robot controller for instructing the robot having the gripper to grip and move an object, wherein the trained neural network gripping learning system is used in an inference mode to calculate gripping commands during live robot operation.
4. The gripping command is the method according to claim 3, including the 3D coordinates of the gripping target point, the approach angle that the gripper should follow, and the rotation angle and width of the gripper or the positions of all finger joints.
5. The 3D shape data of the plurality of objects includes the solid model, surface model or surface point data of each object, and is the method according to claim 1.
6. The operating parameters of the gripper include, for a finger-type gripper, the position of the finger joint, the pivot direction, and the bending angle range of the finger joint, and are the method according to claim 1.
7. Repeatedly performing iterative optimization includes randomly changing the initial pose of the gripper with respect to the object, and is the method according to claim 1.
8. Repeatedly performing iterative optimization includes discretizing the surface of the gripper and the surface of the object into points, calculating contact point pairs and collision pairs based on the current pose of the gripper with respect to the object, repeatedly calculating the optimization model, updating the pose of the gripper, and repeatedly calculating the optimization model until it converges to achieve the gripping quality, and is the method according to claim 1.
9. The optimization model includes an objective function defined to maximize the gripping quality, and a constraint function that defines the contact surface of the gripper and the contact surface of the object as subsets of the surface of the gripper and the surface of the object, respectively, which are discretized into point sets by the pose of the gripper with respect to the object. The contact surface of the gripper and the contact surface of the object are equal to each other, and the contact surface of the gripper and the contact surface of the object have friction defined by the friction cone model, a penalty for the gripper penetrating into the object, and a gripper joint angle within a limited angle, and is the method according to claim 8.
10. The optimization model is convexified and simplified before obtaining a solution. The convexification and the simplification include changing the constraint function from the gripper configuration to the increment of the gripper configuration, defining the distance from each gripper surface point along the object surface perpendicular to the point of the nearest object, and considering only pairs of points within a threshold distance from each other, and is the method according to claim 9.
11. For each selected object, the plurality of grasping poses includes at least 500 grasps, the at least 500 grasps are grouped by similar grasping poses, the average of the grasping poses of each group is calculated, and only the average is used in the physical environment simulation. The method according to claim 1.
12. The data of each simulated grasp output to the grasping database includes the 3D depth map of the pile of the object, the gripper approach angle of the robot, the three-dimensional grasping point, the gripper angle, and the gripper width or knuckle position. The method according to claim 1.
13. Performing the physical environment simulation includes mathematically simulating the dropping of a plurality of objects into a bin, the simulation including gravity, contact between the objects, and the walls of the bin, and the resulting pile of objects in the bin being used in the physical environment simulation. The method according to claim 1. A method including.
14. Performing the physical environment simulation includes calculating one or more synthetic depth maps of the simulated pile of objects, identifying grasping poses in the synthetic depth maps, and simulating grasps using the gripper approach angle of the robot, the three-dimensional grasping point, the gripper angle, and the gripper width or knuckle position. The method according to claim 1.
15. A method for generating a grasping database used in a robot, Preparing an object database including 3D shape data of a plurality of objects, and a gripper database including 3D shape data and operating parameters of one or more grippers; Using a computer having a processor and a memory to perform an initialization including selecting a gripper from the gripper database and selecting a plurality of objects from the object database; Performing iterative optimization to calculate, for each selected object, a plurality of grasping poses by the selected gripper, grouping the grasping poses by similar grasping poses to calculate the average of the grasping poses of each group, and saving only the average; Repeatedly simulating a random pile of objects, repeatedly identifying a grasping pose for applying to the pile of objects using one of the means, and outputting data for each of the simulated grasps to the grasping database, performing a physical environment simulation including the above to generate the grasping database, Training a neural network grasping learning system that is used in an inference mode to calculate a grasping instruction during live robot operation after training, using the grasping database for supervised learning, and providing the grasping instruction to a robot controller that instructs the robot having the gripper to grasp and move an object. A method comprising the above.
16. In a system for a robot that grasps an object, the system comprises: A computer having a processor and a memory, Generating a grasping database, Providing an object database including 3D shape data of a plurality of objects, and a gripper database including 3D shape data and operating parameters of one or more grippers, Performing an initialization including selection of a gripper from the gripper database and selection of a plurality of objects from the object database, Performing iterative optimization to calculate a plurality of grasping poses for each selected object by the selected gripper, Repeatedly simulating a random pile of objects, repeatedly identifying a grasp for applying to the pile of objects using one of the grasping poses, and performing a physical environment simulation including outputting data for each of the simulated grasps to the grasping database to generate the grasping database, A computer configured to train a neural network grasping learning system using the grasping database for supervised learning, During live robot operation, providing a plurality of 3D cameras that provide a depth image of a pile of objects to the computer that calculates a grasping instruction using the neural network grasping learning system after training in an inference mode, A robot controller that communicates with the computer and receives the grasping instruction from the computer, A robot equipped with the gripper that grasps and moves an object based on an instruction from the computer. A system having
17. The system according to claim 16, wherein the grasping instruction includes 3D coordinates of a grasping target point, an approach angle that the gripper should follow, and a rotation angle and width of the gripper or positions of all finger joints.
18. Repeatedly executing iterative optimization includes discretizing the surfaces of the gripper and the object into points, calculating contact point pairs and collision pairs based on the current pose of the gripper with respect to the object, repeatedly calculating an optimization model, updating the pose of the gripper, and repeatedly calculating the optimization model until convergence to achieve grasping quality. The system according to claim 16.
19. The optimization model includes an objective function defined to maximize the grasping quality, and a constraint function that defines the contact surface of the gripper and the contact surface of the object as a subset of the surface of the gripper and a subset of the surface of the object, respectively, discretized into point sets transformed by the pose of the gripper with respect to the object. The contact surface of the gripper and the contact surface of the object are equal to each other, and the contact surface of the gripper and the contact surface of the object have friction defined by a friction cone model, a penalty for the gripper penetrating into the object, and a gripper joint angle within a limit angle. The system according to claim 18.
20. Executing a physical environment simulation includes mathematically simulating the fall of a plurality of objects into a bin, the simulation including gravity, contact between the objects, and the walls of the bin, and the resulting pile of objects in the bin being used in the physical environment simulation. The system according to claim 16.
Citation Information
Patent Citations
Robot simulation device for simulating work takeout process
JP2015168043A
Robot simulation device, robot simulation method, robot simulation program, computer-readable recording medium and recording device
JP2018144153A
Hand control device, hand control method, and hand simulation device
JP2019010713A
Deep machine learning method and apparatus for robotic grasping
JP2019508273A
Planning a Grasp Approach, Position, and Pre-Grasp Pose for a Robotic Grasper Based on Object, Grasper, and Environmental Constraint Data
US20140163731A1