Method and system for robot to pick up object
Through example segmentation and plane representation technology, estimating the picking posture of the robot end effector solves the problem of poor contact between the end effector and the object, improves the picking accuracy and calculation efficiency, and is suitable for autonomous robot systems.
Patent Information
- Application Number
- CN202510118807.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-29
- Filing Date
- 2025-01-24
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art is difficult to efficiently estimate the pick-up posture of robot end effectors of any size, especially when there is no prior information on the object geometry and color, traditional methods are prone to errors and have high calculation costs.
The picking point is estimated by the instance segmentation model, the picking surface is determined using the proximity point in the object mask, reprojected to the plane representation, and the deflection orientation is calculated based on the dimension alignment of the end effector model, and the picking attitude is output to the controller.
It realizes that the accuracy and computing efficiency of pick-up pose estimation are improved without the need for large amount of training data and prior knowledge of 3D models, and is suitable for high-throughput real-time applications.
Smart Images

Figure CN120382478A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure is mainly related to robotics in industrial automation tasks, and in particular to systems and methods for pick pose estimation for robotic picking using end effectors of any size. Background Art
[0002] The goal of the Fourth Industrial Revolution is to drive mass customization to reduce the costs of mass production. This can be achieved by autonomous machines that no longer have to be programmed with detailed instructions (such as waypoints or manually taught paths), but instead use the design information of the product to be produced to automatically define their tasks. Pick-and-place in a bin is a technology that enables autonomous machines. Traditional robotic picking relies on teach-based methods, enabling an operator to pre-determine the robotic poses for pick-up and drop-off positions. In the past decade, advancements in computer vision and deep learning have enabled flexible robotic pick-and-place in a bin, where pre-learned pick and drop positions are no longer required. Camera systems such as RGB-D cameras collect color images and depth maps or point clouds of bins with objects in random configurations. The camera input is then fed into a computer vision algorithm or a deep neural network that has been trained to compute the grasping position or "pick point" on the input. These methods have proven to work well and reliably even without any prior information about the object geometry and color (CAD data is not required).
[0003] Typically, robotic end effectors need to have a relatively large (elliptical) coverage area that contacts the object, making them suitable for picking objects of certain sizes and / or textures. In such cases, it is desirable to estimate the optimal pick pose of the end effector to achieve maximum robustness in picking. Summary of the Invention
[0004] Aspects of the present disclosure provide methods and systems for estimating the deflected orientation pick pose of an end effector of any size for performing robotic picking of an object.
[0005] A first aspect of the present disclosure provides a method for a robot to pick up an object. The method includes acquiring one or more images of a scene via an imaging system, the scene including one or more objects. The method further includes estimating a pick-up point on an object mask by a computing system including one or more processors, the object mask being generated by performing instance segmentation based on the one or more images. The object mask corresponds to an object selected from the one or more objects to be picked up by an end effector or a robot. The method further includes estimating a pick-up pose of the end effector by the computing system, wherein the end effector is modeled as a 2D shape with a specified size. The pick-up pose estimation includes: determining a pick-up surface using neighboring points around the pick-up point in the object mask; re-projecting a set of points defining the extent of the pick-up surface in the object mask with respect to the normal of the pick-up surface to create a planar representation of the pick-up surface; and calculating a deflection orientation based on the alignment of the longer dimension of the end effector model with the longer dimension of the planar representation of the pick-up surface. The method further includes outputting the estimated pick-up pose by the computing system to a controller configured to control the end effector to pick up the selected object.
[0006] Other aspects of the present disclosure provide an autonomous system and a computer program product for implementing the above method.
[0007] Additional technical features and benefits can be achieved through the techniques of the present disclosure. Embodiments and aspects of the present disclosure are described in detail herein and are considered to be part of the claimed subject matter. For a better understanding, reference is made to the detailed description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The foregoing and other aspects of the present disclosure are best understood from the following detailed description when read in conjunction with the drawings. To facilitate easy identification of the discussion of any element or action, the most significant digit or digits of the reference numbers refer to the figure number in which the element or action is first introduced.
[0009] Figure 1 An autonomous system for robotic picking is shown in accordance with one or more embodiments.
[0010] Figure 2A And Figure 2B An example configuration of a robotic end effector suitable for implementing aspects of the present disclosure is shown.
[0011] Figure 3 is a high-level block diagram showing a computer-implemented workflow for estimating a pick-up pose in accordance with one or more embodiments.
[0012] Figure 4 is a flowchart showing a process of a pick-up pose determined by a pick-up point estimator to calculate a deflection orientation in accordance with one or more embodiments.
[0013] Figure 5 Shows a planar representation of a pick surface generated from a point cloud and the alignment of an end effector model therewith.
[0014] Figure 6 Shows a computing system according to the disclosed embodiments that can support performing robotic pickups using end effectors of any size. Detailed Description
[0015] Various techniques for robotic pickup applications using end effectors of any size are described in the present invention. These end effectors can have a rectangular footprint that contacts the object to be picked up. For example, the end effector can include an array of grasping elements (e.g., suction cups, magnetic grippers, etc.) that can define a rectangular footprint. In other examples, the end effector can be defined by a single grasping element having a rectangular footprint (e.g., rectangular, oval, elliptical, etc.). The size of the end effector can be configured based on usage. In such cases, pick points parameterized only by their (X, Y, Z) coordinates and their normal vectors may not be sufficient. To achieve highly robust pickups, it is important that the contact between the end effector and the object surface is maximized. To ensure this, it is necessary to estimate a pick pose further characterized by a deflected orientation. The deflected orientation can define the angular orientation of the end effector in the plane of the pick surface.
[0016] If the 3D model of the object is known in advance, the problem is somewhat straightforward. Conventional computer vision algorithms or self-organizing trained neural networks can be used to perform object pose estimation. Once the pose of the object is known, a pick pose can be derived such that the axis of the end effector is aligned with the axis of the object. If the 3D model of the object is unknown, an object diagnostic instance segmentation model can be used to estimate the object segmentation mask. The pick pose is derived by aligning a rectangular shape (i.e., the shape of the end effector) with the object segmentation mask. However, the segmentation mask can include surfaces other than the pick surface and, in turn, can produce suboptimal pick poses. This technique can also be error-prone, especially when the object is in an abnormal position or in a very cluttered / occluded scene.
[0017] The present disclosure addresses one or more of the disadvantages described herein by providing methods and systems for obtaining a deflected orientation pick pose for end effectors of any size by leveraging pick points and instance segmentation masks in combination with 3D point cloud manipulation and 2D image processing.
[0018] According to the disclosed method, the end effector is modeled as a 2D shape with specified dimensions, which can represent its occupied area on the object to be picked up. The end effector dimensions can define the input of the computer-implemented workflow described in the present invention. In the workflow, pick-up points are estimated on a three-dimensional segmentation mask (referred to herein as the "object mask") of the object to be picked up. The object mask is generated by performing instance segmentation on one or more acquired images of a scene including one or more objects. The object mask can include, for example, a point cloud representation of the selected object. Then, neighboring points around the pick-up points in the object mask are used to determine the pick-up surface. The pick-up surface can represent the contact surface between the selected object and the end effector. Next, a set of points defining the extent of the pick-up surface in the object mask is re-projected with respect to the normal of the pick-up surface to create a planar representation of the pick-up surface. The planar representation generated in this way can be free of camera perspective distortion. Subsequently, a deflected orientation pick-up pose is calculated based on the alignment of the longer dimension of the end effector model with the longer dimension of the planar representation of the pick-up surface. The deflected orientation can represent the rotation angle applied to the end effector to align its longer dimension with the longer dimension of the pick-up surface of the selected object. The estimated pick-up pose is output to a controller configured to control the end effector to pick up the selected object.
[0019] According to the disclosed embodiments, the pick-up pose output to the controller can be defined by a set of parameters, including: the position coordinates of the center of the end effector determined based on the alignment of the end effector model, the normal vector of the pick-up surface, and the deflected orientation defining the angular orientation of the end effector in the plane of the pick-up surface. For example, in one embodiment, the output pick-up pose can explicitly specify the above parameters. In another embodiment, the output pick-up pose can specify a six-degree-of-freedom (6D) pose calculated based on the above parameters using a known transformation.
[0020] Different from the above method, the disclosed method does not rely on deep learning methods that require a large amount of training data and / or prior knowledge of the 3D model of the object, and provides higher accuracy even for a chaotic or abnormal arrangement of objects in the scene. In addition, by re-projecting the 3D object mask cloud into a 2D image including a planar representation of the pick-up surface, the computational cost is significantly reduced (e.g., by implementing image processing via basic 2D computer vision operations), making the solution suitable for high-throughput real-time applications.
[0021] Aspects of the disclosed method can be implemented as software executable by a processor. In some embodiments, aspects of the disclosed method can be appropriately integrated into commercial artificial intelligence (AI)-based automation software products, such as SIMATIC Robot Pick AI developed by Siemens AG TM and so on.
[0022] Turning now to the drawings, Figure 1 illustrates an autonomous system 100 for robotic picking according to one or more embodiments. The autonomous system 100 may be implemented, for example, in a factory setting. Compared to traditional automation, autonomy gives each asset on the factory floor the decision-making and self-control ability to act independently in local problem situations. The autonomous system 100 includes one or more robots, such as robot 102, which may be controlled by a computing system 104 to perform one or more industrial tasks within a physical environment 106. Examples of industrial tasks include assembly, transportation, and the like.
[0023] The computing system 104 may include an industrial PC, or any other computing device, such as a desktop or laptop computer, or an embedded system, etc. The computing system 104 may include one or more processors configured to process information and / or control various operations associated with the robot 102. In particular, one or more processors may be configured to execute application programs for operating the robot 102, such as engineering tools.
[0024] To achieve the autonomy of the system 100, in one embodiment, the application program may be designed to operate the robot 102 to perform tasks in a skill-based programming environment. Compared to conventional automation in which engineers typically participate in programming the entire task from start to finish (usually using low-level code to generate individual commands), in the autonomous system as described in the present invention, physical devices such as the robot 102 are programmed at a higher level of abstraction with skills rather than individual commands. These skills are obtained for high-level abstract behaviors centered on how to modify the physical environment with the programmed physical device. Illustrative examples of these skills include the skill of grasping or picking up an object, the skill of placing an object, the skill of opening a door, the skill of detecting an object, and the like.
[0025] The application program may generate controller code that defines high-level tasks, for example, using the skill functions described above, and the controller code may be deployed to the robot controller 108. According to the high-level controller code, the robot controller 108 may generate low-level control signals for one or more motors to control the movement of the robot 102, such as the angular position of the robot arm, the rotational angle of the robot base, etc., to perform the specified task. In other embodiments, the controller code generated by the application program may be deployed to an intermediate control device such as a programmable logic controller (PLC), and the intermediate control device may then generate low-level control commands for the robot 102. Additionally, the application program may be configured to directly integrate sensor data from the physical environment 106 in which the robot 102 operates. To this end, the computing system 104 may include a network interface facilitating the transfer of live data between the application program and the physical environment 106. The following is combined with Figure 6Describe an example of a computing system applicable to the present application.
[0026] Still referring to Figure 1 , the robot 102 may include a robotic arm or manipulator 110 and a base 112 configured to support the robotic manipulator 110. The base 112 may include wheels 114 or may be configured to move within the physical environment 106. The robot 102 may further include an end effector 116 attached to the robotic manipulator 110. The end effector 116 may include one or more grasping elements 122 configured to grasp (hold) and pick up an object 118. Examples of the grasping elements include vacuum grippers ("suction cups"), magnetic grippers, etc. The one or more grasping elements 122 may be configured such that the end effector 116 has an oblong footprint in contact with the object 118 to be picked up. The object 118 to be picked up may be placed in a bin 120 together with other objects 118. The robotic manipulator 110 may be configured to move to change the position of the end effector 116, for example, to pick up and move the object 118 within the physical environment 106.
[0027] Figure 2A A first exemplary embodiment of the end effector 116A is shown. The end effector 116A includes an array of grasping elements 122A. In this case, each grasping element 122A is a suction cup. In the illustrated embodiment, the grasping elements 122A are identical to each other, each having a cylindrical shape defining a circular contact area. In other embodiments, the array may be formed by different grasping elements or grasping elements having other shapes. The array of grasping elements 122A defines a rectangular footprint of the end effector 116A having a specified length (L) and width (W). Figure 2B A second exemplary embodiment of the end effector 116B is shown. The end effector 116B includes a single (relatively large) grasping element 122B. In this case, the grasping element 122B is a suction cup. The grasping element 122B defines an oblong contact area. In the illustrated embodiment, the grasping element 122B has a rectangular shape having a specified length (L) and width (W) that defines the footprint of the end effector 116B. In other embodiments, the single grasping element 122B may have a different shape, such as oval, ovoid, or any other oblong shape. The dimensions (e.g., length and width) of each type of end effector 116A, 116B are configured depending on the usage.
[0028] Continuing to refer to Figure 1, the robotic pick-up operation may involve using the end effector 116 to grasp the object 118 from the bin 120 in an individualized manner via the robotic manipulator 110. The object 118 may be arranged in the bin 120 in any pose. The objects 118 may be of various types or the same type. The physical environment 106 including the objects 118 placed in the bin 120 is sensed via an imaging system, which may include at least one camera 122. As shown, the camera 122 may be mounted, for example, to the end effector 116. The imaging system including the camera 122 may be used to acquire one or more images of the scene, which may be provided as input to a computing system such as the computing system 104 for estimating the pick-up pose of the end effector 116. The estimated pick-up pose may be output to a controller, such as the robotic controller 108, to control the end effector 116 to pick up the selected object 118. For example, as described above, the pick-up pose may be output to the controller as high-level controller code, and the controller may thereby generate low-level commands to control the movement of the end effector 116.
[0029] Figure 3 A computer-implemented workflow 300 for estimating a pick-up pose according to one or more disclosed embodiments is shown. The various modules described herein, such as the instance segmentation module 304, the object selection module 306, the pick-up point estimation module 308, and the pick-up pose estimation module 310, including their components, may be implemented by a computing system in various ways, such as as hardware and programming. The programming of the modules 304, 306, 308, 310 may take the form of processor-executable instructions stored on a non-transitory machine-readable storage medium, and the hardware may include a processor that executes these instructions. For example, the program may run on an industrial PC or on a smaller device of an autonomous system (e.g., a controller). Additionally, the processing capabilities may be distributed among multiple system components, such as distributed among multiple processors and memories, optionally including multiple distributed processing systems or cloud / network elements.
[0030] Referring to Figure 3 , the disclosed method includes acquiring one or more images 302 of the scene via an imaging system. In one embodiment, the one or more images 302 may include a color intensity image and a depth image of the scene. The color intensity image includes a two-dimensional representation of image pixels, where each pixel includes intensity values of multiple color components. An example of a color intensity image is an RGB color image, which is an image including pixel intensity information in the red, green, and blue channels. The depth image, also referred to as a depth map, includes a two-dimensional representation of image pixels, which includes a depth value for each pixel. The depth value corresponds to the distance of the surface of the scene object from the camera viewpoint. The color intensity image and the depth image may be pixel-aligned. For example, in some embodiments, a single RGB-D camera may be configured to acquire an image of the scene with RGB color and depth channels.
[0031] In another embodiment, one or more images 302 may include a point cloud of a scene. The point cloud may include a set of points in a 3D coordinate system representing a 3D surface or multiple 3D surfaces, where each point position is defined by its Cartesian coordinates in the real-world reference frame 124 (see Figure 1 ).) and is further defined by the intensity values of color components (e.g., red, green, and blue). Thus, the acquired point cloud 302 may include a colored 3D representation of all surfaces in the scene. The point cloud 302 may be acquired, for example, via an RGB-D camera.
[0032] One or more images 302 and the camera intrinsic parameters may define the input of the workflow 300. The camera intrinsic parameters are parameters that allow mapping between pixel coordinates in a 2D image frame and 3D coordinates in the real-world reference frame 124. Generally, the camera intrinsic parameters include the coordinates of the principal point or optical center, and the focal lengths along the orthogonal axes. By applying a sequence of transformations based on the camera intrinsic parameters, the point cloud can be converted into a corresponding color intensity and depth image, and vice versa.
[0033] In a first step, the instance segmentation module 304 performs instance segmentation based on one or more images 302 to detect objects in the scene and thereby compute corresponding object masks. Instance segmentation mainly includes semantic segmentation and object detection, which have the additional feature of identifying object boundaries at the detailed pixel level. Given an input color intensity image, an instance segmentation model of, for example, a trained convolutional neural network can be used to compute an instance segmentation mask corresponding to each object detected in the scene. Examples of instance segmentation models that can be used or are suitable for the purposes of the present invention include instance segmentation using the "Segment Anything Model" (SAM) developed by Meta AI, the "You Look Only Once" (YOLO) model, the Mask Recursive Convolutional Neural Network (Mask R-CNN), etc. Each instance segmentation mask computed by the model may include a flat (2D) mask that includes a set of pixels representing a particular object. The planar instance segmentation mask can be used to segment the pixel-aligned depth map of the scene to generate a 3D object mask for each object in the scene. In one embodiment, the 3D object mask may include a point cloud representation of a particular object, which is obtained from the segmented depth map using the camera intrinsic parameters described above.
[0034] Next, the object selection module 306 can be used to select an object for picking from a list of object masks. Object selection is ideally performed such that the object being picked is isolated (not occluded). In this way, by minimizing the object friction, the chance of a successful pick can be maximized. Also, the topmost object is not accidentally pulled out of the bin. For the above purposes, a pickability metric can be determined for each object mask, and this pickability metric can be used to select the object to be picked. For example, the pickability metric can include a pickability score for each object mask, or a binary label (pickable or non-pickable) for each object mask, or a rank of the object mask, or any combination thereof.
[0035] In one embodiment, a heuristic approach can be used to determine the pickability metric of an object mask. In most cases, the object to be picked is preferably the topmost object, which is typically not occluded. To expose the topmost object, the depth of the object (e.g., obtained from a depth map) can provide the strongest signal. Thus, the heuristic approach can include depth measurement of the object mask. The depth measurement can include, for example, the average or maximum or minimum depth value of the pixels in the mask, or any combination thereof. Also, it may often be desirable to get larger objects out of the way sooner rather than later. Thus, the heuristic approach can also include a size metric of the object mask. For example, the size metric can be defined by the area covered by all the pixels of the object mask. In one embodiment, the heuristic approach can include a combination (e.g., weighted combination) of the depth metric of the object mask, the size of the object mask, and the confidence of the predicted object mask to determine the pickability metric for each object mask. The result of the object selection module 306 can be an object mask, a list of sorted object masks, or a list of individually labeled object masks with binary labels (pickable or non-pickable). Having a list allows for the immediate calculation of pick points for multiple objects, which is beneficial for parallelizing the work or in cases where the top-level object is considered non-pickable due to safety or robotic workspace constraints.
[0036] In an alternative embodiment, a trained neural network or other machine learning model can be used to determine the pickability metric of an object mask. The neural network / machine learning model can similarly provide an output including a binary label for each mask (pickable or non-pickable) or an ordered list of object masks from most pickable to least pickable.
[0037] Next, the pick point estimation module 308 estimates pick points on the object mask of the selected object. In one embodiment, the pick point estimation module 308 may include a grasping neural network to compute a grasping position for the end effector to pick the selected object. The grasping neural network is typically convolutional such that the network can label each pixel of the input image with a type of grasping affordance metric, called a grasping score. The input image typically includes a depth map. The grasping score of a pixel represents the grasping quality at the location defined by the pixel, which typically represents a confidence level for performing a successful grasp (e.g., not dropping the object). Based on the pixel-wise grasping scores, the best grasping position of the end effector can be determined based on defined constraints (e.g., avoiding collision with the bin wall). The grasping neural network can be trained on a dataset that includes depth maps of objects or scenes from various camera positions and ground truth labels that include pixel-level grasping scores for a given type of picker for the end effector. A non-limiting example of a grasping neural network suitable for the purposes of the present invention is disclosed in International Patent Application No. PCT / US2023 / 013550 filed by the present applicant, which is hereby incorporated by reference in its entirety.
[0038] According to an embodiment disclosed by the present invention, the end effector may include, for example, an array of the same grasping elements as shown Figure 2A in. The size (length and width) of the array can be configured based on usage. In this case, a grasping neural network trained on a single type of grasping element (e.g., a cylindrical suction cup) can be used to compute pick points for the individual grasping elements of the array. Given a segmented depth map obtained from the input image 302, the trained grasping neural network can be used to determine the best grasping positions of the individual grasping elements. The best grasping positions or pick points computed on the segmented depth map can be projected onto the 3D space of the real-world reference frame 124 using the depth information from the segmented depth map and the camera intrinsic parameters to locate pick points on the object mask. Thus, the method of the present invention in the embodiment utilizes pick points of only a single grasping element of the array for pick pose estimation, such that the method can be scaled to an array of any size (configurable based on usage) without having to train / re-train the grasping neural network in each case.
[0039] In other embodiments, key points in the instance segmentation mask can be used to model the pick-up points. For example, a neural network can be used to perform key point detection from the color intensity image. The neural network can be embedded in the instance segmentation model or be an independent model. Alternatively, non-deep learning methods can be employed to model the pick-up points. As an example, the centroid of the instance segmentation mask can be used to model the pick-up points. The key points / centroids calculated on the planar instance segmentation mask can be projected onto the 3D space of the real-world reference frame 124 using the depth information from the segmented depth map and the camera intrinsic parameters to locate the pick-up points on the object mask.
[0040] Still referring to Figure 3 , the end effector can be modeled as a 2D shape that appropriately represents the contact occupancy area of the end effector. For example, as shown, the end effector model 312 can include a rectangular shape with a specified length (L) and width (W), which can be adapted to describe an end effector including an array of grasping elements. Other rectangular end effectors that can include single or multiple grasping elements can also be modeled as rectangles. For example, an elliptical or oval end effector can be modeled as a rectangle with lengths and widths specified by the major and minor axes, respectively. In other embodiments, other 2D shapes can be appropriately employed to model the end effector. The dimensions of the end effector model 312 define another input to the workflow 300.
[0041] The pick-up pose estimation module 310 uses the object mask determined at 304, the estimated pick-up points determined at 308, and the specified dimensions (e.g., L, W) of the end effector model 312 as inputs to estimate the pick-up pose of the end effector. As described in detail below, the pick-up pose estimation module 310 can perform a series of operations based on the above inputs to obtain the deflected orientation pick-up pose 314. The pick-up pose 314 can be defined by the coordinates that define the center of the end effector in the real-world reference frame, the normal vector of the pick-up surface of the object, and the angular direction of the end effector in the plane of the pick-up surface (
[0042] Figure 4 Figure 400 shows a process 400 for determining a deflected orientation pick-up pose according to one or more disclosed embodiments. The activity boxes 402-408 of the process 400 can be executed by a computing system including one or more processors. In one embodiment, the activity boxes 402-408 can be executed by the pick-up pose estimation module 310 of the computer-implemented workflow 300 described herein.
[0043] Block 402 involves determining a picked surface using neighboring points around a picked point in an object mask. The object mask may include a point cloud representation of a selected object derived from a segmented depth map of the scene. The picked surface may be defined by a plane. A set of neighboring points may be selected around the picked point in the point cloud to calculate the plane equation.
[0044] The number of proximity points or reach can be determined based on the use case. For example, where the end effector includes an array of gripping elements, the number of proximity points or reach can be determined based on the size of a single gripping element. Figure 2A In the example shown, the set of neighboring points can be selected so that the maximum distance from the pick point does not exceed the radius of the suction cup 122A. In another embodiment, the number of neighboring points or the reach can be determined based on the minimum dimension (eg, width W) of the end effector model 312.
[0045] Given a set of neighboring points, the plane equation can be determined, for example, using a least squares best fit to the points or other regression method. The plane equation can define a picking surface.
[0046] Block 404 includes determining a set of points that define the extent of the picked surface on the object mask (to be reprojected in a subsequent step). Not all points in the object mask necessarily belong to the picked surface. For example, the object may be a tilted box, and several faces of the box may be visible and part of the object mask. The steps aim to find the limits of the picked surface once the plane equation has been determined. In one embodiment, a clustering method based on heuristics combined with distance to the plane, normal classification and other geometric properties can be used for this purpose. The set of points to be reprojected can be obtained by removing all points in the object mask that do not belong to the picked surface, for example, as determined by a clustering method. In addition, in order to determine the set of points to be reprojected, it may be advantageous to remove outliers that contribute to noisy measurements relative to the plane of the picked surface, for example, using a statistical outlier filter.
[0047] Block 406 includes reprojecting the determined set of points relative to the normal of the picked surface to create a planar representation of the picked surface. In one embodiment, the determined set of points in the object mask can first be projected into a depth map. The transformation can be performed using camera intrinsic parameters. The points in the depth map can then be rotated relative to the normal of the picked surface. As a result of the rotation, a 2D image can be generated with a viewing direction perpendicular to the picked surface, i.e., the picked surface is aligned with the camera frame of the 2D image. In this way, camera perspective distortion can be removed.
[0048] refer to Figure 5Describe the above steps. Here, image 502 represents the 2D projection of a point cloud using the camera's intrinsic parameters. Image 502 depicts a scene that includes a box, and the box includes a box located on its right wall. Image 502 is essentially a depth map, which is a 2D representation of 3D points. That is, in addition to the x and y coordinates, each point in image 502 is also characterized by depth information. Reference numeral 504 represents the set of all points in the pick-up surface 506 of the selected object (box). Image 508 represents a 2D image generated by the rotation of points 504 relative to the normal of the pick-up surface 506. Image 508 has an observation direction perpendicular to the pick-up surface 506. That is, the plane of the pick-up surface 506 has been rotated so that it is now aligned with the camera frame of image 508. As shown, the pick-up surface 506, which appears trapezoidal in image 502 due to perspective distortion, appears approximately rectangular in image 508 after the points 504 are re-projected along the pick-up surface normal direction.
[0049] In one embodiment, the planar representation of the pick-up surface can be created by processing a 2D image to generate a contour representing the outline of the pick-up surface. The processing of the 2D image can involve any operation for obtaining an enhanced image, or otherwise extracting useful information to generate the contour. For example, a contour can be generated from the re-projected points by performing basic 2D computer vision operations such as filling, patching, and opening operations, etc., to recover lost points or gaps. The planar representation of the pick-up surface can be created by fitting a primitive shape (typically a rectangle) of the minimum area that includes all the points in the generated contour. In some embodiments, the primitive shape can be directly fitted to the contour without additional operations to recover lost points or gaps. The planar dimensions of the pick-up surface can be calculated, for example, by measuring the dimensions of the fitted primitive shape.
[0050] In Figure 5 the example shown, a contour 510 is generated on the 2D image 508 generated by the re-projection of points 504. As shown, the contour 510 lacks corners. This may occur, for example, due to incorrect depth imaging and / or insufficient points in the point cloud. In the example, the missing corners are recovered by fitting a minimum area rectangle 512 that includes all the points in the contour 510. The rectangle 512 defines the planar representation of the pick-up surface on the 2D image 508.
[0051] Referring again to Figure 4 , the box 408 includes calculating a deflection orientation by aligning the longer dimension of the end effector model with the longer dimension of the planar representation of the pick-up surface. The deflection orientation defines the angular orientation of the end effector in the plane of the pick-up surface. In particular, the deflection orientation can represent the rotation angle applied to the end effector to align its longer dimension with the dimensions of the pick-up surface.
[0052] Continuing Figure 5For an example, rectangle 514 represents the end effector model, which is overlaid on the pick-up surface 506 for illustration purposes. As shown, the rectangle 514 of the end effector model is centered with the rectangle 512 representing the pick-up surface plane, such that the long side 518 of the end effector model 514 is aligned with the long side 520 of the pick-up surface plane representation 512. Based on the alignment of the end effector model, the deflected pick-up pose can be calculated as follows. By transforming the coordinates of 516 from the image reference frame 524 to the real-world reference frame 124, the position of the end effector can be calculated based on the center 516 of the aligned actuator model 514. The approaching direction can be determined by calculating the normal vector of the pick-up surface 506 in the real-world reference frame 124. To calculate the deflected orientation, it is assumed that the end effector stays in a known initial pose in the real-world reference 124 and thus in the image reference frame 524. For example, here it can be assumed that the end effector (represented by the end effector model 514) is stationary such that its longer dimension is initially aligned with the U-axis in the image reference frame 524. The deflected orientation can be determined by aligning its longer dimension with the dimension of the pick-up surface plane representation 512 by applying a rotation angle to the actuator model 514 in the image reference frame 524. In the depicted example, the deflected orientation of the resulting end effector will become 90 degrees.
[0053] Referring to Figure 4 , at block 410, it can be determined whether there is a complete overlap between the end effector model and the pick-up surface plane representation. If a complete overlap is determined at block 410, the calculated pick-up pose can be output to the controller at block 412. The output pick-up pose can include the position coordinates defining the center of the end effector in the real-world reference frame , the normal vector of the pick-up surface and the angular orientation of the end effector in the plane of the pick-up surface ( ). In some embodiments, the pick-up pose can be determined as a 6D pose in the real-world reference frame, which can be calculated using known transformations based on the above parameters.
[0054] If a complete overlap is not determined at block 410, i.e., the end effector is too large for the object, the pick-up pose and the object can be rejected at block 414. In this case, the control can return to workflow 300 to select a different object mask and / or the end effector size can be modified.
[0055] Figure 6Shows an example of a computing system 600 that can support robotic picking using end effectors of any size according to the disclosed embodiments. The computing system 600 can be implemented, for example but not limited to, as an industrial PC with a Linux operating system for performing real-time control of a robot. The computing system 600 includes at least one processor 610, which can take the form of a single or multiple processors. The processor 610 can include one or more CPUs, GPUs, microprocessors, or any hardware device suitable for executing instructions stored on a memory including a machine-readable medium. The computing system 600 also includes a machine-readable medium 620. The machine-readable medium 620 can take the form of one or more media, including any non-transitory electronic, magnetic, optical, or other physical storage device storing executable instructions, such as instance segmentation instructions 622, object selection instructions 624, pick point estimation instructions 626, and pick pose estimation instructions 626, as Figure 6 shown. Thus, the machine-readable medium 620 can be, for example, random access memory (RAM), such as dynamic RAM (DRAM), flash memory, spin transfer torque memory, electrically erasable programmable read-only memory (EEPROM), a storage drive, an optical disc, etc.
[0056] The computing system 600 can execute the instructions stored on the machine-readable medium 620 through the processor 610. Executing the instructions (e.g., instance segmentation instructions 622, object selection instructions 624, pick point estimation instructions 626, and pick pose estimation instructions 626) can cause the computing system 600 to perform any of the technical features described in the present invention, including any features according to the above-mentioned instance segmentation module 304, object selection module 306, pick point estimation module 308, and pick pose estimation module 310.
[0057] The above systems, methods, devices, and logics including the instance segmentation module 304, object selection module 306, pick point estimation module 308, and pick pose estimation module 310 can be implemented in many different ways in many different combinations of hardware, logic, circuits, and executable instructions stored on a machine-readable medium. A product such as a computer program product can include a storage medium and machine-readable instructions stored on the medium, which when executed in an endpoint, computer system, or other device, cause the device to perform operations according to any of the above descriptions, including any features according to the instance segmentation module 304, object selection module 306, pick point estimation module 308, and pick pose estimation module 310. The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network).
[0058] The processing capabilities of the systems, devices, and modules described herein (including the instance segmentation module 304, the object selection module 306, the pick point estimation module 308, and the pick pose estimation module 310) can be distributed among multiple system components, such as among multiple processors and memories, optionally including multiple distributed processing systems or cloud / network elements. Parameters, databases, and other data structures can be stored and managed separately, can be combined into a single memory or database, can be organized logically and physically in many different ways, and can be implemented in many ways, including data structures such as linked lists, hash tables, or implicit storage mechanisms. Programs can be multi-parts of a single program (e.g., subroutines), separate programs, distributed across several memories and processors, or implemented in many different ways, such as in the form of libraries (e.g., shared libraries).
[0059] Although the present disclosure has been described with reference to specific embodiments, it should be understood that the embodiments and variations shown and described herein are for illustrative purposes only. Those skilled in the art can implement modifications to the current design without departing from the scope of the patent claims.
Claims
1. A method for a robot to pick up an object, comprising: Obtaining one or more images of a scene via an imaging system, the scene including one or more objects, Executed by a computing system including one or more processors: Estimating a pick-up point on an object mask, the object mask being generated by performing instance segmentation based on the one or more images, the object mask corresponding to an object selected from the one or more objects to be picked up by an end effector of the robot, Estimating a pick-up pose of the end effector, wherein the end effector is modeled as a 2D shape with a specified size, including: Using neighboring points around the pick-up point in the object mask to determine a pick-up surface, Reprojecting a set of points defining the extent of the pick-up surface in the object mask with respect to the normal of the pick-up surface to create a planar representation of the pick-up surface, and Calculating a deflection orientation based on the alignment of the longer dimension of the end effector model with the longer dimension of the planar representation of the pick-up surface, and Outputting the estimated pick-up pose to a controller configured to control the end effector to pick up the selected object.
2. The method according to claim 1, wherein, The end effector includes an array of grasping elements modeled as a rectangular shape with a specified length and a specified width.
3. The method according to claim 2, wherein Calculating the estimated pick-up point using a grasping neural network based on the one or more images to determine the best grasping position on the object mask for a single grasping element.
4. The method according to any one of claims 1 to 3, wherein The object mask is generated by: Calculating one or more instance segmentation masks for detecting the one or more objects in the scene based on the one or more images, where each instance segmentation mask includes a set of pixels representing a specific object, Using the one or more instance segmentation masks to segment a depth map of the scene obtained from the one or more images to thereby generate a point cloud representation of the selected object.
5. The method according to any one of claims 1 to 4, wherein The scene includes multiple objects, and wherein the method includes selecting an object from the multiple objects by determining a pick-upability metric of the object mask corresponding to each object in the multiple objects to ensure that the selected object to be picked up is not occluded.
6. The method according to any one of claims 2 to 5, wherein Determining the number or reach of the neighboring points around the pick-up point in the object mask based on the size of a single grasping element.
7. The method according to any one of claims 1 to 6, wherein Obtaining the reprojected set of points by removing points in the object mask that do not belong to the pick-up surface based on a clustering method.
8. The method according to any one of claims 1 to 7, wherein, Creating the planar representation of the pick-up surface includes: Projecting the set of points in the object mask into a depth map, and Rotating the points in the depth map with respect to the normal of the pick-up surface to generate a 2D image with an observation direction perpendicular to the pick-up surface.
9. The method according to claim 8, wherein, Creating the planar representation of the pick-up surface further includes processing the 2D image to generate a contour representing the profile of the pick-up surface.
10. The method according to claim 9, wherein, The contour is generated from the reprojected points by filling, or patching, or opening operation, or a combination thereof.
11. The method according to any one of claims 9 and 10, wherein Creating the planar representation of the pick-up surface further includes fitting a primitive shape of minimum area that includes all the points in the contour and thereby estimating the planar dimensions of the pick-up surface.
12. The method according to any one of claims 1 to 11, including outputting the estimated pick-up pose to the controller based on determining a complete overlap between the aligned end-effector model and the planar representation of the pick-up surface.
13. The method according to any one of claims 1 to 12, wherein, The pick-up pose output to the controller is defined by: the position coordinates of the center of the end-effector determined based on the alignment of the end-effector model, the normal vector of the pick-up surface, and the deflection orientation defining the angular orientation of the end-effector in the plane of the pick-up surface.
14. A non-transitory computer-readable storage medium comprising instructions that, when processed by one or more processors, configure the one or more processors to perform the method according to any one of claims 1 to 13.
15. An autonomous system for robotic pick-up, comprising: an imaging system configured to acquire one or more images of a scene, the scene including one or more objects, a robot including an end-effector that can be controlled by a controller, one or more processors, and a memory storing instructions executable by the one or more processors for: estimating pick-up points on an object mask generated by performing instance segmentation based on the one or more images, the object mask corresponding to an object selected from the one or more objects to be picked up by the end-effector of the robot, estimating the pick-up pose of the end-effector, wherein the end-effector is modeled as a 2D shape with specified dimensions, including: using neighboring points around the pick-up points in the object mask to determine the pick-up surface, re-projecting a set of points defining the extent of the pick-up surface in the object mask with respect to the normal of the pick-up surface to create a planar representation of the pick-up surface, and calculating the deflection orientation based on the alignment of the longer dimension of the end-effector model and the longer dimension of the planar representation of the pick-up surface, and outputting the estimated pick-up pose to the controller to control the end-effector to pick up the selected object.