Method and system for performing object pickup by robot with placement notification
By combining the image information of the picking scene and placement range in the robot picking system, the object-area pairing cost is calculated, and the inefficiency problem caused by independent operation of the picking and placement process in traditional systems is solved, and more efficient object selection and placement are achieved, and space utilization is optimized.
Patent Information
- Application Number
- CN202510184963.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-20
- Filing Date
- 2025-02-19
- Publication Date
- 2025-08-22
AI Technical Summary
Traditional robot pick-up and placement systems lack effective interactions in compact box packaging applications, resulting in inefficient space utilization and unoptimized object layout.
By obtaining images of the pick scene and placement range, using instance segmentation to calculate the object and placement area mask, and combining placement constraints to calculate the object-area pairing cost, select the most suitable object-area pairing to achieve efficient pickup and placement.
It improves the overall efficiency of the object selection and placement process, reduces the search space of the placement area, optimizes the space utilization, and is suitable for high-throughput applications.
Smart Images

Figure CN120516673A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to robotics in industrial automation tasks, and in particular, to systems and methods for implementing place-informed picking for robotic pick and place operations. Background Art
[0002] Box picking and packing are common operations in industrial warehouse automation, traditionally performed by humans due to the randomness and variability of the objects being handled. With recent advances in machine learning, 3D vision, and robotics, these tasks are increasingly being handled by autonomous robotic systems to achieve flexibility and higher performance.
[0003] The autonomous robotic picking system may include one or more RGB-D cameras that capture color images and depth maps or point clouds of a picking scene with randomly configured objects. The camera input may be passed to a computer vision algorithm or deep neural network that has been trained to select objects and calculate the robotic grasping position or "pickup point" on the selected object based on the camera input. Once the selected object has been successfully picked up, information about the picked object (e.g., size) may be passed to a placement system. The placement system may then determine a placement pose for placing the picked object in the placement box based on the size of the picked object, state information of the placement box (e.g., obtained from an image of the placement box), and considerations such as high space utilization and packaging stability.
[0004] In the state of the art described, the gripping system and the placement system typically have limited interaction. This can potentially lead to inefficient space utilization and other issues with optimally arranging objects within the placement box, particularly for compact box packaging applications. An improved system is needed. Summary of the Invention
[0005] Aspects of the present disclosure address and overcome one or more of the shortcomings described herein by providing methods, systems, and computer program products that enable a robotic picking system to make an intelligent decision about which object to pick from a picking scene by utilizing information about the placement range.
[0006] A first aspect of the present disclosure provides a method for performing placement-informed robotic object picking. The method includes acquiring a first image of a picking scene including a plurality of objects, and acquiring a second image of a placement range configured to accommodate objects selectively picked from the picking scene by a robotic end-effector. The method includes computing object masks based on the first image by performing instance segmentation, wherein each object mask represents a specific object detected in the picking scene. The method also includes computing placement region masks based on the second image by clustering locations in the second image based on height levels from a ground surface of the placement range, wherein each placement region mask represents a surface at a particular height level. The method also includes computing costs for corresponding object-region pairings based on the computed object masks and placement region masks, each object-region pairing defining a pairing between an object mask and a placement region mask, the cost being defined at least in part by one or more placement constraints. The method also includes selecting an object to be picked from the picking scene by the robotic end-effector by selecting the object-region pairing based on the computed cost.
[0007] This approach allows for better integration between the picking process and the placing process, enabling more efficient object selection and placement.
[0008] Other aspects of the present disclosure provide autonomous systems and computer program products for implementing the above methods.
[0009] Additional technical features and benefits can be achieved through the technology of the present disclosure. The embodiments and aspects of the present disclosure are described in detail herein and are considered to be part of the subject matter claimed. For a better understanding, please refer to the detailed description and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other aspects of the disclosure are best understood when the following detailed description is read in conjunction with the accompanying drawings.To easily identify discussion of any element or action, the most significant digit in a reference number refers to the figure number in which the element or action is first introduced.
[0011] Figure 1 A pick and place scenario is shown in which aspects of the present disclosure may be suitably implemented.
[0012] Figure 2 An autonomous system configured to perform placement-informed robotic picking of objects is schematically illustrated in accordance with one or more embodiments.
[0013] Figure 3 A pick and place workflow is shown in accordance with one or more implementations.
[0014] Figure 4The object mask computed from the picked bin image is shown.
[0015] Figure 5 The placement area mask computed from the pick bin image is shown.
[0016] Figure 6 is a flow diagram illustrating a process performed by an object selector according to one or more implementations.
[0017] Figure 7 The range utilization of the object mask over the region mask is shown.
[0018] Figure 8 A computing system that can support placement-informed robotic picking of objects in accordance with disclosed embodiments is shown. DETAILED DESCRIPTION
[0019] It should be recognized that currently, picking and placing are often considered separate processes in robotic pick and place pipelines and, therefore, are addressed separately. For example, the picking process can analyze the image of the pick bin to select the optimal object for picking. This decision is guided by picking constraints, such as the gripper type and stable grasping conditions (e.g., avoiding picking occluded objects, side collisions between the gripper and the bin, etc.). This approach can significantly improve bin picking efficiency in terms of cycle time and pick success rate. If the subsequent placement task involves placing the object on a conveyor or simply placing it in a bin or container, the decision about which object to select can be made without considering any placement constraints. In such use cases, the interdependence between the picking and placing processes may be minimal. However, in certain use cases, particularly for compact bin packaging, the subsequent placement task may be somewhat influenced by the picking process. This is because the picking process determines which object to select from the pick bin, while the placement process is responsible for estimating the optimal placement pose based on the selected object, taking into account placement constraints such as range utilization and stability to achieve efficient packaging.
[0020] Figure 1 An exemplary pick and place scenario for a compact bin packaging use case is shown. The task involves picking objects from a pick bin and placing them in a place bin in a singulated manner to achieve compact packaging of the placed objects. Figure 1The placement ranges H1, H2, H3, H4, etc. in the upper right corner represent corresponding surfaces at specific height levels from the ground of the placement box (e.g., derived from the placement box's segmented depth map). In the scenario shown, if decisions were simply guided by picking constraints (typically picking larger, unoccluded objects first), the picking process could select Object 1 from the picking box. The placement process would then be constrained to placing Object 1 on surface H2. This is because surface H2 provides the most solid foundational support for Object 1. However, when Object 1 is located on surface H2, surface H1 (at a lower height level) becomes inaccessible because it is partially obscured by Object 1. Conversely, if the picking process had already selected Object 2, it could be appropriately placed on surface H1, promoting efficient packing and effective use of space in the placement box.
[0021] The above scenario highlights the drawbacks of operating independently for the pick and place processes. The lack of effective communication between these two processes can lead to inefficient space utilization and potential issues with optimal object placement, especially in compact case packaging. The proposed method addresses some of these issues.
[0022] Overall, the proposed method enables a robotic picking system to make informed decisions about which objects to pick from a picking scene by leveraging information about a placement scope. Unlike the prior art, which only considers an image of a picking bin to make object selection decisions, the proposed method utilizes a first image of the picking scene and a second image of a placement scope as input. The picking scene can include multiple objects, which can be contained in a picking bin (as described herein) or otherwise arranged (e.g., on a table). The placement scope can be configured to receive objects selectively picked from the picking scene by a robotic end-effector. The placement scope can include, for example, a placement bin (as described herein), or alternatively, a tote, tray, or container. The placement scope can be initially empty and subsequently loaded one by one by the robotic end-effector with objects picked from the picking scene. Objects can initially be placed on the floor of the placement scope and eventually stacked on top of other objects. The first and second images can indicate the instantaneous states of the picking scene and the placement scope, respectively.
[0023] According to the proposed method, an object mask is calculated by performing instance segmentation based on a first image. A placement region mask is calculated by clustering locations in a second image based on their height level from the ground (or base) of the placement range. A "mask" generally refers to a set of pixels (e.g., representing an object or region) obtained through image segmentation. Based on the calculated object and placement region masks, a cost is calculated for each object-region pairing, where each object-region pairing defines a pairing between the object mask and the placement range mask. The cost of the object-region pairing is defined at least in part by the placement constraints. For example, the cost can be determined based on a heuristic method that includes one or more cost components representing one or more placement constraints. Examples of placement constraints can include range utilization (minimizing unused space), stability of the placed object (e.g., avoiding collisions, rollovers, etc.), and so on. Objects to be picked from the picking scene are selected by selecting an object-region pairing based on the calculated cost. For example, the object-region pairing with the highest (or lowest) cost can be selected, depending on how the cost is defined. This method can ensure a more balanced and efficient operation that optimizes both object selection and placement.
[0024] The innovative aspect of the proposed method lies in integrating the segmented placement zones (masks) from the placement range into the object selection process. This integration enables accurate selection of the most suitable objects to be picked, which can also improve packaging efficiency and, in turn, the overall performance of the end-to-end pick and place process. For example, by incorporating a range utilization cost component into the cost used for object-region pairing, the proposed method can be used to achieve higher space utilization within the placement range, as the objects being picked can be selected taking into account the available free area within the placement range. Due to the higher space utilization, the number of bins required to place a given set of objects can be significantly reduced.
[0025] Furthermore, according to the proposed method, the object selection process not only selects objects but also identifies paired placement regions. The placement region masks (output by the object selection process) from the selected object-region pairs can be directly used to calculate the placement poses needed to place the selected object within the placement range by the robot end-effector. By reducing the search space for placement regions, the time required for the downstream placement process to calculate the optimal placement pose can be significantly reduced. Consequently, the overall cycle time required to perform real-time pick and place operations can be reduced, making the proposed method suitable for high-throughput applications.
[0026] Aspects of the disclosed method can be implemented as software executable by a processor. In some embodiments, aspects of the disclosed method can be suitably integrated into commercial artificial intelligence (AI)-based automation software products, such as the SIMATIC Robot Pick AI developed by Siemens AG. TMwait.
[0027] Turning now to the disclosed embodiments, Figure 2 An autonomous system 200 is schematically illustrated, configured to perform placement-informed robotic object picking, according to one or more embodiments. Autonomous system 200 can be implemented, for example, in a factory environment. In contrast to traditional automation, autonomy empowers each asset on the factory floor with decision-making and self-control capabilities, enabling it to act independently when local issues arise. Autonomous system 200 includes one or more robots, such as robot 202, which can be controlled by a computing system 204 to perform one or more industrial tasks. Examples of industrial tasks include assembly and transportation.
[0028] The computing system 204 may include an industrial PC, or any other computing device, such as a desktop or laptop computer, or an embedded system, etc. The computing system 104 may include one or more processors configured to process information and / or control various operations associated with the robot 202. Specifically, the one or more processors may be configured to execute applications for operating the robot 202, such as an engineering tool.
[0029] To achieve autonomy in the system 200, in one embodiment, an application can be designed to operate the robot 202 to perform tasks in a skill-based programming environment. In contrast to traditional automation, where engineers are typically involved in programming the entire task from start to finish, typically using low-level code to generate individual commands, in autonomous systems such as those described in the present invention, physical devices, such as the robot 202, are programmed at a higher level of abstraction using skills rather than individual commands. These skills are derived for higher-level abstract behaviors, centered around how the programmed physical device changes the physical environment. Illustrative examples of skills include a skill to grasp or pick up an object, a skill to place an object, a skill to open a door, a skill to detect an object, and the like.
[0030] The application can generate controller code that defines the task at a high level, for example, using the skill function described above, which can be deployed to the robot controller 208. Based on the high-level controller code, the robot controller 208 can generate low-level control signals for one or more motors to control the movement of the robot 202, such as the angular position of the robot arm, the rotation angle of the robot base, etc., to perform the specified task. In other embodiments, the controller code generated by the application can be deployed to an intermediate control device, such as a programmable logic controller (PLC), which can then generate low-level control commands for the robot 202 to be controlled. In addition, the application can be configured to directly integrate sensor data from the physical environment in which the robot 202 operates. To this end, the computing system 204 may include a network interface to facilitate the transmission of live data between the application and various sensors (such as cameras 222, 226). The following is a description of the controller code generated by the application in conjunction with the accompanying drawings. Figure 8 An example of a computing system suitable for use with the present application is described.
[0031] The robot 202 may include a robotic arm or manipulator 210 and a base 212 configured to support the robotic manipulator 210. The base 212 may include wheels 214 or may be otherwise configured to move within the physical environment 206. The robot 202 may also include an end effector 216 attached to the robotic manipulator 210. The end effector 216 may include a gripper configured to grasp (hold) and pick up an object 218. Examples of end effectors include vacuum grippers (suction cups), antipodal grippers (fingers or claws), magnetic grippers, and the like. The robotic manipulator 210 may be configured to move to change the position of the end effector 216, thereby enabling the object 218 to be picked up and moved within the physical environment.
[0032] A robotic pick and place operation may involve the robotic manipulator 210 using the end effector 216 to pick objects 218 one by one from a pick bin 220 and place them in a place bin 224. Objects 218 may be arranged in random poses within the pick bin 220. Objects 218 may be of various types or of the same type. The placement operation may involve placing objects 218 in the place bin 224 in an orderly manner to achieve efficient packaging. In this case, the picking scene, including the pick bin 220 containing the cluttered objects 218, may be perceived via at least one first camera 222. Similarly, the place bin 224, including any objects 218 placed therein, may be perceived via at least one second camera 226. Cameras 222 and 226 may include, for example, RGB-D cameras. In some embodiments, the pick bin and place bin may be perceived by a single camera that can move between a first position above the pick bin 220 and a second position above the place bin 224.
[0033] Images acquired via cameras 222 and 226 can be provided as input to a computing system, such as computing system 204. Based on the input images, computing system 204 can select an object 218 to be picked up from pick bin 220 and estimate a pick point on the selected object 218 using the methods described herein. The estimated pick point can be output to a controller, such as robot controller 208, to control end effector 216 to pick up the selected object 218. For example, as described above, the pick point can be output to the controller as high-level controller code, which can generate low-level commands to control the movement of end effector 216. After successfully picking up the selected object 218, a placement pose can be estimated by computing system 204 for placing the selected object 218 in drop bin 224. The estimated placement pose can also be output to controller 208 for controlling end effector 216 to properly place the selected object 218 within drop bin 224.
[0034] Figure 3 A computer-implemented pick and place workflow 300 is shown, in accordance with one or more embodiments. The various modules described herein, such as the instance segmentation module 304, the drop zone extraction module 310, the object selection module 314, and the pick point estimation module 316, including their components, can be implemented by a computing system in various ways, for example, as hardware and programming. The programming for modules 304, 310, 314, and 316 can take the form of processor-executable instructions stored on a non-transitory, machine-readable storage medium, and the hardware can include a processor for executing these instructions. For example, the program can be run on an industrial PC or on a smaller device (e.g., a controller) in an autonomous system. Furthermore, processing power can be distributed across multiple system components, such as across multiple processors and memories, optionally including multiple distributed processing systems or cloud / network elements.
[0035] refer to Figure 3, the proposed method comprises acquiring a first image 302 of a picking scene comprising, for example, a plurality of objects arranged in a picking box. According to the disclosed embodiment, the first image 302 may comprise an image set comprising an intensity image and a corresponding depth image of the picking scene. The method further comprises acquiring a second image 308 of a placement range, for example, the placement range comprising a placement box configured to accommodate objects picked from the picking scene. According to the disclosed embodiment, the second image 308 may comprise an image set comprising an intensity image and a corresponding depth image of the placement range. However, for the purpose of implementing the proposed method, it is sufficient that the second image 308 comprises at least a depth image of the placement range. In an embodiment, the images 302, 308 are preferably acquired as a top view of the picking box and the placement box, respectively. The first image 302 and the second image 308 may define the input of the workflow 300.
[0036] An intensity image comprises a two-dimensional representation of image pixels, where each pixel comprises a single intensity value (monochrome image) or intensity values for multiple color components (color image). An example of a color intensity image is an RGB color image, which is an image that includes pixel intensity information in red, green, and blue channels. A depth image, also called a depth map, comprises a two-dimensional representation of image pixels, which includes a depth value for each pixel. The depth value corresponds to the distance of the surface of a scene object from the camera viewpoint. The intensity image and the corresponding depth image of the scene can be aligned pixel-wise. To this end, an RGB-D camera can be used, which can be configured to acquire images with red-green-blue (RGB) color and depth (D) channels.
[0037] In some embodiments, one or both of images 302, 308 can be obtained by acquiring a point cloud of the corresponding scene. A point cloud can include a set of points in a 3D coordinate system representing a 3D surface or multiple 3D surfaces, where each point position is defined by Cartesian coordinates in the real-world reference frame 230 and also by intensity values of color components (e.g., red, green, and blue). Thus, the point cloud can include a colored 3D representation of all surfaces in the corresponding scene. The point cloud can be converted into intensity (RGB) and depth images by applying a series of transformations based on camera intrinsic parameters. Camera intrinsic parameters are parameters that allow for a mapping between pixel coordinates in a 2D image frame and 3D coordinates in the real-world reference frame 230. Typically, camera intrinsic parameters include the coordinates of the principal point or optical center, and the focal length along orthogonal axes.
[0038] Based on the first image 302, the instance segmentation module 304 can perform instance segmentation to detect objects in the captured scene and thereby calculate corresponding object masks 306. Specifically, an intensity image of the captured scene can be utilized for this purpose. Instance segmentation essentially includes semantic segmentation and object detection, with the added feature of identifying object boundaries at a detailed pixel level. Given an input intensity image (e.g., an RGB color image), an instance segmentation model, such as a trained convolutional neural network, can be used to calculate an instance segmentation mask (referred to as an "object mask") corresponding to each object detected in the captured scene. Examples of instance segmentation models that can be used or adapted for this purpose include instance segmentation using the "Segment Anything Model" (SAM) developed by Meta AI, the "You Look Only One" (YOLO) model, the Mask R-CNN, and the like. Each object mask 306 calculated by the instance segmentation model can include a set of pixels associated with a specific object.
[0039] For illustration, refer to Figure 4 , input image 302 depicts a picking bin 220 containing multiple objects, including objects 218a, 218b, 218c, etc. Segmentation output 400 depicts an object mask computed using the instance segmentation model for each object detected in picking bin image 302. For example, object mask 306a includes a set of pixels associated with object 218a, object mask 306b includes a set of pixels associated with object 218b, and object mask 306c includes a set of pixels associated with object 218c.
[0040] Continue to refer Figure 3 Based on the second image 308 , the placement area extraction module 310 can calculate multiple placement area masks 312 by segmenting areas forming surfaces at the same height. Placement area masks 312 can be calculated by clustering locations in the image space of the second image 308 based on their height level from the ground (or base) of the placement range. This can be achieved by utilizing state-of-the-art clustering algorithms, such as Density-Based Spatial Clustering of Applications with Noise (DBSCAN) and k-means. Thus, each placement area mask 312 can represent a surface or plane at a specific height level.
[0041] The second image 308 may include a depth map of the placement range, wherein each pixel includes a depth value. The depth value of a pixel corresponds to the distance of the surface represented in the pixel from the camera viewpoint, which can be converted to a height from the ground of the placement range. To achieve this, the camera can be appropriately positioned to capture a top view of the ground of the placement range. If the camera is positioned at an angle relative to the ground of the placement range, the camera image can be appropriately reprojected to calculate the height from the depth value using a known transformation. Each placement area mask 312 calculated by the clustering algorithm can be defined by a group of adjacent pixels having depth values corresponding to a specific height level from the ground of the placement range. The height level can represent a height value within a defined tolerance band.
[0042] For illustration, refer to Figure 5 , the input image 308 depicts a placement box 224 in which a plurality of objects have been placed, including objects 218p, 218q, 218r, 218s, etc. The segmentation output 500 depicts five placement area masks 312a, 312b, 312c, 312d, 312e calculated by the placement area extraction module 310 based on the depth map of the input image. Each placement area mask 312a-e includes a set of neighboring pixels at a particular height level that define a corresponding surface. As shown, the surface defined by each placement area mask 312a-e can encompass a single object, multiple adjacently placed objects of the same height, or no objects (i.e., the ground of the placement box).
[0043] Still refer to Figure 3 , the calculated object mask 306 and placement region mask 312 can be provided as input to the object selection module 314. The object selection module 314 can calculate a corresponding cost for each of a plurality of object-region pairs. Each object-region pairing can include a pairing between an object mask 306 representing an object from the picking scene and a placement region mask 312 representing a region on the placement range. The selection of an object to be picked from the picking scene by the robot end effector can be performed based on the calculated cost of the object-region pairing. Unlike state-of-the-art methods in which object selection is based solely on picking constraints, in the disclosed embodiment, object selection is based on a cost calculated for each object-region pairing, which includes at least one placement constraint. This can ensure a more balanced and efficient operation of the system while optimizing object selection and downstream placement processes.
[0044] Figure 6A process 600 for selecting an object based on placement constraints is shown according to one or more embodiments. Activity blocks 602-610 of process 600 may be performed by a computing system including one or more processors. In one embodiment, activity blocks 602-610 may be performed by object selection module 314 of computer-implemented workflow 300 described herein. Figure 6 It is not intended to indicate that the activity blocks of process 600 are to be performed in any particular order, or that all activity blocks of process 600 must be included in every case. Furthermore, process 600 may include any suitable number of additional operations.
[0045] At block 602, a pickability metric can be calculated for each object mask 306. The pickability metric can indicate a success rate for picking a specific object associated with the object mask 306. The pickability measurement can guide object selection to be performed so that the object being picked is not obscured. In this way, the chance of a successful pick can be maximized by minimizing object friction. At the same time, it can be ensured that the topmost object is not accidentally pulled out of the bin. The pickability metric for the object mask 306 can include, for example, a pickability score for the object mask, or a binary label for the object mask ("pickable" or "unpickable"), or a ranking of the object mask, or any combination thereof.
[0046] The pickability metric for each object mask 306 can be calculated using depth information from the corresponding depth map. For example, the object mask 306 calculated by the instance segmentation model can be used to segment the pixel-aligned depth map of the picked scene. The pickability metric for the object mask 306 can then be calculated based on the corresponding segmented depth map. This can be achieved in a variety of ways.
[0047] For example, in one embodiment, a heuristic method may be used to calculate a pickability metric for an object mask using the corresponding segmented depth map. In most cases, the object to be picked is preferably the topmost object that is typically unoccluded. To detect the topmost object, the object's depth (e.g., derived from the depth map) can provide the strongest signal. Therefore, the heuristic method may include a depth metric for the object mask. The depth metric may include, for example, the average depth value, the maximum depth value, the minimum depth value, or any combination thereof, of the pixels in the object mask. Furthermore, it may often be desirable to remove larger objects as early as possible. Therefore, the heuristic method may also include a size metric for the object mask. The size metric may be defined, for example, by the range covered by all pixels of the object mask. In one implementation, the heuristic method may include a combination (e.g., a weighted combination) of the depth metric for the object mask, the size of the object mask, and the confidence of the predicted object mask to determine a pickability metric for each object mask. The result may be a single object mask, a ranked list of object masks, or a list of individually labeled object masks with a binary label ("pickable" or "unpickable").
[0048] In another embodiment, a trained neural network or other machine learning model may be used to calculate a pickability metric for the object masks 306 using the corresponding segmented depth maps. The neural network / machine learning model may also provide an output that includes a binary label for each mask ("pickable" or not "pickable") or a ranking of object masks from most pickable to least pickable.
[0049] At block 604, a search space for object-region pairings can be determined. To achieve an optimal solution for selecting objects, an exhaustive search space for object-region pairings can be determined. In this case, the search space can include all possible object-region pairings from the entirety of the calculated object mask 306 and the calculated placement region mask 312. However, in practice, particularly in use cases involving a large number of objects, it may be desirable to constrain the search space to improve computational efficiency and cycle time. This can be achieved by selecting object-region pairings only from a subset of the calculated object mask 306 and / or a subset of the calculated placement region mask 312.
[0050] In one embodiment, a pickability metric for the object masks 306 (e.g., as determined at block 602) can be utilized to select a subset of the computed object masks 306 for use in determining the search space. For example, if the pickability metric comprises a binary label ("pickable" or "not pickable"), this step can be relatively simple, wherein only object masks 306 labeled "pickable" can be selected. If the pickability metric comprises a score or ranking, a cutoff score / rank can be defined to select the subset of the computed object masks 306. Additionally or alternatively, the search space can be constrained by ranking the object masks 306 based on size and / or ranking the placement area masks 312 based on size. The size of the mask can be defined in terms of pixel area. Thus, the search space can be determined by selecting a subset of the computed object masks 306 based on pixel area and / or a subset of the computed placement area masks 312 based on pixel area.
[0051] Before calculating the cost of an object-region pairing, it may be valuable to confirm whether the object can be placed anywhere within the region. At block 606, an object fit check can be performed to determine whether the kernel defined by the object mask 306 fits within the region defined by the placement region mask 312 of the object-region pairing. The object's kernel can be defined, for example, as a two-dimensional array with all entries assigned a value of "1," where the size of the array is equal to the size (in pixels) of the object mask 306. In one embodiment, the object fit check can involve a convolution operation, in which the object's kernel is convolved with the placement region mask 312 to determine whether the kernel can fit completely within the placement region mask 312 (without colliding with other regions). The convolution operation can be performed for different orientations of the object's kernel (e.g., 0 degrees and 90 degrees). Before performing this operation, it can be ensured that the object's kernel and the placement region mask have the same scale factor / resolution. Block 606 can ensure that the object can indeed fit within the potential placement region before performing any further calculations.
[0052] At block 608, a cost can be calculated for each object-region pairing in the search space for which object fit has been determined, i.e., the object kernel can fit within the object-region pairing within the region of the placement region mask. For example, the cost can be determined based on a heuristic method that includes one or more cost components representing one or more placement constraints. Examples of placement constraints can include range utilization (minimizing unused space), stability of the placed object (e.g., avoiding collisions, tumbling, etc.), etc.
[0053] In one embodiment, the heuristic cost for an object-region pairing can include a first cost component representing a first placement constraint, i.e., extent utilization. When a new object is placed in a region and the size of the region exceeds the size of the object, there is remaining, unused space. In order to achieve compact packing, objects and regions should be selected in a way that minimizes this unused space. Therefore, the first cost component can be defined accordingly so that it indicates the utilization of the extent of the placement region mask 312 by the object mask 306 in a given object-region pairing. For example, the first cost component can be defined by an area factor calculated as the ratio of the pixel area of the object mask 306 to the pixel area of the placement region mask 312 in a given object-region pairing. That is:
[0054] Area factor = pixel area of object mask / pixel area of region mask (1)
[0055] The area factor provides a heuristic estimate of the spatial efficiency of placing a specific object in a specific area, thereby guiding the object selection process towards maximizing space utilization. Figure 7 In the illustrated scenario, there may be two object-region pairs, namely, a first pair comprising an object mask 306p and a placement region mask 312q, and a second pair comprising an object mask 306p and a placement region 312r. Based on equation (1), the second object-region pair 306p, 312r will have a higher area factor and, therefore, a higher cost, which may encourage the object selection process to prioritize the second object-region pair 306p, 312r over the first object-region pair 306p, 312q.
[0056] In another embodiment, the heuristic cost for an object-region pairing may further include a second cost component representing a second placement constraint, namely, stability. The stability constraint can ensure that the placed object remains in place. Experimental observations have shown that stability generally depends on altitude level. Therefore, the second cost component may indicate the altitude level of the placement region mask 312 for a given object-region pairing.
[0057] Each placement area within the placement range has a specific height level. Experimental observations have shown that filling objects at lower height levels before moving to higher height levels minimizes potential collisions and tumbles and also achieves better space utilization. Therefore, a second cost component can be defined such that lower height levels carry higher costs (and vice versa). Thus, the second cost component can encourage the object selection process to prioritize objects that can be accommodated at lower height levels before moving to objects at higher height levels.
[0058] In some embodiments, the heuristic cost may include a weighted combination of the first and second cost components, such as a weighted sum. Depending on the use case, the weights assigned to the respective cost components may be selected. For example, in a use case involving packaging of fragile objects, stability constraints may be of paramount importance. Therefore, in such a use case, a relatively higher weight may be assigned to the second cost component. Conversely, for example, in a use case involving packaging of substantially flat objects, space utilization may be more important. In such a use case, a relatively higher weight may be assigned to the first cost component. In various embodiments, depending on the use case, the heuristic cost may additionally or alternatively include one or more other cost components representing one or more other placement constraints.
[0059] Still refer to Figure 6 , at block 608, the cost of each object-region pairing can be additionally combined with a pickability metric for the object mask 306 in the given object-region pairing. The pickability metric can be determined as described in conjunction with block 602. The pickability metric can be incorporated into the heuristic cost in several possible ways. For example, the pickability metric (e.g., a score or ranking) can be incorporated into the above-mentioned heuristic cost by weighted addition. For another example, in the case of binary labeling ("pickable" = 1 or "unpickable" = 0), the pickability metric can be incorporated into the above-mentioned heuristic cost as a multiplier. In some embodiments, particularly where the pickability metric has been used to constrain the search space (e.g., by selecting only object masks with the binary label "pickable"), it may not be necessary to subsequently incorporate the pickability metric in the heuristic cost. In this case, the heuristic cost can be calculated using only the placement constraints, as described above.
[0060] At block 610, an object to be picked from the picking scene is selected by selecting an object-region pairing based on the calculated heuristic cost. For example, according to disclosed embodiments, the object-region pairing with the highest cost may be selected. In other embodiments, depending on how the heuristic cost is defined, the object-region pairing with the lowest cost may be selected.
[0061] Reference again Figure 3 The output of the object selection module 314 may include the selected object-region pairing, including the object mask 306 Sel and paired placement area mask 312 Sel .
[0062] Object Mask 306 SelThe pick point estimation module 316 can be used to estimate the pick point of the robot end effector. In one embodiment, the pick point estimation module 316 can include a grasping neural network to calculate the grasp position at which the end effector picks up the selected object. The grasping neural network is often convolutional so that the network can label each pixel of the input image with some type of grasp affordance metric, known as a grasp score. In this case, the input image can include using the object mask 306 Sel A segmented depth map of the picked scene is computed. A grasp score for a pixel can indicate the quality of the grasp at the location defined by the pixel, which generally represents a confidence level for performing a successful grasp (e.g., not dropping the object). Based on the pixel-by-pixel grasp score, the optimal grasp position for the end effector can be determined as the pick point based on defined constraints (e.g., avoiding collisions with the bin walls). A grasping neural network can be trained on a dataset that includes depth maps of the object or scene from various camera positions and ground truth labels that include pixel-by-pixel grasp scores for a given type of gripper for the end effector.
[0063] In other embodiments, the object mask 306 may be used Sel The key points in the image are used to model the picked points. Key point detection can be performed, for example, from the intensity image using a neural network, which can be embedded in the instance segmentation model or a standalone model. Alternatively, a non-deep learning approach can be used to model the picked points. For example, the object mask 306 Sel The centroid of can be used to model the picking point.
[0064] The output 318 of the pick point estimation module 316 may include the coordinates (X, Y, Z) of the pick point in the real-world reference frame 230. To calculate the coordinates (X, Y, Z), the object mask 306 may be transformed into the real-world reference frame 230 using the depth information from the segmented depth map and the camera intrinsic parameters. Sel The pick point calculated in the two-dimensional image space of the real-world reference frame 230 is projected into the three-dimensional space of the real-world reference frame 230. Output 318 may also include an approach vector calculated based on the pick point, which specifies the approach direction. For example, if the pick point is located on a flat surface, the approach direction can be calculated as the normal vector of the surface. Output 318, including the pick point coordinates (X, Y, Z) and the approach vector, can be provided to the controller to control the robot end effector to pick up the selected object.
[0065] Placement area mask 312 for the selected object-area pair Sel Can be utilized by downstream placement pose estimation process. After successfully picking the selected object, the placement area mask 312 SelAs well as information about the selected object, such as estimated planar dimensions and grasp offset, is provided as input to the placement pose estimation process. The placement pose estimation process can employ known techniques for calculating an optimal placement pose based on the above inputs. However, by limiting the search space of available areas in the placement range to the placement region mask 312 already calculated by the object selection module 314, the placement pose estimation process can be performed. Sel , which significantly reduces the time required to calculate the optimal placement pose, thus helping to reduce the overall cycle time.
[0066] Figure 8 An example of a computing system capable of supporting placement-informed robotic picking of objects in accordance with the disclosed embodiments is shown. The computing system 204 may be implemented as, for example, but not limited to, an industrial PC with a Linux operating system for performing real-time control of the robot. The computing system 204 includes at least one processor 810, which may take the form of a single processor or multiple processors. The processor 810 may include one or more CPUs, GPUs, microprocessors, or any hardware device suitable for executing instructions stored on a memory including a machine-readable medium. The computing system 204 also includes a machine-readable medium 820. The machine-readable medium 820 may take the form of one or more media, including any non-volatile electronic, magnetic, optical, or other physical storage device for storing executable instructions, such as instance segmentation instructions 822, placement zone extraction instructions 824, object selection instructions 826, and pick point estimation instructions 828, as described herein. Figure 8 Thus, the machine-readable medium 820 may be, for example, a random access memory (RAM) such as dynamic RAM (DRAM), flash memory, spin transfer torque memory, electrically erasable programmable read-only memory (EEPROM), a storage drive, an optical disk, or the like.
[0067] The computing system 204 can execute instructions stored on the machine-readable medium 820 via the processor 810. Executing the instructions (e.g., the instance segmentation instructions 822, the placement area extraction instructions 824, the object selection instructions 826, and the pick point estimation instructions 828) can cause the computing system 204 to perform any of the technical features described herein, including any of the features of the instance segmentation module 304, the placement area extraction module 310, the object selection module 314, and the pick point estimation module 316 described above.
[0068] The systems, methods, devices, and logic described above, including the instance segmentation module 304, the placement region extraction module 310, the object selection module 314, and the pick point estimation module 316, can be implemented in a variety of different ways, in a variety of different combinations of hardware, logic, circuitry, and executable instructions stored on a machine-readable medium. A product, such as a computer program product, can include a storage medium and machine-readable instructions stored on the medium, which, when executed in an endpoint, computer system, or other device, causes the device to perform operations according to any of the above descriptions, including any features according to the instance segmentation module 304, the placement region extraction module 310, the object selection module 314, and the pick point estimation module 316. The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network, or downloaded to an external computer or external storage device.
[0069] The systems, devices, and modules described herein, including the processing capabilities of the instance segmentation module 304, the placement region extraction module 310, the object selection module 314, and the pick point estimation module 316, can be distributed across multiple system components, such as across multiple processors and memories, optionally including multiple distributed processing systems or cloud / network elements. Parameters, databases, and other data structures can be stored and managed separately, combined into a single memory or database, logically and physically organized in a variety of ways, and implemented in a variety of ways, including data structures such as linked lists, hash tables, or implicit storage mechanisms. Programs can be part of a single program (e.g., subroutines), separate programs distributed across multiple memories and processors, or implemented in a variety of different ways, such as as libraries (e.g., shared libraries).
[0070] Although the present disclosure has been described with reference to specific embodiments, it should be understood that the embodiments and variations shown and described are for illustrative purposes only. Modifications to the present design may be implemented by those skilled in the art without departing from the scope of the patent claims.
Claims
1. A method for performing placement-informed robotic picking of an object, comprising: acquiring a first image (302) of a picked-up scene (220) including a plurality of objects (218), acquiring a second image (308) of a placement range (224) configured to accommodate an object (218) selectively picked up by a robotic end effector (216) from the picking scene (220), Based on the first image (302), object masks (306) are calculated by performing instance segmentation, wherein each object mask (306) represents a specific object detected in the picked scene (220), Based on the second image (308), calculating placement area masks (312) by clustering locations in the second image (308) based on height levels from the ground of the placement range (224), wherein each placement area mask (312) represents a surface at a specific height level, Based on the computed object mask (306) and the placement region mask (312), computing costs for corresponding object-region pairings (306, 312), each object-region pairing (306, 312) defining a pairing between the object mask (306) and the placement region mask (312), the costs being defined at least in part by one or more placement constraints, and By selecting an object-region pairing based on the calculated cost (306 Sel , 312 Sel ), thereby selecting an object (218) to be picked up by the robot end effector (216) from the picking scene (220).
2. The method according to claim 1, wherein The second image (308) includes a depth map, and wherein each calculated placement area mask (312) is defined by a set of neighboring pixels having depth values corresponding to a particular height level from the ground of the placement range (224).
3. The method according to any one of claims 1 and 2, wherein The cost of the object-region pairing (306, 312) includes a first cost component representing a first placement constraint of the one or more placement constraints, the first cost component indicating utilization of the extent of the placement region mask (312) by the object mask (306) in a given object-region pairing (306, 312).
4. The method according to claim 3, wherein: The first cost component is defined by a ratio of a pixel area of the object mask (306) to a pixel area of the placement region mask (312) in the given object-region pair (306, 312).
5. The method according to any one of claims 3 and 4, wherein: The cost of the object-region pairing further includes a second cost component representing a second placement constraint of the one or more placement constraints, the second cost component indicating a height level of the placement region mask (312) in the given object-region pairing (306, 312).
6. The method according to claim 5, wherein: The cost of the object-region pairing (306, 312) includes a weighted combination of the first cost component and the second cost component.
7. The method according to any one of claims 1 to 6, comprising, before calculating the cost of the object-region pairing (306, 312), performing a check to determine whether a kernel defined by the object mask (306) fits within a region defined by the placement region mask (312) of the object-region pairing (306, 312).
8. The method according to any one of claims 1 to 7, comprising calculating for each object mask (306) a pickability metric indicative of picking success, wherein The cost of the object-region pairing (306, 312) incorporates the pickability metric of the object mask (306) in a given object-region pairing (306, 312).
9. The method according to any one of claims 1 to 8, comprising calculating for each object mask (306) a pickability metric indicative of picking success, wherein A search space of object-region pairs (306, 312) for computing the cost is determined using a subset of the computed object masks (306) selected based on the computed pickability metric.
10. The method according to any one of claims 8 and 9, in, The first image (302) comprises an intensity image and a depth map of the picked-up scene (220), wherein the object mask is calculated by performing instance segmentation based on the intensity image (306), and wherein, for each object mask (306), the pickability metric is calculated using depth information obtained from the depth map of the picked scene (220).
11. The method according to any one of claims 1 to 10, wherein A search space for object-region pairs (306, 312) for computing the cost is determined using a subset of the computed object masks (306) and / or a subset of the computed placement region masks (312) selected based on the size-ordered object masks (306) and / or the size-ordered placement region masks (312).
12. The method according to any one of claims 1 to 11, further comprising: Using the selected object-region pairing (306 Sel , 312 Sel ) in the object mask (306 Sel ) to estimate the pick-up point of the robot end effector (318), and The estimated pick-up point (318) is output to the controller (208) to control the robot end effector (216) to pick up the selected object (218).
13. The method according to any one of claims 1 to 12, wherein Using the selected object-region pairing (306 Sel , 312 Sel ) in the placement area mask (312 Sel ) to calculate a placement posture for placing the selected object (218) in the placement range (224) by the robot end effector (216).
14. A non-transitory computer-readable storage medium comprising instructions which, when executed by one or more processors, configure the one or more processors to perform the method according to any one of claims 1 to 13.
15. An autonomous system (200) configured to perform placement-informed robotic picking of an object, comprising: A robot (202) comprising an end effector (216), One or more cameras (222, 226) configured to acquire a first image (302) of a pickup scene (220) including a plurality of objects (218), and acquire a second image (308) of a placement range (224) configured to accommodate objects (218) selectively picked up from the pickup scene (220) by the end effector (216), one or more processors (810), and A memory (820) storing instructions, wherein the instructions are executable by the one or more processors (810) to: Based on the first image (302), object masks (306) are calculated by performing instance segmentation, wherein each object mask (306) represents a specific object detected in the picked scene (220), Based on the second image (308), locations in the second image (308) are clustered based on height levels from the ground of the placement range (224) to calculate placement area masks (312), wherein each placement area mask (312) represents a surface at a specific height level, Based on the calculated object mask (306) and placement region mask (312), calculating a cost for corresponding object-region pairings (306, 312), wherein each object-region pairing (306, 312) defines a pairing between the object mask (306) and the placement region mask (312), the cost being defined at least in part by one or more placement constraints; and By selecting an object-region pairing based on the calculated cost (306 Sel , 312 Sel ) to select an object (218) to be picked up by the end effector (216) from the picking scene (220).