Robotic system and method of grasping an object in a quantity of randomly placed objects
Patent Information
- Application Number
- PCT/CA2026/050396
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-14
- Filing Date
- 2026-03-13
- Publication Date
- 2026-09-17
Smart Images

Figure CA2026050396_17092026_PF_FP_ABST
Abstract
Description
ROBOTIC SYSTEM AND METHOD OF GRASPING AN OBJECT IN A QUANTITY OF RANDOMLY PLACED OBJECTS FIELD
[0001] The application relates to robotic systems and more specifically relates to robotic systems involving computer vision to automate robotized object manipulation.BACKGROUND
[0002] Automation of tasks using robots to pick and place objects may be convenient and efficient in at least some circumstances. For instance, the automation of operations performed by a robot to repeat sequences of movements or a predetermined trajectory with precision and speed has already been and continues to be a frequently discussed topic in the industry. One may think of applications on finely tuned assembly or manufacturing lines and the like. Although existing robotic systems are satisfactory to a certain degree, there always remain room for improvement.SUMMARY
[0003] In some applications, the objects to be manipulated by an automated robotic system have a known position and orientation relative to a pick-up area. However, some other applications which are less deterministic in nature could also benefit from the use of automated robotic systems. For instance, when the objects to be manipulated are not stacked in an orderly fashion, but rather stacked in a disorderly pile randomly placed objects. However, existing automated robotic systems are not equipped to face such challenges. Indeed, risks of collision and / or possible interference with nearby objects, in addition to the fact that not all of the objects in the pile are “graspable,” i.e., not all objects of the pile do possess a position and / or orientation which would allow grasping by an end effector or a robot, are factors which can impede the use of conventional automated robotic systems in these indeterministic situations. There is thus a need in the industry to improve existing automated robotic systems to provide automated robotic systems able to pick up objects even when they are in a pile of randomly placed objects, and do so in a manner which is computationally efficient thereby allowing real time or quasi-real-time operation of the automated robotic systems. Additionally or alternately, there is a need for improvements in the operation of a robot in applicationswhere the tasks to be performed require a selection and grasping of objects with randomized positions and orientations.
[0004] In accordance with a first aspect of the present disclosure, there is provided a method of grasping randomly placed objects using a robot, the method comprising: using a camera facing a plurality of randomly placed objects, capturing a given image of a plurality of randomly placed objects; using a computing device, segmenting the given image into a plurality of image segments, and comparing the image segments to a plurality of template images representing an object matching the randomly placed objects, each template image showing the object in a corresponding one of a plurality of different orientations, said comparing including identifying some image segments at least partially showing the object in any of the plurality of different orientations; processing the some image segments until a given image segment showing a graspable object is found, said processing including determining whether the corresponding image segments show an exposed grasping area of the object in a desired grasping orientation, and determining a position and orientation of the graspable object from the given image segment; and instructing the robot to perform a step of grasping, with the end effector, the graspable object shown in the given image segment based on the position and orientation of the graspable object.
[0005] Further in accordance with the first aspect of the present disclosure, the method can for example further comprise, after said grasping, an end effector moving away from the plurality of randomly placed objects to position the graspable object at a remote drop off location.
[0006] Still further in accordance with the first aspect of the present disclosure, the method can for example further comprise the camera capturing another given image of the randomly placed objects which have been repositioned upon removal of the graspable object, and performing another iteration of said method to grasp another graspable object.
[0007] Still further in accordance with the first aspect of the present disclosure, said capturing the other given image can for example be performed immediately after removal of the graspable object and prior to positioning of the graspable object at a remote drop off location.
[0008] Still further in accordance with the first aspect of the present disclosure, the given image can for example include at least one of a two-dimensional image and a depth image.
[0009] Still further in accordance with the first aspect of the present disclosure, the two-dimensional image can for example be a red-green-blue (RGB) image.
[0010] Still further in accordance with the first aspect of the present disclosure, the given image can for example be a two-dimensional image, said processing being performed in a decreasing order of size of the object in the corresponding image segments.
[0011] Still further in accordance with the first aspect of the present disclosure, the given image can for example be a depth image, said processing being performed in an increasing order of depth of the object in the corresponding image segments.
[0012] Still further in accordance with the first aspect of the present disclosure, said processing can for example further include: comparing the object shown in a given image segment to the template images until a matching template image is found, determining a coarse orientation of the object based on an orientation of the matching template image, wherein said processing is further based on the coarse orientation of the object.
[0013] Still further in accordance with the first aspect of the present disclosure, the method can for example further comprise estimating a position and orientation of the object based on the given image and the coarse orientation, generating a three-dimensional model of the object using the position and orientation, rendering a test image showing the object at the estimated position and orientation, and comparing the object of the test image to the object of the given image segment.
[0014] Still further in accordance with the first aspect of the present disclosure, said processing can for example include rejecting the given image segment when the object of the test image and the object of the given image segment do not match to one another.
[0015] Still further in accordance with the first aspect of the present disclosure, said processing can for example further include rendering a test image showing the object in saidposition and orientation, and determining an overlap between the test image and the given image segment.
[0016] Still further in accordance with the first aspect of the present disclosure, said rendering can for example include graphically displaying the test image over the given image segment.
[0017] Still further in accordance with the first aspect of the present disclosure, said determining the overlap can for example include an intersection over union determination between the test image as rendered and the given image segment.
[0018] Still further in accordance with the first aspect of the present disclosure, the method can for example further comprise one of: rejecting the given image segment when the overlap is below an overlap threshold, and confirming the orientation as the desired grasping orientation when the overlap is above the overlap threshold.
[0019] Still further in accordance with the first aspect of the present disclosure, the overlap threshold can for example be is at least 65%, preferably at least 75%, and most preferably at least 85%.
[0020] Still further in accordance with the first aspect of the present disclosure, said processing can for example include rejecting a given image segment upon determining that a robot path leading to the object amidst the plurality of randomly placed objects is obstructed.
[0021] Still further in accordance with the first aspect of the present disclosure, said processing can for example include rejecting a given image segment upon determining that a grasping area of the object in the given image segment is occluded.
[0022] Still further in accordance with the first aspect of the present disclosure, the template images can for example have been generated using a computer model of the object.
[0023] Still further in accordance with the first aspect of the present disclosure, said comparing can for example include comparing the image segments to a plurality of descriptor features of the object in a plurality of different orientations.
[0024] Still further in accordance with the first aspect of the present disclosure, said descriptor features can for example have been generated based on the template images.
[0025] Still further in accordance with the first aspect of the present disclosure, when the given image is captured, the camera can for example have a first viewpoint relative to the randomly placed objects, and wherein said step of grasping the graspable object by the robot can for example include moving towards the graspable object along a path at least partially originating from the first viewpoint.
[0026] In accordance with a second aspect of the present disclosure, there is provided a robotic system for grasping randomly placed objects, the robotic system comprising: a camera facing a pick-up area encompassing a plurality of randomly placed objects, the camera capturing a given image of a plurality of randomly placed objects; a robot having an end effector movable within the pick-up area and configured for performing a grasping motion at a given position and orientation of the pick-up area; and a controller communicatively coupled to the camera and to the robot, the controller: a segmentation module configured for segmenting the given image into a plurality of image segments, and comparing the image segments to a plurality of template images representing an object matching the randomly placed objects, each template image showing the object in a corresponding one of a plurality of different orientations, said comparing including identifying some image segments at least partially showing the object in any of the plurality of different orientations; a graspable object finding module configured for processing the some image segments until a given image segment showing a graspable object is found, said processing including determining whether the corresponding image segments show an exposed grasping area of the object in a desired grasping orientation, and determining a position and orientation of the graspable object from the given image segment; and a robot controlling module configured for instructing the robot to perform a step of grasping the graspable object shown in the given image segment based on the position and orientation of the graspable object; wherein the robot grasps the graspable object at the position and orientation thereof in response to said instructing.
[0027] Further in accordance with the second aspect of the present disclosure, said processing can for example further include: comparing the object shown in a given image segment to the template images until a matching template image is found, determining acoarse orientation of the object based on an orientation of the matching template image, wherein said processing is further based on the coarse orientation of the object.
[0028] Further in accordance with the second aspect of the present disclosure, the robotic system can for example further comprise estimating a position and orientation of the object based on the coarse orientation, generating a three-dimensional model of the object using the position and orientation, rendering a test image showing the object at the estimated position and orientation, and comparing the object of the test image to the object of the given image segment.
[0029] In accordance with a third aspect of the present disclosure, there is provided a method of grasping randomly placed objects using a robot, the method comprising: using a camera facing a plurality of randomly placed objects, capturing a given image of a plurality of randomly placed objects; using a computing device, segmenting the given image into a plurality of image segments; using the plurality of image segments, performing a grasping routine including: identifying a given image segment showing the object; determining a position and orientation of the object in the given image segment; rendering a test image of the object in said position and orientation over the given image segment; determining an overlap between the test image and the given image segment; and upon determining that the overlap exceeds an overlap threshold, instructing the robot to perform a step of grasping the object shown in the given image segment based on the position and orientation of the object.
[0030] Further in accordance with the third aspect of the present disclosure, the method can for example further comprise performing another iteration of the grasping routine for another one of the plurality of image segments when the overlap is below the overlap threshold.
[0031] All technical implementation details and advantages described with respect to a particular aspect of the present invention are self-evidently mutatis mutandis applicable for all other aspects of the present invention.
[0032] Many further features and combinations thereof concerning the present improvementswill appearto those skilled in the art following a reading of the instant disclosure.DESCRIPTION OF THE FIGURES
[0033] In the figures,
[0034] Fig. 1 is a side elevation view of an example of a robotic system, showing a pile of randomly placed objects in a container, a camera, a robot and a controller, in accordance with one or more embodiments;
[0035] Fig. 1A is a top plan view of the container of Fig. 1 containing the pile of randomly placed objects, taken along viewpoint 1A-1A of Fig. 1 , in accordance with one or more embodiments;
[0036] Fig. 1B is a perspective view of one of the randomly placed objects of Fig. 1A, overlaid with an annotation showing a corresponding position and orientation having six degrees of freedom including three for position (translational movements along axes x, y, and z) and three for orientation (rotational movements about perpendicular axes a, p, and y), in accordance with one or more embodiments;
[0037] Fig. 2 is a schematic view of the robotic system of Fig. 1 , showing a segmentation module, a graspable object finding module, and a robot controlling module, in accordance with one or more embodiments;
[0038] Fig. 3 is a schematic view of an example template image generation module of the controller of Fig. 1 , in accordance with one or more embodiments;
[0039] Fig. 4 is a flow chart of an example of a method for grasping randomly placed objects using a robot, in accordance with one or more embodiments;
[0040] Fig. 5A is an example of a red-green-blue (RGB) image of a pile of randomly placed objects, in accordance with one or more embodiments;
[0041] Fig. 5B is the RGB image of Fig. 5A over which image segments have been overlaid for ease of visualization purposes, in accordance with one or more embodiments;
[0042] Figs. 6A-6D are schematic views each showing a corresponding image segment, and outputs of a comparison of the corresponding image segment to template images of theobject, showing that only the image segments of Figs. 6B-6D at least partially show the object to be manipulated, in accordance with one or more embodiments;
[0043] Fig. 7 is a block diagram of an example of a method for grasping randomly placed objects using a robot, showing a pose estimating module and a test image rendering module, in accordance with one or more embodiments
[0044] Fig. 8A is the RGB image of Fig. 5A over which is overlaid a first test image including a rendering of an object positioned at a first coarse position and oriented at a first coarse orientation, showing a failed comparison, in accordance with one or more embodiments;
[0045] Fig. 8B is the RGB image of Fig. 5B over which is overlaid a second test image including a rendering of an object positioned at a second coarse position and oriented at a second coarse orientation, showing a successful comparison, in accordance with one or more embodiments;
[0046] Fig. 9 is a schematic view of an example of a computing device of the controller of the robotic system of Fig. 1 , in accordance with one or more embodiments;
[0047] Fig. 10 is a perspective view of an example of an end effector mounted to an effector end of a robotic arm, in accordance with one or more embodiments;
[0048] Fig. 11 s a close-up perspective view of tool end members of the end effector of Fig.10, in accordance with one or more embodiments;
[0049] Fig. 12 is another close-up perspective view of the tool end members of Fig. 11 , in accordance with one or more embodiments;
[0050] Fig. 13 is a perspective view of a tool end member of Figs. 11-12, in accordance with one or more embodiments;
[0051] Fig. 14 is a front elevation view of the tool end member of Fig. 13, in accordance with one or more embodiments;
[0052] Fig. 15 is a side elevation view of the tool end member of Fig. 13, in accordance with one or more embodiments; and
[0053] Figs. 16A, 16B, 16C, 16D, 16E and 16F are frames showing the end effector of Fig.10 executing an operating sequence to reach for an object to be grasped within a pile of cluttered objects in a pick-up area, in accordance with one or more embodiments.DETAILED DESCRIPTION
[0054] Fig. 1 shows an example of a robotic system 100 for grasping object(s) 13 in a pile of randomly placed objects 13, in accordance with an embodiment. The randomly placed objects 13 can be contained in a container or resting on a support surface, depending on the embodiment. Accordingly, the objects 13 in the pile may be randomly positioned, randomly oriented, or a combination of both. In this example, the objects 13 to be manipulated are directed to many occurrences of the same object. More specifically, the objects share a common size, color, material, and the like. The objects 13 may be identical to one another. However, manipulating randomly placed objects that are different from one another can also be envisaged in some other embodiments.
[0055] In the illustrated embodiment, objects 13 of the pile may be moved from a pickup area 102 to a drop-off area 104. As illustrated, the robotic system 100 has a camera 106 facing the pickup area 102. The camera 106 thus has a field of view 108 encompassing at least a portion of the pile of randomly placed objects 13. Images captured by the camera 106 can show at least some randomly placed objects 13 of the pile. As depicted, the robotic system 100 has a controller 110 which is communicatively coupled to the camera 106 to receive the images from the camera 106 and process the images using computing modules which are stored on a non-volatile memory of the controller 106 and executable by a processor of the controller 106.
[0056] A robot 112 is also provided to handle the object 13 and move it from the pickup area 102 to the remote drop off area 104. The robot 112 typically has an end effector (not shown) which is movable between the pickup area 102 and the drop-off area 104. The robot 112 can be provided in the form of a robotized arm in some preferred embodiments, an example of which will be described further below with reference to Figs. 10-16F. It is intendedthat the robotic system 100 is not limited to only one robot or robot arm, as it can have more than one robot or robot arm configured for picking up, in either a simultaneous or sequential manner, objects 13 from the pile of randomly placed objects.
[0057] Fig. 1 A shows a top plan view of the pile of randomly placed objects 13. As depicted, objects 13 that are closer to the camera 106 may appear bigger in size and have a larger area, whereas objects 13 that are farther from the camera 106 may appear smaller in size and have a smaller area. Each of the objects 13 have a respective unknown position and orientation within the pickup area 102. These position and orientation can include for instance six degrees of freedom to fully determine the position and orientation of any given objects in the pile. The six degrees of freedom can include three degrees of freedom determining position (translational movements along axes x, y, and z) and three degrees of freedom determining orientation (rotational movements about perpendicular axes a, p, and y). Fig. 1 B shows a single object of the pile of Fig. 1A on which has been overlaid an annotation showing an unknown position and orientation x1 , y1 , z1 , a1 , p1 , and y1.
[0058] As will be described in fuller detail below, it is a purpose of the robotic system 100 to find the actual position and orientation of a graspable object within the pile of randomly placed objects. In this disclosure, the term “orientation” is used to refer to the rotational positioning of the object regardless of its translational position. In other words, the term “orientation” is meant to refer to the rotational coordinates a, p, and y of a given object 13. The term “position” is used to refer to translational position of the object regardless of its rotational position. As such, the term “position” is meant to refer to the translational coordinates x, y, and z of a given object 13. The term “pose” is used to refer to a combination of the position and of the orientation of a given object. Accordingly, the term “pose” may refer to the translational and rotational coordinates of a given object, e.g., including translational coordinates x, y and z and rotational coordinates a, p, and y. According to this nomenclature, Fig. 1 B thus shows the object 13 at the given pose x1 , y1 , z1 , a1 , p1 , and y1 .
[0059] Referring now to Fig. 2, which shows a schematic view of the system 100, the camera 106 captures a given image 120 of the randomly placed objects 13 of Fig. 1. The robotic system 100 is not limited to only one camera, as it can have more than one camera capturing images from different viewpoints, and / or capturing different types of images whichmay be processed to find a graspable object. For instance, in this specific embodiment, a red-green-blue (RGB) camera 106a and a depth camera 106b are used concurrently to simultaneously capture a RGB image and a depth image, respectively, of some or all of the randomly placed objects in the pile for further processing. For instance, in some embodiments, the camera 106 is a two-dimensional camera capturing two-dimensional images. Examples of such two-dimensional images can include, but are not limited to, color images, RGB images, black and white images, plenoptic images and the like. The camera 106 may also be a three-dimensional camera capturing three-dimensional images such as depth images, and the like. Examples of such three-dimensional images can include, but are not limited to, depth images, stereoscopic images, reconstructed plenoptic images, and the like.
[0060] The controller 110 is communicatively coupled to the camera 106 and to the robot 112. Such communication may be wired or wireless, or a combination of both depending on the embodiment. As shown, the controller 110 has a segmentation module 122, a graspable object finding module 124 and a robot controlling module 126. The segmentation module 122 is configured for segmenting the given image 120 into image segments 130. After the segmentation of the given image 120, the graspable object finding module 126 is configured for comparing the image segments 130 to template images 132 (and / or associated descriptor features) representing a virtual object matching the object 13. Each template image 132 represents the object 13 in a corresponding one of a plurality of different orientations (e.g., different a, p, and y). The template images 132 can be used to generate descriptor features 133 (e.g., dino features, SAM-6D) which can be used in the comparison in addition or in substitution to the template images 132. As best shown in Fig. 3, the template images 132 may have been generated by a template image generation module 140 of the controller 110 using a computer model 142 of the object to be manipulated. The template images 132 may have been generated by another, external controller, in some embodiments. The generation of the template images 132 may be done only once for each object type, and may be stored (e.g., cached) in a computer memory accessible to the controller 110. The descriptor features 133 may be generated directly from the computer model of the object and / or directly from the template images 132. Referring back to Fig. 2, some image segments at least partially showing the object 13 in any of the plurality of different orientations are identified in the comparison and are then processed until a given image segment showing a graspable objectis found. The processing of the given image 120 involves determining whether the corresponding image segments 130 show an exposed grasping area of the object 13 in a desired grasping orientation, and, when the exposed grasping area of an object is found in an image segment, determining a position and orientation 134 of the graspable object. When a graspable object is found, the robot controlling module 128 is configured for instructing the robot 112 to perform a step of grasping, with an end effector and via a grasping instruction 136, the graspable object shown in the given image segment 130 based on the position and orientation 134 of the graspable object. The grasped object may be moved to a remote drop off area as discussed above.
[0061] Fig. 4 shows a flow chart of a method 400 of grasping randomly placed objects using a robot. The method 400 can be performed using the robotic system 100 described with reference to Fig. 1 , or any other suitable robot or robotic system, depending on the embodiment.
[0062] At step 404, the camera captures a given image showing randomly placed objects. The camera can include one or more camera. Accordingly, the given image can include one or more given images which are to be processed to find graspable objects in the pile of randomly placed objects.
[0063] At step 406, using a segmentation module, the given image is segmented into image segments. The segmentation of the given image can involve any existing image segmentation algorithms including, but not limited to, thresholding image segmentation algorithms, edgebased image segmentation algorithms, region-based image segmentation algorithms, clustering-based image segmentation algorithms, neural network-based image segmentation algorithms, and graph-based image segmentation algorithms.
[0064] At step 408, using a graspable object finding module, the image segments are compared to template images. The template images can be produced at step 402 of the method during which the template images are generated using a computer model of the object of interest. This step can be performed by a template image generation module and can take fewer than 10 minutes, preferably fewer than 5 minutes, and most preferably fewer than 4 minutes to compute. The step 408 includes a step of identifying some image segments whichat least partially show the object in any of the different orientations of the template images. In some embodiments, the step 408 of comparing includes comparing the image segments to descriptor features of the object in a plurality of different orientations. The descriptor features may have been generated from a computer model of the object, from the template images, or a combination of both.
[0065] In some embodiments, the image segments showing an object are kept or labeled accordingly whereas the image segments failing to show an object are rejected, discarded from the lot of image segments, cleared from the controller’s memory, and / or labelled accordingly. In some other embodiments, the image segments showing an object are referred to as object segments. Any other suitable way of manipulating, naming, labelling the image segments is meant to be encompassed, as long as it enable the next steps of the method 400 to be applied only to the image segments showing an object or a portion thereof for computational efficiency purposes.
[0066] At step 410, those image segments, i.e., the ones at least partially showing the object, are processed until a given image segment showing a graspable object is found. The step 410 can include a step 412 of determining whether each of the corresponding image segments show an exposed grasping area of the object in a desired grasping orientation, and if any, a step 414 of determining an actual position and orientation of the graspable object.
[0067] At step 416, using a robot controlling module, a robot is instructed to perform a step of grasping, with an end effector for the robot, the graspable object showing in the given image segment based on the position and orientation of the graspable object determined at step 414. After grasping the graspable object in the pile of randomly placed objects, the method 400 can include a step of moving away from the pile to position the grasped object at a remote drop off area. The grasped object may be dropped off at a specific position of the remote drop off area, and in a specific orientation or stance, depending on the embodiment.
[0068] In some embodiments, a subsequent object is to be picked up from the pile of randomly placed objects. In these embodiments, the removal of the first graspable object can move at least some surrounding objects, resulting in a new configuration of the randomly placed objects of the pile. In these embodiments, the method 400 can include a step ofcapturing yet another given image of the randomly placed objects which have been repositioned upon removal of the first graspable object, and a step of repeating the steps 406 to 416 of the method 400, and so forth, until all of the graspable objects in the pile have been satisfactorily handled by the robot.
[0069] In some embodiments, the method 400 has a step of ranking the image segments showing an object ora portion thereof prior to their processing. This ranking step can prioritize image segments which are larger in area, bigger in size and / or closer to the camera as they are statistically more prone of showing exposed grasping areas. The ranking step thereby allows the method 400 to converge and find a graspable object faster than if the method 400 were to be executed without any ranking step. Indeed, image segments that are smaller in area or size, or deeper, may show objects which are at the bottom of the pile, or obstructed, thereby being ungraspable. In contrast, image segments that are larger in area, greater in size, or shallower, may show objects that are at the top of the pile, or unobstructed, thereby being more likely to be graspable. Typically, when the images are two-dimensional images, the image segments can be ranked in a decreasing order of size or area (from bigger to smaller) of the object in the corresponding image segments. Additionally or alternately, the image segments can be ranked in an increasing order of depth (from the shallowest to the deepest) when the images are two-dimensional images. It is intended that this step of ranking does not necessarily involve computer vision and can therefore be performed by the computer at relatively advantageous computational cost. Indeed, it was found that performing the steps 404 to 416 of the method 400 using a satisfactory controller can take fewer than 10 seconds, preferably fewer than 5 seconds, and preferably fewer than 3 seconds, which can indeed enable real-time or quasi-real-time operation of the robotic system 100.
[0070] Fig. 5A shows an example image 520 showing a number of randomly placed objects 13. As depicted, there is no regular arrangement of the objects, as they all have different poses, i.e., different positions and orientations. While all being identical objects, some have a position and an orientation which is more easily graspable by an end effector of a robot. Those more graspable objects should preferably be grasped first, thereby reducing manipulation error risks. Fig. 5B shows the image 520 over which image segments 530 outputted by a segmentation module have been overlaid at the right positions. As shown, some imagesegments do not show the object, such as image segment 530a, whereas some other image segments show the object, such as image segments 530a, 530b, and 530b. All of the image segments 530 are compared to the template images of the object of interest to determine whether each image segment shows at least a portion of the object. Fig. 6A shows that upon comparing the image segment 530a to the template images, no match can be found. Accordingly, it is determined that image segment 530a does not contain an object or a portion thereof, and can therefore be rejected (e.g., removed from the processing pipeline, cleared from the controller memory, labelled accordingly). In contrast, Figs. 6B, 6C and 6D show that upon comparing the image segments 530b, 530c, and 530d, respectively, matches could be found with corresponding ones of the template images. Since the orientation of the objects shown in these image segments differ from one another, the matching template images are also different from one another. As such, the image segments 530b, 530c, and 530d are kept for further processing. It is noted that reducing the number of image segments using this template image-based comparison is cheap in terms of computational costs and can greatly reduce the computational requirements which can in turn allow for real time or quasi-real-time operation of the robotic system 100.
[0071] As discussed above, the remaining image segments may be ranked prior to being further processed. If the illustrated image segments 530b, 530c, and 530d were to be ranked by decreasing area order, the image segment 530b would be further processed first, then the image segment 530c would be processed second, and the image segment 530d would be further processed third, and so forth. Any other type of ranking or prioritizing may be applied to accelerate the rate or odds of swiftly finding a graspable object in the image segments.
[0072] In some embodiments, the processing of an image segment can include a step of determining a coarse orientation of the object based on an orientation of the matching template image. In these embodiments, the processing of the image segment can be based on the coarse orientation of the object shown in the respective image segment. More specifically, as each template image has a corresponding orientation, and because a match is found between an image segment and a corresponding template image, a coarse orientation of the object shown in the image can be inferred from the orientation of the matched template image. The coarse orientation can be used in determining whether a grasping area of the object is likelyto be exposed. For instance, if it is known that the orientation 1 and the orientation 42 of the image segments 530b and 530c, respectively, are prone to exposing the grasping area of the object, then the chances of the image segments 530b and 530c showing an exposed grasping area are more likely. The image segments 530b and 530c may thus be kept for further processing. However, in contrast, if it is known that the orientation 88 of the image segment 530d is not particularly favourable for grasping purposes, the image segment 530d can be rejected as the object it shows may be being ungraspable or not easily graspable.
[0073] Moreover, the processing of the image segment can include a step of determining whether an end effector path leading to the object shown in the corresponding image segment is obstructed, and if so, the given image segment may be rejected or labeled accordingly. An image segment may also be rejected for further processing purposes if it is determined that the grasping area of the object shown in the given image segment is occluded in any way. Referring now to the image segment 530c of Fig. 5B, it is known that the object shown in the image segment 530c has an orientation which is favourable for grasping. However, since the image segment 530c has an area which is smaller than an expected area for an object having that orientation, it can be determined that the object is either too deep in the pile or obstructed by surrounding object(s). In this case, the object shown in Fig. 530c appear to be both too deep and obstructed by surrounding objects.
[0074] In some embodiments, the processing of the image segment can include a step of estimating a position and orientation of the object based on the coarse orientation, generating a three-dimensional model of the object using the position and orientation, rendering a test image showing the object at the estimated position and orientation, and comparing the object of the test image to the object of the given image segment. Such embodiments may be performed by a pose estimating module and a test image rendering module of a controller, an example of which is shown in Fig. 7. For instance, Fig. 8A shows a test image showing an object at an estimated position and orientation. It is clear that upon comparing the image segment of the image 520 to the test image that a match can’t be found therebetween. In these circumstances, the estimated position and orientation of the object may be rejected and estimated again, or focus can change to another image segment which may be more easily analyzable. Indeed, in the case of Fig. 8B, a test image showing an object at an estimatedposition and orientation is overlaid to the corresponding image segment. As depicted, a match is found between the two. In these circumstances, the position and orientation of the object shown in the corresponding image segment is known with a satisfactory degree of certainty, which can allow the determination of whether a grasping area thereof is exposed for grasping purposes.
[0075] The following paragraphs set forth implementations in which a test image is generated and compared to the representation of the object within the corresponding image segment. This comparison serves as a verification mechanism to ensure that a grasping operation is carried out only when the computed position and orientation of the object exhibit a level of accuracy sufficient to enable reliable robotic manipulation.
[0076] In certain embodiments, the processing performed by the computing device further includes generating a test image representing the object as it would appear when positioned and oriented according to the parameters previously determined forthat object. The rendering of this test image allows the computing device to create a visual or computational model of the expected appearance of the object in the scene. Once rendered, the computing device evaluates the relationship between the test image and the original image segment corresponding to the object, thereby enabling a comparison that serves to validate or refine the estimated position and orientation.
[0077] In some implementations, the act of rendering the test image includes graphically displaying the test image directly overthe corresponding image segment. The superimposition of the test rendering onto the selected region of the acquired image facilitates a more precise assessment of the match between the predicted object representation and the actual visual data. This graphical overlay may be generated using conventional image processing techniques or more advanced rendering methods capable of accurately modeling object contours, edges, or shading. In some other embodiments, instead of visually displaying the test image, the computing device may mathematically superimpose the test image over the image segment in memory and evaluate pixel-wise differences, similarity scores, or contour alignment. Alternatively or additionally, the test image may be converted into a binary mask or segmentation mask, which is then compared to a mask derived from the image segment usinglogical AND / OR operations to compute overlap. In these last instances, the graphical display can be omitted.
[0078] In certain embodiments, determining the overlap between the rendered test image and the selected image segment includes performing an intersection-over-union (loU) calculation. The loU metric provides a quantitative measure of the extent to which the rendered test image and the corresponding portion of the original image occupy the same area within the scene. By computing the ratio of the intersecting region to the union of both regions, the computing device obtains a normalized value indicative of the accuracy of the predicted object localization.
[0079] Based on the determined overlap value, the computing device may further take an automated decision with respect to the suitability of the selected image segment. When the overlap value falls below a predetermined threshold, the computing device may classify the selected segment as unsuitable and reject it accordingly. Conversely, when the overlap value exceeds the threshold, the computing device may confirm the previously determined orientation as the correct or desired orientation for a grasping operation. This conditional evaluation ensures that the robot is instructed to grasp objects only when a sufficiently reliable match has been detected between expected and actual object positioning.
[0080] In further embodiments, the overlap threshold used for determining suitability may be set to a high-confidence level to ensure robust operation. For example, the threshold may be at least 65%, and in preferred embodiments at least 75%, while in more stringent implementations the threshold may be set at 85% or higher. Such high-precision thresholds promote accurate pose validation and reduce the likelihood of failed or incorrect grasping attempts by the robotic manipulator.
[0081] In another aspect, this disclosure also presents automated object-handling systems and more specifically to a method enabling a robotic manipulator to grasp objects that are distributed in a random or unstructured arrangement. The method disclosed herein can leverage image processing techniques performed on a digital image of a scene containing multiple objects, enabling precise localization and orientation estimation prior to executing a grasp.
[0082] In one implementation, a camera is positioned such that its field of view encompasses a plurality of objects arranged in a non-organized or scattered manner. The camera acquires an image of this scene. The image is transmitted to a computing device configured to perform segmentation operations. Through segmentation, the computing device divides the acquired image into multiple discrete image segments. Each image segment may contain visual content corresponding to a full object, a portion of an object, or background.
[0083] The computing device then executes a grasping routine that operates on the segmented portions of the image. During this grasping routine, the computing device selects a given one of the image segments for evaluation. The given image segment is analyzed to determine whether it depicts an object suitable for grasping. This may involve pattern recognition, classification algorithms, or feature extraction steps sufficient to distinguish an object from the surrounding environment.
[0084] Once an image segment showing an object is identified, the computing device determines the spatial characteristics of the object within that image segment. These characteristics include at least the estimated position and orientation of the object relative to the frame of the acquired image. Using this position and orientation information, the computing device generates or renders a test image representing the object as it should appear when correctly aligned with the determined spatial parameters. This test image is then computationally overlaid or otherwise compared to the actual selected segment through an overlap evaluation.
[0085] The overlap evaluation assesses the degree of similarity or correspondence between the test image and the selected image segment. The computing device determines an overlap value indicative of how well the rendered test image matches the object representation in the actual image segment. When this overlap value exceeds a predetermined threshold, the computing device considers the identified object and its estimated parameters sufficiently accurate for manipulation. In response, the computing device transmits control instructions to the robot, enabling the robot to perform a grasping step targeting the object based on the determined position and orientation.
[0086] If, however, the computed overlap does not exceed the threshold, the grasping routine does not proceed with grasping forthat image segment. Instead, the computing device continues the grasping routine by selecting another of the previously generated image segments and repeating the routine described above.
[0087] This iterative approach ensures that the robot attempts to grasp only those objects for which a confidence threshold, defined by the overlap measurement, has been satisfied. For instance, this iterative approach ensures that situations where the image has been oversegmented or under-segmented are identified and rejected prior to the actual grasping operation, thus saving time and enhancing grasping efficiency. Over-segmentation occurs when an image is segmented into too many segments, whereas under-segmentation occurs when an image is segmented in too few segments. Both these segmentation error types can adversely affect the grasping operation if not identified in time. Through these operations, the method disclosed herein can enable robust grasping of individual objects from a collection of randomly distributed items, even in settings where precise object placement or uniform orientation cannot be guaranteed.
[0088] In this disclosure, the term “render” is meant to encompass any generation of a visual, mathematical, and / or computational representation of an object, scene, or model based on underlying data, such as estimated position, orientation, geometry, appearance parameters, 2D or 3D models, and the like.
[0089] In this disclosure, the term “grasping area” refers to an area of the object which can be grasped by an end effector of the robot. The size, shape or form of the object can dictate the size, shape or form of its grasping area. An object can have one or more grasping areas. Additionally or alternately, the type of end effector can also have an influence on the size, shape or form of the grasping area(s). For instance, in the example illustrated above, the object has a generally cylindrical shape. Accordingly, the grasping area may be a cylindrical surface located about a middle section of the object, provided that the end effector is sized and shaped to grasp a cylindrical surface. Moreover, another grasping area of the cylindrical object may involve its hollow bore inside which a finger element of the end effector can be inserted for grasping purposes. It is understood that since the grasping of an object is objectdependent, and also optionally end effector-dependent, the robotic system and methoddisclosed herein can factor in the form, size and shape of the object to be manipulated and / or the form, size and shape of the end effector which is meant to grasp the object of interest. Accordingly, the step of determining whether the corresponding image segments show an exposed grasping area can factor in object data including information relative to the known potential grasping areas of the object of interest and / or end effector data including information relative to the known size, shape, form and specific grasping motion of the end effector. However, as discussed above, it may not be sufficient for a grasping area to be exposed in the image segment, the grasping area may preferably also have a grasping orientation which would allow the end effector to satisfactorily grasp the object as it lies in the pile of randomly placed objects.
[0090] Referring now to Fig. 9, the controller can be provided as a combination of hardware and software components. The hardware components can be implemented in the form of a computing device 900, an example of which is described with reference to Fig. 9. The computing device 900 can have a processor 902, a memory 904, and I / O interface 906. Instructions 908 for grasping randomly placed objects can be stored on the memory 904 and accessible by the processor 902.
[0091] The processor 902 can be, for example, a general-purpose microprocessor or microcontroller, a digital signal processing (DSP) processor, an integrated circuit, a field-programmable gate array (FPGA), a reconfigurable processor, a programmable read-only memory (PROM), a programmable logic controller (PLC), or any combination thereof.
[0092] The memory 904 can include a suitable combination of any type of computer-readable memory that is located either internally or externally such as, for example, randomaccess memory (RAM), read-only memory (ROM), compact disc read-only memory (CDROM), electro-optical memory, magneto-optical memory, erasable programmable readonly memory (EPROM), and electrically-erasable programmable read-only memory (EEPROM), Ferroelectric RAM (FRAM) or the like.
[0093] Each I / O interface 906 enables the computing device 900 to interconnect with one or more input devices, such as mouse(s), keyboard(s), sensor(s), camera(s), or with one ormore output devices such as monitor screen(s), robot(s), robot controller(s), accessible memory system(s), and internal or external networks).
[0094] Each I / O interface 906 enables the controller to communicate with other components, to exchange data with other components, to access and connect to network resources, to server applications, and perform other computing applications by connecting to a network (or multiple networks) capable of carrying data including the Internet, Ethernet, plain old telephone service (POTS) line, public switch telephone network (PSTN), integrated services digital network (ISDN), digital subscriber line (DSL), coaxial cable, fibre optics, satellite, mobile, wireless (e.g., Wi-Fi, WiMAX), SS7 signaling network, fixed line, local area network, wide area network, and others, including any combination of these.
[0095] The computing device 900 and any software application that can be run by the computing device 900 are meant to be examples only. Other suitable embodiments of the controller can also be provided, as it will be apparent to the skilled reader.
[0096] The following paragraphs show an example of a robot grasping an object using an end effector. In this case, the end effector is specifically suited to the shape and size of the object of interest. However, in some other embodiments, the end effector may be a conventional gripper, for instance. Fig. 10 shows an example of an end effector 10 mounted at an effector end 11 of a robotic arm 12. The robotic arm 12 may have joints 12A and links 12B which, when articulated, can perform a displacement of the effector end 11 with several degrees of freedom in rotation and in translation. The robotic arm 12 may be a collaborative robot, also known as cobot. The robotic arm 12 may also be fully automated, or be operable in both a collaborative mode and an automated mode. Such type of robotic arm 12 may be easily programmable and may interact in an environment shared with humans relatively reliably and safely.
[0097] In an exemplary use scenario as illustrated in Fig. 10 and further illustrated and described with reference to Figs. 16A-16F herein below, grasping a selected object 13 from a pile of objects 13 in a receptacle or on a surface may involve pushing aside other nearby objects to declutter a surrounding space in the vicinity of the selected object 13 to be grasped, which may result in the shift of the selected object 13 with respect to the end effector 10 asthe end effector 10 approaches the selected object 13. As described herein, some operating sequences to displace the end effector 10 towards the selected object 13 to be grasped and to actuate the end effector 10 to open and close grasping members of the end effector 10 may be better adapted to limit collisions with nearby objects and allow repeatable and precise grasping of objects in the presence of other nearby objects within a pick up area 14.
[0098] According to the present disclosure, in an application, the end effector 10 is adapted to grasp an object 13 that is a flanged component. The flanged component can be a cylinder, elliptical cylinder, or convex prism (as some possibilities), with a portion of the selected object 13, along an outer surface thereof, having a positive relief (e.g., protrusion, ridge, flange, or other types of projections). An example of such a component can be seen in Fig. 10 and Figs.16A-16F. In the example shown, the selected object 13 has, at least, a flange 13A extending along the outer surface of the selected object 13, about a central axis 13X (Fig. 16D) of the selected object 13. Components having this or a similar outer geometry can be pipe / hose connectors / fittings or other similar hardware components. The flange 13A may define predetermined contact points or a predetermined surface area to be targeted by the end effector 10 (or automation system 1 including the end effector 10, robotic arm 12 and related controller) as a portion of the object 13 identified as the preferred grasping zone. The grasping zone on the object 13 can be determined based on several criteria, for example a geometry, a surface texture, a surface relief / structure, or other surface features identified (or preidentified) by an operator, a vision device (not shown) and a controller 40 of the automation system 1 , specifications (e.g., dimension, shape, etc.) or a virtual model (e.g., CAD model) of the object 13, for example. The component shown is only provided as an exemplary component that will be used as a reference throughout the present disclosure, to present an exemplary use scenario of the end effector 10. Similar use scenarios for the end effector 10, with other components, or different predetermined grasping zones or structures could be contemplated.
[0099] The design of the end effector 10, and more specifically the tool end members 30 (described later) configured to contact the object 13 to grasp may allow to compensate for misalignment between the selected object 13 to grasp and the end effector 10. Such misalignment may occur, in pick up scenarios where the selected object 13 is surrounded byother nearby objects and, by approaching the selected object 13 to be grasped, the tool end members 30 inserted into a pile of cluttered objects push nearby objects in the vicinity of the selected object 13, which may cause the selected object 13 to move (position and / or orientation changes) as it is being grasped.
[0100] Referring to Fig. 10, an exemplary end effector 10 and tool end members 30 will now be described. The end effector 10 is mountable to the effector end 11 . One or more actuated joints 12A at the mounting interface between the end effector 10 and the effector end 11 may allow multiple degrees of freedom movement of the end effector 10 (e.g., rotational and / or translational degrees of freedom). The end effector 10 is configured to perform prehensile movements to grasp an object as an exemplary object 13 shown in Fig. 10. The end effector 10 has an actuation mechanism 20. The actuation mechanism 20 may include a combination of links, bars and joints to define articulations configured for permitting a controlled motion of movable ends 21 of the end effector 10. In an embodiment, the actuation mechanism 20 has a four-bar linkage configuration. Other configurations could be contemplated. The actuation mechanism 20 is operable to move the movable ends 21 one with respect to another, between a closed position (Fig. 16A) and a release or open position (Fig. 16F). A displacement of the movable ends 21 from the closed position to the open position (or vice versa), may follow a linear or arcuate trajectory, as some possibilities. The end effector 10 allows for an intermediary position, which may be referred to as a grasping position, between the closed position and the open position.
[0101] Tool end members 30 extend at respective movable ends 21 of the actuation mechanism 20. In the closed position (Fig. 16A), the tool end members 30 are at a minimum distance from each other that is smaller than a widthwise dimension of the selected object 13. In the closed position, the end effector 10, at the movable ends 21 , is configured to define a relatively narrow end, such as an arrow-like end to facilitate an insertion of the tool end members 30 in gaps / spaces between adjacent components in the pick up area 14. The tool end members 30 could contact each other in the closed position, in variants. In the open position (Fig. 16F), the minimum distance between the tool end members 30 is greater than a widthwise dimension of the selected object 13, allowing the selected object 13 to be inserted in a space, which may be referred to as a grasping area, between the tool end members 30.Preferably, the grasping area is exposed to the end effector for a grasping operation to be possible. The tool end members 30 are adapted to engage the selected object 13 in the grasping position. In the grasping position, the tool end members 30 may engage opposite sides of the selected object 13 to hold it, firmly enough to prevent any shift of the selected object 13 in a transverse direction with respect to an engagement direction of the tool end members 30 onto the selected object 13. Load sensors of the end effector 10 and / or robotic arm 12 may provide feedback information to a controller of the robotic arm 12 and / or end effector 10 on a grasping load applied by the end effector 10 onto the selected object 13 during grasping, to monitor and / or control the grasping load onto the selected object 13. In the open position, the selected object 13 may be inserted in a space between the movable ends 21 I tool members 30, which will now be described according to some embodiments.
[0102] Referring to Figs. 11-12, the tool end members 30 include a first tool end member 31 and a second tool end member 32 facing the first tool end member 31 . The first tool end member 31 and the second tool end member 32 are configured to engage with the selected object 13 (Fig. 10) to be grasped. The first tool end member 31 and second tool end member 32 may be identical / mirrored versions of each other, though this is optional.
[0103] Referring to Figs. 11-12, the first tool end member 31 and the second tool end member 32 have a connector interface 33 adapted to connect with respective ones of the movable ends 21 of the actuation mechanism 20. In an embodiment, the connector interface 33 and the movable ends 21 define a male-female connection between a base of the first tool end member 31 and the second tool end member 32 and respective movable ends 21 , for engaging each one of the first tool end member 31 and the second tool end member 32 with a respective one of the movable ends 21. Referring to Fig. 13, the connector interface 33 includes a recess forming the female part of the male-female connection. The connector interface 33 could form the male part of the male-female connection in variants. Other connector interface 33 could be contemplated (e.g., other shapes, connection types).
[0104] Now referring to Figs. 13-15, the first tool end member 31 and the second tool end member 32 have respective inner sides 34 configured to face towards one another. The tool end members 31 ,32 have an outer side 35 opposite the inner side 34. The outer side 35 may be referred to as the back end side of the tool end members 31 , 32, and the inner side 34 maybe referred to as the object contacting side of the tool end members 31 , 32. The outer side 35 defines a rear surface of the tool end members 31 , 32. As can be seen in Fig. 12, the rear surfaces of the first tool end member 31 and the second tool end member 32 taper towards the tip of the tool end members 31 , 32. Such tapering may facilitate an insertion of the tool end members 31 , 32 on either sides of the selected object 13 to be grasped, in the presence of other nearby objects in the cluttered pile of objects 13 where the pick up of the selected object 13 may occur. The rear surface may contact, and push aside nearby objects as the tool end members 31 , 32 approach the selected object 13 to be grasped, as will be further described later.
[0105] The inner side 34 of each one of the first tool end member 31 and the second tool end member 32 may be configured to facilitate the engagement with a predetermined grasping zone on the selected object 13. The inner side 34 may be dimensioned, sized and / or shaped so as to receive a portion of the selected object 13 that has been determined, either by an operator or by computer processing, to be a preferred or possible grasping zone on the selected object 13 to provide a reliable and repeatable grasping of multiple objects 13 as part of repeated operation sequences of the robotic arm 12. In at least some embodiments, as shown, the inner side 34 of each one of the first tool end member 31 and the second tool end member 32 has alignment surfaces 36 defining a funnel to guide an engagement of the first tool end member 31 and the second tool end member 32 with a portion of the selected object 13. In at least some embodiments, the alignment surfaces 36 may be arcuate. As shown, the alignment surfaces 36 are concave. The alignment surfaces 36 may generally face in opposite directions. The alignment surfaces 36 each have a surface vector (normal vector projecting from the surface 36) extending in transverse directions one with respect to the other. The surface vector of the alignment surfaces 36 may project towards a median plane MP of the tool end members 31 , 32. The median plane MP may correspond to a symmetrical plane of the space 37S (Fig. 13, described below) or between the finger portions 37 (described below). The surface vector of the alignment surfaces 36 may project in directions intersecting the median plane MP and / or intersecting each other at a location between the tool end members 31 , 32, in the closed position and / or the grasping position. In some embodiments, the alignment surfaces 36 may be flat surfaces inclined so as to generally face inwardly with respect to each lateral sides of the tool end members 31 , 32. The surface vector projectingfrom the alignment surfaces 36 may point in opposite directions, at a relative angle from each other a side each other.
[0106] As shown, finger portions 37 extend from the base to a tip 39 of the tool end member 31 , 32. The finger portions 37 include respective ones of the alignment surfaces 36. The finger portions 37 are laterally spaced from one another to define a space 37S (Fig. 13) configured to receive a portion, such as the flange 13A, of the object 13, when the end effector 10 is in the grasping position. The finger portions 37 defines spaced apart ridge surfaces 37R at the inner side 34 of the tool end members 31 , 32. In at least some use scenarios, the ridge surfaces 37R may contact the outer surface of the object 13, in the grasping position. In at least some embodiments, the profile of the ridge surfaces 37R correspond to the profile of the object 13, at least along part of the outer surface thereof. The profile of the ridge surfaces 37R could be a negative of the object 13 to be grasped, though this is optional. In the embodiment shown, the ridge surfaces have a concave profile adapted to follow a profile of the outer surface of the generally cylindrical object 13. In some embodiments, as shown, the ridge surfaces 37R may be recessed with respect to an innermost surface of the tool end members 31 , 32 (Fig. 15).
[0107] The first tool end member 31 and the second tool end member 32 each define a groove 38 on the inner sides 34 of the first tool end member 31 and the second tool end member 32. The alignment surfaces 36 extend along sides of the groove 38. The alignment surfaces 36 are joined by a groove end 38E. As shown, an end surface 38S defines the groove end 38E. The end surface 38S extends between the alignment surfaces 36. In at least some use scenarios, the end surface 38S may contact a portion of the object 13, in the grasping position. Grasping load onto the object 13 may be transmitted to the object 13 by the end surface 38S of the groove end 38E of respective ones of the tool end members 31 , 32, in at least some cases. The end surface 38S of the groove 38 and the alignment surfaces 36 may be segments of a continuous surface delimiting the end and sides of the groove 38, in at least some embodiments. The surfaces 36, 38S may be distinct surfaces joined together at connecting edges between adjacent ones of such surfaces, as another possibility. In some variants, the finger portions 37 could be separate from each other, so as to not being joined together by a wall or surface extending between the alignment surfaces 36. The finger portions37 may be structurally connected to each other only via a common base, and projecting from that common base independently from each other. While finger portions 37 not joined together by a common wall between them along a substantial portion of their length could be contemplated, interconnection between the finger portions 37 may provide more sturdiness to the tool end members 31 , 32.
[0108] The tool end members 31 , 32 have a tip 39, at a free end thereof. Each finger portion 37 may define parts of the tip 39. In the embodiment shown, the tip 39 defines a notch 39N between the alignment surfaces 36. The groove end 38E may be delimited in part by the notch 39N. The notch 39N has a depth extending from the tip 39 towards the base of the finger portions 37. The notch 39N may be sized and / or shaped to receive a portion of the object 13 - the portion of the object 13 forming part of or corresponding to the predetermined grasping zone on the object 13 - during a grasping operation. The notch 39N may serve as a guard to limit transverse shifts of the object 13 with respect to the tip 39 of the tool end members 31 , 32 during a grasping operation, and / or facilitate the maintaining of the alignment of the tip 39 of the tool end members 31 , 32 with the portion of the object 13 forming part of or corresponding to the predetermined grasping zone on the object 13, as will now be described with reference to Figs. 16A-16F.
[0109] In at least some embodiments, a controller 40 (Fig. 10) which may be separate or the same controller as that of the robotic arm 12, can include one or more processing units and transitory computer-readable memory communicatively coupled to the processing unit(s) and comprising computer-readable program instructions executable by the processing unit(s) for operating the end effector 10 and / or the robotic arm 12 in the manner described herein.
[0110] The controller 40 can be programmed with a set of instructions which, when executed by one or more processing unit(s) of the controller 40, can cause the following sequence of operation : displace the end effector 10 towards the selected object 13 with a first tool end member 31 and a second tool end member 32 of the end effector 10 in a closed position, trace a surface profile of the selected object 13 with a tip 39 of respective ones of the first tool end member 31 and the second tool end member 32 as the first tool end member 31 and the second tool end member 32 move away from each other towards an open position until the first tool end member 31 and the second tool end member 32 are placed on oppositesides of the selected object 13, and grasp the selected object 13 between the first tool end member 31 and the second tool end member 32 by moving the first tool end member 31 and the second tool end member 32 in a gripping position between the closed position and the open position.
[0111] Fig. 16A shows the end effector 10 with the tool end members 31 , 32 generally aligned with a flange 13A of the selected object 13 within the pick up area 14, here shown as a bin containing a plurality of objects that are similar or identical to the selected object 13. The tool end members 31 , 32 are in the closed position, thereby defining a narrow end of the end effector 10 to facilitate their insertion between closely positioned objects in the vicinity of the selected object 13 targeted for the grasping. Reaching towards the selected object 13 with the first tool end member 31 and the second tool end member 32 in the close position may facilitate the approach towards the selected object 13, when the object 13 is in a cluttered space or pile of objects, and allow the tips 39 of the first tool end member 31 and the second tool end member 32 to get in closer proximity with the selected object 13 while limiting tip collisions of one or both of the tool end members 31 , 32 with other nearby objects. Colliding with nearby objects could prevent one or all tool end members 31 , 32 from making contact with the selected object 13. The collision between the tool end member(s) 31 ,32 and nearby objects could cause the robotic arm 12 to go into a safety state and abort the grasping task.
[0112] Once the tool end members 31 , 32 have reached a position as shown in Fig. 16A where the tip 39 of the respective tool end members 31 , 32 are in close proximity with the selected object 13, the actuation mechanism 20 may be operated to move the tool end members 31 , 32 away from each other. As the tool end members 31 , 32 are being distanced from each other towards the open position, the end effector 10 is operated to maintain the tip 39 of the respective tool end members 31 , 32 at a constant and close distance from the outer surface of the selected object 13. The displacement of the tip 39 of the respective tool end members 31 , 32 may therefore trace a profile of the selected object 13 while the respective tool end members 31 , 32 gain a position where the selected object 13 is situated between the tool end members 31 , 32. Opening the end effector 10 by tracing the profile of the selected object 13 with the tip 39 of the first tool end member 31 and the second tool end member 32 as they move on opposite sides of the selected object 13 may allow to push aside nearbyobjects and insert the first tool end member 31 and the second tool end member 32 in between objects (e.g., sliding the back of the tool end members 31 , 32 against nearby objects) to position the first tool end member 31 and the second tool end member 32 about the selected object 13. Displacing the end effector 10 towards the selected object 13 may therefore include pushing aside nearby objects in the vicinity of the selected object 13 with a rear surface of at least one of the first tool end member 31 and the second tool end member 32 to create an insertion space 15 about the selected object 13 for the first tool end member 31 and the second tool end member 32 to at least partially surround the selected object 13 with the first tool end member 31 and the second tool end member 32. This is illustrated in the frames of the operation sequence of Figs. 16B-16C.
[0113] Displacing the end effector 10 towards the selected object 13 may include aligning the tip 39 of respective ones of the first tool end member 31 and the second tool end member 32 with a positive relief at an outer surface of the selected object 13. The positive relief may be a flange 13A (Fig. 16B) extending along the outer surface of the selected object 13, about a central axis 13X (Fig. 16D) of the selected object 13. Such positive relief is provided as an example, as one or more other portions of the object 13 to be grasped could form part of the grasping zone identified. Such identified grasping zone may vary depending on the component shapes, dimensions, sizes, for example.
[0114] Tracing the surface profile of the selected object 13 with the tip 39 of respective ones of the first tool end member 31 and the second tool end member 32 may include engaging a notch 39N (Figs. 14-15) at the tip 39 of respective ones of the first tool end member 31 and the second tool end member 32 with the flange 13A at the outer surface of the selected object 13, and displacing the tool end members 31 , 32 so that the tip 39 of respective ones of the first tool end member 31 and the second tool end member 32 follows the profile of the outer surface of the selected object 13 so as to maintain the flange 13A within the notch 39N as the first tool end member 31 and the second tool end member 32 close up onto the selected object 13.
[0115] Figs. 16C-16D show frames of the grasping operation sequence where the tool end members 31 , 32 gain the grasping position to hold the selected object 13 between them. Grasping the selected object 13 between the first tool end member 31 and the second toolend member 32 may include engaging the groove 38 (Figs. 14-15) of respective ones of the first tool end member 31 and the second tool end member 32 with the flange 13A of the selected object 13. Engaging the groove 38 with the flange 13A may include aligning the flange 13A between alignment surfaces 36, which may be referred to as tunneled surfaces, of at least one of the first tool end member 31 and the second tool end member 32, the tunneled surfaces 36 delimiting opposite sides of the groove 38 from the tip 39 to a base of the at least one of the first tool end member 31 and the second tool end member 32. Grasping the selected object 13 between the first tool end member 31 and the second tool end member 32 may include capturing the flange 13A of the selected object 13 between finger portions 37 of the first tool end member 31 and the second tool end member 32, the respective finger portions 37 projecting from a base of a respective tool end member 31 , 32. In an embodiment, as described above the finger portions 37 may be joined together along a substantial portion of their length, with the groove 38 defined between the finger portions 37 from the base to the tip 39 of the tool end members 31 , 32. As the tool end members 31 , 32 are being positioned on either sides of the selected object 13 prior to grasping, other nearby objects in the vicinity of the selected object 13 could be contacted by the tool end members 31 , 32. In some cases, this may cause a dynamic where the objects may move and / or cause movement of the selected object 13 during grasping (or immediately prior to gaining the grasping position). The grasping of the selected object 13 between the first tool end member 31 and the second tool end member 32 may cause a change in a position and / or orientation of the selected object 13 with respect to the first tool end member 31 and the second tool end member 32. For example grasping the selected object 13 between the first tool end member 31 and the second tool end member 32 may thus include causing a change in the orientation of the selected object 13 (twist) and / or the selected object 13 may get pushed down, slightly, by the first tool end member 31 and / or second tool end member 32 as the tip 39 of the tool end members 31 , 32 brush up against the selected object 13 while moving down for being positioned on either sides of the selected object 13 priorto grasping. In the first case (i.e., twist), the notch 39N may limit the amount of twist, and the alignment surfaces 36, may help to realign the selected object 13 in the desired grasping orientation with respect to the tool end members 31 , 32 during the grasping. In the second case (i.e., pushed down), the curvature of the finger portions 37 and ridge surfaces 37R, as seen in Fig. 15, may assist in scooping up the selected object 13.
[0116] Figs. 16D-16E show the tool end members 31 , 32 having completed the grasping of the selected object 13 and the end effector 10 transporting the grasped selected object 13. Fig. 16F shows the tool end members 31 , 32 in the open position, releasing the selected object 13 into a drop-off area. The operation sequence described herein may be repeated on and on, by selecting another object within the pick up area 14 (Fig. 16A).
[0117] Selecting and picking up objects within a cluttered pile of objects in the pick up area 14 require dynamic adjustment of the sequence of operation of the robotic arm 12 and end effector 10. The system 1 may be adapted to map the object pick up area 14 to detect the presence, position and orientation of at least one object of a plurality of objects to be grasped, select one object from the plurality of objects based on computed selection criteria, and generate a robot control sequence which, upon execution, allow to place the end effector in a suitable position to align with a grasping portion I grasping points on the selected object 13 and reach for the selected object 13 within the plurality of objects to grasp the selected object 13. Since the position and orientation of each object within the pick up area 14 may be different, the system 1 may require to dynamically adjust the operation sequence to reach towards an object 13 that is identified as graspable from one operation sequence to the other.
[0118] As can be understood, the examples described above and illustrated are intended to be exemplary only. The scope is indicated by the appended claims.
Claims
WHAT IS CLAIMED IS:1 . A method of grasping randomly placed objects using a robot, the method comprising:using a camera facing a plurality of randomly placed objects, capturing a given image of a plurality of randomly placed objects;using a computing device,segmenting the given image into a plurality of image segments, and comparing the image segments to a plurality of template images representing an object matching the randomly placed objects, each template image showing the object in a corresponding one of a plurality of different orientations, said comparing including identifying some image segments at least partially showing the object in any of the plurality of different orientations;processing the some image segments until a given image segment showing a graspable object is found, said processing including determining whether the corresponding image segments show an exposed grasping area of the object in a desired grasping orientation, and determining a position and orientation of the graspable object from the given image segment; andinstructing the robot to perform a step of grasping the graspable object shown in the given image segment based on the position and orientation of the graspable object.
2. The method of claim 1 further comprising, after said grasping, an end effector of the robot moving away from the plurality of randomly placed objects to position the graspable object at a remote drop off location.
3. The method of claim 1 or 2 further comprising the camera capturing another given image of the randomly placed objects which have been repositioned upon removal of thegraspable object, and performing another iteration of said method to grasp another graspable object.
4. The method of claim 3 wherein said capturing the other given image is performed immediately after removal of the graspable object and prior to positioning of the graspable object at a remote drop off location.
5. The method of any one of claims 1 to 4 wherein the given image includes at least one of a two-dimensional image and a depth image.
6. The method of any one of claims 1 to 5 wherein the given image is a two-dimensional image, said processing being performed in a decreasing order of size of the object in the corresponding image segments.
7. The method of any one of claims 1 to 6 wherein the given image is a depth image, said processing being performed in an increasing order of depth of the object in the corresponding image segments.
8. The method of any one of claims 1 to 7 wherein said processing further includes: comparing the object shown in a given image segment to the template images until a matching template image is found, determining a coarse orientation of the object based on an orientation of the matching template image, wherein said processing is further based on the coarse orientation of the object.
9. The method of claim 8 further comprising estimating a position and orientation of the object based on the given image and the coarse orientation, generating a three-dimensional model of the object using the position and orientation, rendering a test image showing the object at the estimated position and orientation, and comparing the object of the test image to the object of the given image segment.
10. The method of claim 9 wherein said processing includes rejecting the given image segment when the object of the test image and the object of the given image segment do not match to one another.
11. The method of any one of claims 1 to 10 wherein said processing further includes rendering a test image showing the object in said position and orientation, and determining an overlap between the test image and the given image segment.
12. The method of claim 11 wherein said rendering includes graphically displaying the test image over the given image segment.
13. The method of claim 11 wherein said determining the overlap includes an intersection over union determination between the test image as rendered and the given image segment.
14. The method of claim 11 further comprising one of: rejecting the given image segment when the overlap is below an overlap threshold, and confirming the orientation as the desired grasping orientation when the overlap is above the overlap threshold.
15. The method of claim 14 wherein the overlap threshold is at least 65%, preferably at least 75%, and most preferably at least 85%.
16. The method of any one of claims 1 to 15 wherein said processing includes rejecting a given image segment upon determining that an end effector path leading to the object amidst the plurality of randomly placed objects is obstructed.
17. The method of any one of claims 1 to 16 wherein said processing includes rejecting a given image segment upon determining that a grasping area of the object in the given image segment is occluded.
18. A robotic system for grasping randomly placed objects, the robotic system comprising:a camera facing a pick-up area encompassing a plurality of randomly placed objects, the camera capturing a given image of a plurality of randomly placed objects;a robot having an end effector movable within the pick-up area and configured for performing a grasping motion at a given position and orientation of the pickup area; anda controller communicatively coupled to the camera and to the robot, the controller:a segmentation module configured for segmenting the given image into a plurality of image segments, and comparing the image segments to a plurality of template images representing an object matching the randomly placed objects, each template image showing the object in a corresponding one of a plurality of different orientations, said comparing including identifying some image segments at least partially showing the object in any of the plurality of different orientations;a graspable object finding module configured for processing the some image segments until a given image segment showing a graspable object is found, said processing including determining whether the corresponding image segments show an exposed grasping area of the object in a desired grasping orientation, and determining a position and orientation of the graspable object from the given image segment; anda robot controlling module configured for instructing the robot to perform a step of grasping, with the end effector, the graspable object shown in the given image segment based on the position and orientation of the graspable object;wherein the robot grasps the graspable object at the position and orientation thereof in response to said instructing.
19. The robotic system of claim 18 wherein said processing further includes: comparing the object shown in a given image segment to the template images until a matchingtemplate image is found, determining a coarse orientation of the object based on an orientation of the matching template image, wherein said processing is further based on the coarse orientation of the object.
20. The robotic system of claim 19 further comprising estimating a position and orientation of the object based on the coarse orientation, generating a three-dimensional model of the object using the position and orientation, rendering a test image showing the object at the estimated position and orientation, and comparing the object of the test image to the object of the given image segment.
21. A method of grasping randomly placed objects using a robot, the method comprising:using a camera facing a plurality of randomly placed objects, capturing a given image of a plurality of randomly placed objects;using a computing device,segmenting the given image into a plurality of image segments; using the plurality of image segments, performing a grasping routine including:identifying a given image segment showing the object; determining a position and orientation of the object in the given image segment;rendering a test image of the object in said position and orientation over the given image segment;determining an overlap between the test image and the given image segment; andupon determining that the overlap exceeds an overlap threshold, instructing the robot to perform a step of grasping the object shownin the given image segment based on the position and orientation of the object.
22. The method of claim 21 further comprising performing another iteration of the grasping routine for another one of the plurality of image segments when the overlap is below the overlap threshold.