Random bin picking and precise placement with a multiview system

The multiview system with 2D cameras and AI-based pose estimation addresses the challenge of precise placement in random bin picking by improving accuracy and reducing costs through triangulation-based refinement, enhancing industrial efficiency.

WO2025177020A1PCT designated stage Publication Date: 2025-08-28ABB (SCHWEIZ) AG
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2024/051619
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-20
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Current random bin picking systems face challenges in achieving precise placement of industrial parts due to inaccurate pose estimation, leading to inefficient cycle times and high costs associated with 3D vision systems, while low-cost RGB or grayscale cameras suffer from poor performance.

Method used

A multiview system using two 2D cameras captures images from different altitudes, employing AI-based pose estimation and triangulation-based refinement to determine and adjust the pose of picked parts, reducing reliance on intermediate fixtures.

Benefits of technology

This approach enhances pose estimation accuracy, reduces system costs, and minimizes cycle time by refining the pose of parts using secondary 2D image sets and triangulation, ensuring precise placement on conveyor belts or molds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024051619_28082025_PF_FP_ABST
    Figure IB2024051619_28082025_PF_FP_ABST
Patent Text Reader

Abstract

Random bin picking with a multi-view system includes capturing a first set of images of a set of parts. A pose estimation model is used to predict a first coarse position for a target part. The first coarse position for the target part is refined by determining a position of the target part in a combined coordinate system generated by combining information from the first set of images. A robotic arm picks up the target part and moves the robotic arm with the target part to a preset location. A second set of images of the target part are captured and a second coarse position is predicted for the target part using the pose estimation model. The second coarse position is refined for the target part for placement by determining a position of the target part in a second combined coordinate system generated based on the second set of images.
Need to check novelty before this filing date? Find Prior Art

Description

RANDOM BIN PICKING AND PRECISEPLACEMENT WITH A MULTIVIEW SYSTEMFIELD

[0001] The present disclosure relates to a random bin picking system. In particular, the present disclosure relates to precise picking and placement with a multiview system in robotic arm-based random bin picking applications.BACKGROUND

[0002] For industrial machine tending applications, the robot arms are required to accurately pick an industrial part that is randomly positioned in a standard container for subsequent processing including blow off, wash, deburring, gauging, inspection, and / or precisely placing the part on the conveyer belt. In order to accurately place the mechanical part after the initial picking process, an accurate pose estimation of the mechanical part is required. Most current random bin picking applications are based on a three-dimensional (3D) vision system. Such vision systems can achieve highly accurate pose estimation of industrial parts for picking automation but are expensive to be integrated into the random bin picking system. Alternatively, a low-cost single red-green-blue (RGB) or grayscale camera may be used for such applications, but those systems suffer from limited picking performance due to poor pose estimation accuracy. In some random bin picking applications, precise placement is required. Due to the grasping contacts and interactions between the grippers and mechanical parts, the pose of the picked part cannot be accurately determined for precise placement. Current industrial implementation involves using an intermediate fixture for the robot arm to place the picked part into and repick the part after the fixture corrects the pose. This repose step affects the cycle time of such pick-and-place system.SUMMARY

[0003] A first aspect of the present disclosure provides a method for random bin picking with a multi view system, the method comprising: capturing a first set of images of a set of mechanical parts randomly distributed in a bin; determining a subset of target parts from the set of mechanical parts based on the captured images; predicting a first coarse position for each target part of the subset of target parts using an Al-based pose estimation model; refining the predicted first coarse position for each target part of the set of target parts,wherein refining the predicted coarse position for each target part comprises determining a position of each target part in a combined coordinate system generated by combining information from the first set of images; providing instructions to a robotic arm to pick up a target part from the set of target parts and move the robotic arm with the target part to a preset location; capturing a second set of images of the target part as the target part is held by the robotic arm at the preset location; predicting a second coarse position for the target part using the pose estimation model; and refining the second predicted coarse position for the target part, wherein refining the second predicted coarse position for the target part for precise placement comprises determining a position of the target part in a second combined coordinate system generated by combining information from the second set of images.

[0004] According to an implementation of the first aspect, the method further comprises placing the target part on a conveyor belt for further processing based on the refined second predicted coarse position.

[0005] According to an implementation of the first aspect, the method further comprises placing the target part in a pre-mold for further processing based on the refined second predicted coarse position.

[0006] According to an implementation of the first aspect, the first set of images are captured simultaneously using a first camera and a second camera, wherein the first camera and the second camera are color cameras.

[0007] According to an implementation of the first aspect, the first camera and the second camera are arranged at a different altitude with respect to the bin.

[0008] According to an implementation of the first aspect, the prediction of the second coarse position and refining the second coarse position is performed using a known position of a tool center tip of the robotic arm.

[0009] A second aspect of the present disclosure provides a system for random bin picking with a multiview system, the system comprising: a controller, wherein the controller is configured to: capture a first set of images of a set of mechanical parts randomly distributed in a bin; determine a subset of target parts from the set of mechanical parts based on the captured images; predict a first coarse position for each target part of the subset of target parts using an Al-based pose estimation model; refine the predicted first coarse position for each target part of the set of target parts, wherein refining the predicted coarse position for each target part comprises determining a position of each target part in a combined coordinate system generated by combining information from the first set of images; provide instructions to a robotic arm to pick up a target part from the set of targetparts and move the robotic arm with the target part to a preset location; capture a second set of images of the target part as the target part is held by the robotic arm at the preset location; predict a second coarse position for the target part using the pose estimation model; and refine the second predicted coarse position for the target part, wherein refining the second predicted coarse position for the target part for precise placement comprises determining a position of the target part in a second combined coordinate system generated by combining information from the second set of images.

[0010] According to an implementation of the second aspect, the controller is further configured to place the target part on a conveyor belt for further processing based on the refined second predicted coarse position.

[0011] According to an implementation of the second aspect, the controller is further configured to place the target part in a pre-mold for further processing based on the refined second predicted coarse position.

[0012] According to an implementation of the second aspect, the first set of images are captured simultaneously using a first camera and a second camera, wherein the first camera and the second camera are color cameras.

[0013] According to an implementation of the second aspect, the first camera and the second camera are arranged at a different altitude with respect to the bin.

[0014] According to an implementation of the second aspect, the prediction of the second coarse position and refining the second coarse position is performed using a known position of a tool center tip of the robotic arm.BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Subject matter of the present disclosure will be described in even greater detail below based on the exemplary figures. All features described and / or illustrated herein can be used alone or combined in different combinations. The features and advantages of various embodiments will become apparent by reading the following detailed description with reference to the attached drawings, which illustrate the following:

[0016] FIG. 1 illustrates a simplified system for a random bin picking system, according to one or more examples of the present disclosure;

[0017] FIG. 2 illustrates an exemplary system diagram related to a multiview triangulation-based refinement technique, according to one or more examples of the present disclosure;

[0018] FIG. 3 illustrates exemplary diagrams related to a process of pose estimation after picking up a piece for precise placement, according to one or more examples of the present disclosure;

[0019] FIG. 4 is a simplified block diagram of one or more devices or systems within the exemplary environment of FIG. 1, according to one or more examples of the present disclosure; and

[0020] FIG. 5 illustrates a process performed by a controller as part of a random bin picking system, according to one or more examples of the present disclosure.DETAILED DESCRIPTION

[0021] Examples of the presented application will now be described more fully hereinafter with reference to the accompanying FIGS., in which some, but not all, examples of the application are shown. Indeed, the application may be exemplified in different forms and should not be construed as limited to the examples set forth herein; rather, these examples are provided so that the application will satisfy applicable legal requirements. Where possible, any terms expressed in the singular form herein are meant to also include the plural form and vice versa, unless explicitly stated otherwise. Also, as used herein, the term “a” and / or “an” shall mean “one or more” even though the phrase “one or more” is also used herein. Furthermore, when it is said herein that something is “based on” something else, it may be based on one or more other things as well. In other words, unless expressly indicated otherwise, as used herein “based on” means “based at least in part on” or “based at least partially on.”

[0022] The present disclosure focuses on solving the random bin picking and precise placement problem using a multiview vision system. The present disclosure provides an AI- based pose estimation vision system that consists of two or more two-dimensional (2D) cameras and performs precise pose estimation of mechanical parts for pick-and-place applications using robotic arms. Images from each of the 2D cameras are provided to the pose estimation vision system to perform different pose estimations of target mechanical parts using the different images for picking. Subsequently, a triangulation-based pose and / or pose refinement is performed for precise picking of the mechanical part.

[0023] In some cases, due to the contacts between a picking end effector of the robotic arm and the mechanical part, the pose of the picked mechanical part may shift slightly, which may cause unsuccessful placement of the part for subsequent machine tending processes.Conventionally, the pose discrepancy of the mechanical part is fixed by placing the part in / on a fixture at an intermediate location and picking it up again. However, this process may have a negative impact on cycle time and requires an elaborate setup. While achieving satisfactory performance and maintaining the low system cost, this system adopts a multiview system based on cost-effective 2D cameras for random bin picking.

[0024] In order to avoid unsuccessful placements, the present disclosure describes performing an adjustment of the pose estimation of the mechanical part after it is picked up in order to adjust for any change in the pose of the mechanical part after it is picked up and held by the robotic arm and thereby provide successful placement of the mechanical part. In some embodiments, the pose estimation adjustment is performed using secondary 2D image sets. The secondary 2D image sets are captured after the part is picked up and attached to the picking end effectors. The secondary 2D image sets are used to predict a second coarse pose of the mechanical part with deep learning -based algorithms. The second coarse pose is refined using a triangulation-based pose refiner module. In order to predict the second coarse pose and perform the second fine refinement of the second coarse pose, a tool center point (TCP) location of the robotic arm is determined. The TCP location is used as a reference point which helps in determining the second coarse pose and subsequently refining the second coarse pose, and placing the mechanical part.

[0025] FIG. 1 illustrates a simplified diagram for a random bin picking system, according to one or more examples of the present disclosure. System 100 includes a controller 104, a vision system 108, a robotic arm 102, and a memory 106. The controller 104 instructs the robotic arm 102 to precisely pick up a mechanical part that has been randomly placed in a bin and then to place that mechanical part in a location for further processing. The controller 104 determines a first coarse position of the mechanical part randomly placed in the bin and performs a first pose refinement of the first coarse position using a pose refiner module. In some embodiments, the first coarse position is obtained via an Al -based deep learning neural network. The Al-based deep learning neural network receives a plurality of RGB or monochrome 2D images of the mechanical parts in the bin. A plurality of mechanical parts are determined from the plurality of 2D images. A 6 degree-of-freedom (6D) pose of each detected mechanical part in the bin is predicted using the Al-based deep learning neural network. The predicted 6D pose may be expressed in 3-dimensional (3D) translation and 3D orientation. The Al-based deep learning model is applied to each view from each image of the plurality of images. In some examples, the predicted 6D pose for each part is refined. The refinement of the predicted pose is based on multiview triangulation. In order to performmultiview triangulation, a set of 3D points are pre-defined or sampled from a CAD mesh model of each detected mechanical part in the bin. In some cases, the CAD mesh model may be stored in memory 106. The 3D points of the CAD mesh model are converted to 2- dimensional points in an image space. With the coarse pose of the mechanical part in each view, 2-dimensional (2D) points (pixel coordinates) in an image space of the 3D points selected from the CAD model are computed via back re-projection. A multiview triangulation is applied to each point in the set of sampled points. Using the multiview triangulation on the 2D points, the 3D coordinates in the camera view of the selected points are computed. The refined pose of the target part is easily computed with the 3D coordinates in the camera view using a rigid body constraint.

[0026] The coarse pose determination and pose refinement may be used to efficiently pick up the industrial part using a robotic arm for further assembly processing.

[0027] The pose estimation vision process as shown by system 100 is implemented in two stages. In a first stage, using a vision system 108, images are taken of a bin in which parts are randomly distributed. In some cases, the vision system 108 may include at least two 2D cameras for taking images of the bins that are filled with randomly distributed parts. In some other cases, the at least two cameras of the vision system 108 may be arranged at different angles and different altitudes to capture images of the industrial part in the bin from different angles. For example, the first camera 202 may be arranged at a first height above the bin and the second camera 204 may be arranged at a second height above the bin, wherein the first height and the second height are different.

[0028] In some embodiments, a first set of images are taken by the at least two cameras of the vision system 108. From the first set of images of the bin of parts, target mechanical parts are identified to be picked up. In some cases, an off-the-shelf MaskRCNN model may be used to identify the target parts. The MaskRCNN model may be retrained or fine-tuned with generated synthetic dataset given ground truth information such as the part identifiers (IDs), 2D bounding boxes, etc. Other state-of-the-art object detection models such as YOLO variants may also be used.

[0029] For each target mechanical part, a first coarse pose is estimated using a neural network model 110. Then, the first set of images from the at least two cameras of the vision system are combined to refine the first coarse pose for each target mechanical part using a pose refinement module. Using the first refined pose that is generated by refining the first coarse pose, the controller 104 instructs the robotic arm 102 to pick up one of the target mechanical parts from the bin. After the target mechanical part is picked up, the robotic arm102 moves with the target mechanical part to a preset location. The vision system 108 may take a second set of images of the mechanical part, which is now held by the robotic arm 102. Using the second set of images, a second coarse pose of the part is determined which is further refined by combining information from second set of images. The controller 104 may adjust the second coarse pose of the mechanical part while it is held by the robotic arm 102 based on the second coarse pose determination and the output of the pose refinement module. Once the second coarse pose of the part is refined in the robotic arm 102, the robotic arm places the mechanical part for processing in premade molds or on a conveyer belt.

[0030] As noted, in the first stage the vision system 108 captures a first set of images that are related to real-world random bin scenes. In some embodiments, the bin may include a collection of industrial parts that are randomly distributed throughout the bin. The vision system 108 may be used to photograph the bin. The first set of images that are captured by the at least two cameras of the vision system are used to estimate a coarse position of each target mechanical part in the bin using an Al -based neural network 110. Subsequently, the first set of images is combined to perform a first pose refinement. As discussed above, the first coarse position is obtained via an Al-based deep learning neural network. The Al-based deep learning neural network receives a 2D image of the mechanical parts in the bin and predicts a 6-degree-of-freedom (6D) pose of each detected mechanical part in the bin. The estimated 6D pose may be expressed in 3-dimensional (3D) translation and 3D orientation. A set of 3D points are pre-defined or sampled from a mesh model of each detected mechanical part in the bin. With the coarse pose of the mechanical part in each view, the 2-dimensional (2D) pixel coordinates in image space of the 3D points selected in a CAD model are computed via back re-projection. A multiview triangulation is applied to each point in the set of selected points. Thus, the 3D coordinates in the camera view of the selected points are computed. The refined pose of the target part is easily computed with the 3D coordinates in the camera view using a rigid body constraint. The controller 104 then instructs the robotic arm 102 to pick up the mechanical part based on the first coarse pose and the first pose refinement.

[0031] In some embodiments, the 6D pose estimation including two steps may be performed twice. A first 6D pose estimation step may be used to determine a coarse position and the second 6D pose estimation step may be used to refine the determined coarse position. The first 6D pose estimation may be used to perform the picking. The second 6D pose estimation may be used to determine the pose of the picked part and place it to a mold or conveyor belt.

[0032] The Al-based neural network model 110 that may be used to perform the first coarse pose estimation of each target part that is randomly distributed in a bin may be stored in memory 106. The Al-based neural network 110 is applied by the controller to the images of the bin full of parts captured by the vision system 108. The neural network model 110 is used to predict a first coarse pose for a mechanical part in the collection of industrial parts that are randomly distributed in the bin. In such cases, the neural network model 110 may take the images of the bin captured by the vision system 108 and process the images to predict the first coarse poses of parts that are randomly distributed in the bin.

[0033] The neural network 110 may be trained using a training dataset 112 stored in memory 106. The training dataset 112 may include computer-aided design (CAD) mesh models for the mechanical parts that are randomly distributed in the bin. In some embodiments, a large physically based rendered (PBR) synthetic dataset can be generated along with ground truth labels for training 2D-based vision models. Physically based rendered synthetic images are photorealistic and better reflect the real lighting and reflections. Physically based rendering is a method of shading and rendering technique that provides more accurate representations of how light interacts with different material properties. The ground truth labels may include a part identifier (ID), a 2D bounding box of the part, part visibility in the view, and the pose of the part that consists of translation in xyz- axis and its orientation represented by a 3x3 rotation matrix.

[0034] A coarse pose estimation model, implemented using a neural network 110, may be a one-stage pose estimator or two-stage pose estimator that consists of an object detector and a pose predictor. In some cases, mechanical parts are detected by a deep learning-based object detector that is fine-tuned or retrained to recognize the picking targets. The object detector used here is based on MaskRCNN. However, the object detector can be easily changed to other you only look once (YOLO) based algorithms. In some cases the part detector is trained with a dataset of 2D images and ground truth labels consisting of the identifier of the mechanical part and a 2D bounding box.

[0035] In some embodiments, the object detector may first detect an object from the images captured by the vision system 108. Then the pose predictor may be used to determine a pose of the object detected by object detector. In some embodiments, the pose predictor may be a coarse pose predictor implemented using the neural network 110. In some embodiments, the object detector and the coarse pose estimator may be combined into a single stage pose estimator, implemented by the neural network 110.

[0036] After the first coarse pose estimation, a multiview triangulation is applied to refine the grasping pose and / or object pose for precise picking of the mechanical part. In some embodiments, the multiview triangulation involves combining the data acquired by the different cameras of the vision system 108. In some cases, combining the information from the first set of images may involve combining the coordinate systems of the various images of the first set of images. The position of the mechanical parts in the bin may then be generated in a common coordinate system. The combination of the information from the different images may be helpful in ascertaining a depth of the target mechanical part in the bin.

[0037] Based on the first refined pose, the controller 104 instructs the robotic arm 102 to pick up the mechanical part from the bin. Once the part is picked up, the robotic arm 102, while holding the mechanical part, moves to a preset position for secondary 2D imaging. The different cameras of the vision system 108 may take a second set of images of the target mechanical part as it is held in the robotic arm 102. A second coarse pose of the target mechanical part is determined while it is held by the robotic arm 102. The second coarse pose may be determined by providing the second set of images to the neural network model 110. In some embodiments, the second coarse position of the mechanical part may be represented in a six-degree-of-freedom (6D) pose consisting of 3D translation and 3D rotation. The information from the second set of images is then combined to perform a second pose refinement. In some embodiments, the second coarse position and second pose refinement of the second pose of the mechanical part is determined with respect to a known tool center point (TCP) of the robotic arm 102. The known tool center point position helps to determine the pose range for synthetic data generation for training a close-up view Al model that estimates the coarse pose of the part when picked in the robot arm. With the part being in the known TCP position, the part association step may be eliminated in the multiview vision system. It can be two different Al models in coarse pose estimation when parts are in the bin and when parts are picked in the robot arm. In some embodiments, the second estimated pose is not determined with respect to the TCP.

[0038] Once the pose of the mechanical part in the robotic arm 102 is refined, the controller 102 instructs the robotic arm 102 to place the mechanical part for subsequent process. In some embodiments, the mechanical part may be placed in pre-molds or conveyor belts for further processing.

[0039] FIG. 2 illustrates an exemplary system diagram related to a multiview triangulation-based refinement technique, according to one or more examples of the presentdisclosure. System 200 of FIG. 2 includes a first camera 202 and a second camera 204. The first camera 202 and the second camera 204 may be part of the vision system 108 as described in relation to FIG. 1. In some embodiments, both the first camera 202 and the second camera 204 are color cameras or monochrome cameras. As discussed above, the first camera 202 and the second camera 204 may be arranged at different heights above the bin so that the depth of the mechanical parts arranged in the bin is captured better. The first camera 202 and the second camera 204 are separated by a fixed distance ‘b’ 206. Both the first camera 202 and the second camera 204 are used to simultaneously capture the images of the mechanical part randomly distributed in the bin. A coordinate system 208 is associated with the image captured by the first camera 202 and a second coordinate system 210 is associated with the image captured by the second camera 204. The images of the first camera 202 and the second camera 204 are used to determine a coarse pose associated with a mechanical part in the two images. In some embodiments, the images of the first camera 202 and the second camera 204 are provided to the neural network model 110 to determine the coarse position of the mechanical part. The coordinate system of the first image 208 and the coordinate system of the second image 210 are in 2D image space. The combined coordinate system 212 that is generated using the coordinate system of the first image 208 and the coordinate system of the second image 210 is composed of 3D coordinates in a main camera view. The combined coordinate system 212 is used to perform a pose refinement of the coarse position of the mechanical part determined in the previous step. In order to perform the pose refinement of the coarse position, a position of the mechanical part in the bin may be generated in the combined coordinate system. Using a transformation matrix that describes the camera position in robot base, the part pose can then be described in the common robot base (the common 3D coordinate system).

[0040] FIG. 3 illustrates exemplary diagrams related to a process of pose estimation after picking up a piece for precise placement, according to one or more examples of the present disclosure. A first image 302 of FIG. 3 depicts a random distribution of mechanical parts in a bin. The second image 304 depicts an arrangement of a robotic arm 310 over the bin and a vision system 312. The robotic arm 310 is similar to the robotic arm 102 and the vision system 312 is similar to the vision system 108. The vision system 312 may include the first camera 202 and the second camera 204 as shown in FIG. 2. As discussed above, the vision system 108 captures a first set of images of the random distribution of mechanical parts in the bin. The first set of images are used by a neural network 110 to determine a first coarse position of a set of target parts from the random distribution of mechanical parts in the bin.Subsequently, the information from the first set of images is combined to perform a first pose refinement of the first coarse position to generate a first refined position. The controller 104 instructs the robotic arm 102 to pick up a target part from the set of target parts using the first refined position.

[0041] Image 306 shows the robotic arm 310 after it picks up a mechanical part from the bin. After the mechanical part is picked up, it is moved to a position 314 for a secondary pose estimation. In order to perform the secondary pose estimation, a second set of images are captured by the vision system 312. After the target part is picked up using the first refined position. The controller 102 uses the second set of images captured by the vision system 108 to perform a second coarse position determination and a second pose refinement. Image 308 shows the position of the robotic arm 310 after the second pose refinement of the mechanical part held in the robotic arm 310 is performed.

[0042] FIG. 4 is a block diagram of an exemplary system or device 400 within the system 100 such as the controller 104. The system 400 includes a processor 104, such as a central processing unit (CPU), and / or logic, that executes computer executable instructions for performing the functions, processes, and / or methods described herein. In some examples, the computer executable instructions are locally stored and accessed from a non-transitory computer readable medium, such as storage 410, which may be a hard drive or flash drive. Read Only Memory (ROM) 406 includes computer executable instructions for initializing the processor 404, while the random-access memory (RAM) 408 is the main memory for loading and processing instructions executed by the processor 404. The network interface 412 may connect to a wired network or cellular network and to a local area network or wide area network. The system 400 may also include a bus 402 that connects the processor 404, ROM 406, RAM 408, storage 410, and / or the network interface 412. The components within the system 400 may use the bus 402 to communicate with each other. The components within the system 400 are merely exemplary and might not be inclusive of every component within the controller 104. Additionally, and / or alternatively, the system 400 may further include components that might not be included within every entity of system 400. For instance, in some examples, the controller 104 might not include a network interface 412. In some embodiments, the system 400 may include a computing accelerator (GPU) 414 for running the artificial intelligence (Al) deep learning neural network model.

[0043] FIG. 5 illustrates a process performed by a controller as part of a random bin picking system, according to one or more examples of the present disclosure. However, it will be recognized that any of the following blocks may be performed in any suitable orderand that the process 500 may be performed in any environment and by any suitable computing device and / or controller.

[0044] At 502, the vision system 108 captures images of at least one mechanical part in a bin. In some embodiments, the vision system 108 may include a first camera 202 and a second camera 204. The first camera 202 and the second camera 204 may be color cameras or monochrome cameras that may be installed at different altitudes relative to the bin to accurately measure the depth of the mechanical parts placed in the bin. In some embodiments, more than two cameras may be used to capture the images of the at least one mechanical part in the bin.

[0045] At 504, a first coarse pose of each target part is estimated. In some embodiments, the first set of images captured by the first camera 202 and the second camera 204 are provided to a neural network 110. The neural network 110 selects target parts and determines a first coarse position of the mechanical parts in the bin.

[0046] At 506, the mechanical part is associated in the multiview system. In some embodiments, the information from the first camera 202 and the second camera 204 are combined. As discussed in FIG. 2, the coordinates of images captured by the first camera 202 and the coordinates of the images captured by the second camera 204 are combined to generate a combined coordinate system. The first coarse position of the one or more mechanical parts in the bin may then be generated in the combined coordinate system.

[0047] At 508, the first coarse position of the target mechanical part is refined. In some embodiments, the combined information from multiple camera views is used to perform a refinement of the first coarse position.

[0048] At 510, the controller 104 instructs the robotic arm 102 to pick up the mechanical part based on the first refined pose of the mechanical part. In some embodiments, after the part is picked up, the robotic arm 102 moves to a preset position with the mechanical part for secondary pose estimation.

[0049] At 512, the pose of the mechanical part is estimated and refined after it is picked up. In some embodiments, a secondary pose estimation is performed after the part is picked up by the robotic arm 102 and the robotic arm 102 is moved to a preset location. The secondary pose estimation is performed using a second set of images of the robotic arm 102 and the mechanical part captured from the vision system 108. Using the second set of images that are captured by the vision system 108, a second coarse position of the mechanical part as it is held by the robotic arm 102 is determined. Then, the images captured by the visionsystem 108 are combined to perform a second refinement of the pose of the mechanical part as it is held by the robotic arm 102.

[0050] At 514, the mechanical part is placed on pre-made molds or conveyor belt for further processing. For example, the controller 104 uses second refined pose of the mechanical part as it is held by the robotic arm 102 direct the robot arm 102 to place the mechanical part on pre-made molds or conveyor belt in the correct pose for further processing.

[0051] While subject matter of the present disclosure has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive. Any statement made herein characterizing the invention is also to be considered illustrative or exemplary and not restrictive as the invention is defined by the claims. It will be understood that changes and modifications may be made, by those of ordinary skill in the art, within the scope of the following claims, which may include any combination of features from different embodiments described above.

[0052] The terms used in the claims should be construed to have the broadest reasonable interpretation consistent with the foregoing description. For example, the use of the article “a” or “the” in introducing an element should not be interpreted as being exclusive of a plurality of elements. Likewise, the recitation of “or” should be interpreted as being inclusive, such that the recitation of “A or B” is not exclusive of “A and B,” unless it is clear from the context or the foregoing description that only one of A and B is intended. Further, the recitation of “at least one of A, B and C” should be interpreted as one or more of a group of elements consisting of A, B and C, and should not be interpreted as requiring at least one of each of the listed elements A, B and C, regardless of whether A, B and C are related as categories or otherwise. Moreover, the recitation of “A, B and / or C” or “at least one of A, B or C” should be interpreted as including any singular entity from the listed elements, e.g., A, any subset from the listed elements, e.g., A and B, or the entire list of elements A, B and C.

Claims

CLAIMSWhat is claimed is:

1. A method for random bin picking with a multiview system, the method comprising: capturing a first set of images of a set of mechanical parts randomly distributed in a bin; determining a subset of target parts from the set of mechanical parts based on the captured images; predicting a first coarse position for each target part of the subset of target parts using an Al-based pose estimation model; refining the predicted first coarse position for each target part of the set of target parts, wherein refining the predicted coarse position for each target part comprises determining a position of each target part in a combined coordinate system generated by combining information from the first set of images; providing instructions to a robotic arm to pick up a target part from the set of target parts and move the robotic arm with the target part to a preset location; capturing a second set of images of the target part as the target part is held by the robotic arm at the preset location; predicting a second coarse position for the target part using the pose estimation model; and refining the second predicted coarse position for the target part, wherein refining the second predicted coarse position for the target part for precise placement comprises determining a position of the target part in a second combined coordinate system generated by combining information from the second set of images.

2. The method of claim 1, wherein the method further comprises placing the target part on a conveyor belt for further processing based on the refined second predicted coarse position.

3. The method of claim 1, wherein the method further comprises placing the target part in a pre-mold for further processing based on the refined second predicted coarse position.

4. The method of claim 1, wherein the first set of images are captured simultaneously using a first camera and a second camera, wherein the first camera and the second camera are color cameras.

5. The method of claim 6, wherein the first camera and the second camera are arranged at a different altitude with respect to the bin.

6. The method of claim 1, wherein the prediction of the second coarse position and refining the second coarse position is performed using a known position of a tool center tip of the robotic arm.

7. A system for random bin picking with a multi view system, the system comprising: a controller, wherein the controller is configured to: capture a first set of images of a set of mechanical parts randomly distributed in a bin; determine a subset of target parts from the set of mechanical parts based on the captured images; predict a first coarse position for each target part of the subset of target parts using an Al-based pose estimation model; refine the predicted first coarse position for each target part of the set of target parts, wherein refining the predicted coarse position for each target part comprises determining a position of each target part in a combined coordinate system generated by combining information from the first set of images; provide instructions to a robotic arm to pick up a target part from the set of target parts and move the robotic arm with the target part to a preset location; capture a second set of images of the target part as the target part is held by the robotic arm at the preset location; predict a second coarse position for the target part using the pose estimation model; and refine the second predicted coarse position for the target part, wherein refining the second predicted coarse position for the target part for precise placement comprises determining a position of the target part in a second combined coordinate system generated by combining information from the second set of images.

8. The system of claim 7, wherein the controller is further configured to place the target part on a conveyor belt for further processing based on the refined second predicted coarse position.

9. The system of claim 7, wherein the controller is further configured to place the target part in a pre-mold for further processing based on the refined second predicted coarse position.

10. The system of claim 7, wherein the first set of images are captured simultaneously using a first camera and a second camera, wherein the first camera and the second camera are color cameras.

11. The system of claim 10, wherein the first camera and the second camera are arranged at a different altitude with respect to the bin.

12. The system of claim 7, wherein the prediction of the second coarse position and refining the second coarse position is performed using a known position of a tool center tip of the robotic arm.

Citation Information

Patent Citations

  • Pick and place systems and methods

    CA3202375A1

  • System for imaging and orienting seeds and method of use

    US20150321353A1

  • In-hand pose refinement for pick and place automation

    US20230071384A1

  • Reliable robotic manipulation in a cluttered environment

    US20230339118A1