Method and system for determining pose prediction confidence in bin picking applications
The system improves robotic grasping accuracy by using image and depth sensors with refined pose estimation algorithms, addressing the limitations of conventional systems in cost and performance, enabling efficient handling of randomly positioned objects.
Patent Information
- Application Number
- PCT/IB2024/055447
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-04
- Publication Date
- 2025-12-11
AI Technical Summary
Conventional robotic systems for randomly positioned object grasping face challenges due to high cost and complexity of 3D vision camera setups and limited pose prediction performance of smaller systems, requiring complex parameter tuning.
A system utilizing image sensors, depth sensors, and a robotic arm with a processor and memory to generate initial pose estimations, refine them using ICP registration, and determine grasp accuracy based on threshold comparisons, employing algorithms like stereo matching and ICP to ensure precise robotic manipulation.
Enhances pose prediction accuracy for robotic grasping, allowing efficient and reliable handling of randomly positioned objects, reducing the need for costly and complex camera systems.
Smart Images

Figure IB2024055447_11122025_PF_FP_ABST
Abstract
Description
METHOD AND SYSTEM FOR DETERMINING POSEPREDICTION CONFIDENCE IN BIN PICKING APPLICATIONSFIELD
[0001] The present disclosure relates to a method and system for determining pose prediction accuracy of an object for motion planning of a robot arm to grasp the object randomly positioned in a bin.BACKGROUND
[0002] Facilities which implement and utilize robotic systems for interacting with objects (e.g., mechanical parts) for further processing may require automatically selecting and retrieving randomly positioned and oriented objects from a container or bin. For example, manufacturers may use a number of robotic arms or manipulators to retrieve objects randomly positioned in a container to place it on a conveyer belt or perform subsequent processing such as assembly, blow off, wash, deburring, gauging, and inspection in a facility during manufacture to more efficiently carry out the manufacturing process. The robotic systems may be required to identify particular parts or objects in a given space of a view in order to grasp the object or perform other operations. Conventional methods of identifying objects in a space for robotic manipulation may include a three dimensional (3D) vison camera for predicting a pose and location of an object within a space. However, such camera system rigs may require a large foot print within a facility to properly implement and are expensive. Other conventional systems may utilize smaller or less technology proficient image capturing systems that have limited pose prediction performance and require complex parameter tuning onsite to properly identify the object and the pose of the object.SUMMARY
[0003] An embodiment of the present disclosure provides a system including an image sensor, a depth sensor, a robotic arm, a processor, and a memory including instructions that, when executed with the processor, cause the system to, at least receive an image of a plurality of objects in a bin captured by the image sensor, the plurality of objects and the bin being a proximal distance from the image sensor, the depth sensor, and the robotic arm, receive depth information for the plurality of objects captured by the depth sensor, generate an initial pose estimation for an object of the plurality of objects using a pose estimation algorithm and theimage, generate a depth map for the object using the depth information, render a depth for the object using the initial pose estimation for the object and camera intrinsics of the image sensor, where the camera intrinsics indicate imaging parameters that are used to capture the image, determine a value that represents an accuracy of the depth for the object in the bin based on comparing the depth map for the object and the rendered depth for the object, and instruct the robotic arm to grasp the object based on comparing the value with a first threshold.
[0004] In an embodiment of the system, the instructions are configured to implement an iterative closest point (ICP) registration algorithm to generate an updated pose estimation for the object based on the value being between the first threshold and a second threshold.
[0005] In an embodiment of the system, the instructions are configured to render an updated depth for the object using the updated pose estimation for the object and the camera intrinsics, and determine an updated value that represents the accuracy of the depth for the object in the bin based on comparing the depth map for the object and the updated rendered depth for the object.
[0006] In an embodiment of the system, the instructions are configured to suppress transmission of instructions to the robotic arm to grasp the object based on comparing the value or the updated value with the second threshold or a third threshold.
[0007] In an embodiment of the system, wherein the instructing the robotic arm to grasp the object is further based on the initial pose estimation.
[0008] In an embodiment of the system, wherein the camera intrinsics include a focal length and camera principal point.
[0009] In an embodiment of the system, wherein the image is a color image or a mono color image.
[0010] In an embodiment of the system, wherein the first threshold is specified by a user through a configuration file or a graphical user interface.
[0011] In an embodiment of the system, wherein the first threshold is based at least in part on physical properties of the object or a grasping tolerance of the robotic arm.
[0012] In an embodiment of the system, wherein generating the initial pose estimation for the object includes using an artificial neural network trained with images of objects.
[0013] Another embodiment of the present disclosure provides a system including two image sensors, a robotic arm, a processor, and a memory including instructions that, when executed with the processor, cause the system to, at least receive images of a plurality of objects in a bin from multiple views captured by the two image sensors, the plurality ofobjects and the bin being a proximal distance from the two image sensors and the robotic arm, generate an initial pose estimation for an object of the plurality of objects using a pose estimation algorithm and an image of the images, generate a depth map for the object using a stereo matching algorithm and the images, render a depth for the object using the initial pose estimation for the object and camera intrinsics of the two image sensors, where the camera intrinsics indicate imaging parameters that are used to capture the image, determine a value that represents an accuracy of the depth for the object in the bin based on comparing the depth map for the object and the rendered depth for the object, and instruct the robotic arm to grasp the object based on comparing the value with a first threshold.
[0014] In an embodiment of the system, wherein the instructions are configured to use a triangulation technique to refine the initial pose estimation.
[0015] In an embodiment of the system, wherein generating the initial pose estimation for the object includes using the pose estimation algorithm and the images.
[0016] In an embodiment of the system, wherein the depth map for the object is generated using block matching, semi-global matching, an artificial intelligence deep learning based stereo matching algorithm, or multi-view matching algorithm.
[0017] In an embodiment of the system, wherein the instructions are configured to implement an iterative closest point (ICP) registration algorithm to generate an updated pose estimation for the object based on the value being between the first threshold and a second threshold.
[0018] In an embodiment of the system, wherein the instructing the robotic arm to grasp the object is further based on the initial pose estimation.
[0019] Another embodiment of the present disclosure provides a computer-implemented method for determining pose prediction accuracy including receiving, from two image sensors, images of a plurality of objects in a bin from multiple views that are captured by the two image sensors, generating an initial pose estimation for an object of the plurality of objects using a pose estimation algorithm and an image of the images, generating a depth map for the object using a stereo matching algorithm and the images, rendering a depth for the object using the initial pose estimation for the object and camera intrinsics of the two image sensors, where the camera intrinsics indicate imaging parameters that are used to capture the image, determining a value that represents an accuracy of the depth for the object in the bin based on comparing the depth map for the object and the rendered depth for the object, and instructing a robotic arm to grasp the object based on comparing the value with afirst threshold, the robot arm being a proximal distance from the two image sensors, the plurality of objects, and the bin.
[0020] In an embodiment, the computer-implemented method further includes implementing an iterative closest point (ICP) registration algorithm to generate an updated pose estimation for the object based on the value being between the first threshold and a second threshold.
[0021] In an embodiment, the computer-implemented method further includes suppressing grasping of the object based on the value being higher than a second threshold and a number of objects in the bin.
[0022] In an embodiment of the computer-implemented method instructing the robotic arm to grasp the object is further based on the initial pose estimation.BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The present disclosure will be described in even greater detail below based on the exemplary figures. The disclosure is not limited to the exemplary embodiments. All features described and / or illustrated herein can be used alone or combined in different combinations in embodiments of the disclosure. The features and advantages of various embodiments of the present disclosure will become apparent by reading the following detailed description with reference to the attached drawings which illustrate the following:
[0024] FIG. 1 illustrates an example architecture of a system for determining pose prediction accuracy including a computer system, one or more image sensors, a robot arm, and a plurality of objects within a bin according to embodiments described herein;
[0025] FIG. 2 illustrates an example diagram for determining pose prediction accuracy according to embodiments described herein;
[0026] FIG. 3 illustrates examples of thresholds for comparison to determined depths of an object for motion planning according to embodiments described herein;
[0027] FIG. 4 illustrates an example diagram for determining pose prediction accuracy according to embodiments described herein;
[0028] FIG. 5 illustrates an example diagram for determining pose prediction accuracy according to embodiments described herein;
[0029] FIG. 6 illustrates a flow chart for determining pose prediction accuracy according to embodiments described herein; and
[0030] FIG. 7 illustrates a simplified block diagram of one or more devices or systems for determining pose prediction accuracy according to embodiments described herein.DETAILED DESCRIPTION
[0031] Embodiments of the present disclosure provide a method and system for determining pose prediction or estimation accuracy of a vision-based pose predictor in robotic arm-based random bin picking applications. While the present disclosure is described primarily in connection with machines, systems, or components operated in an industrial setting or environment, as would be recognized by a person of ordinary skill in the art, the disclosure is not so limited and inventive features apply to other components or systems of other environments.
[0032] High speed automated assembly lines involve robots that are responsible for bringing together different industrial parts to build an object. Machine tending robots of highspeed assembly lines may include tool-tips that may be configured to pick up industrial parts (objects) from different bins and place the parts picked up on a conveyer belt for further processing. In some embodiments, further processing of industrial parts may include other robot arms picking up the part for assembly, blow off, wash, deburring, gauging, inspection, etc. High speed automated assembly lines are generally used in assembling automobiles, airplanes, printed circuit boards, and the like. For the high speed automated assembly lines to function smoothly, the industrial parts that are randomly positioned in a bin may be required to be correctly picked up. In order for the industrial parts to be picked up in the correct manner, a perception system implemented by a computer system described herein may be required to identify the industrial parts that are randomly placed in the bin and a pose estimation model may precisely predict a six-degree-of-freedom (6DoF) pose in space of the industrial parts in order to correctly grasp the industrial parts using the robot arm and correctly place the industrial parts on a conveyer belt or perform some other subsequent processing. 6DoF may refer to the process of determining the precise position and orientation of an object in a three-dimensional (3D) space. This may involve estimating three translational movements (e.g., forward / back, up / down, left / right) and three rotational movements (pitch, yaw, roll) of the object relative to a given reference frame. This disclosure provides systems and methods for determining the pose prediction accuracy when a visionbased pose prediction model is used to predict the 6DoF poses for industrial parts randomlydistributed in a bin based on real images captured by cameras that may be further refined according to embodiments described herein.
[0033] A pose prediction model can be designed with feature-based computer vision techniques or learning-based Al techniques using industrial image cameras.
[0034] Embodiments of the present disclosure describe a process to verify or determine an accuracy of a pose estimation for an object using vision-based pose prediction systems at runtime in robotic -arm-based random bin picking applications. Determining an accuracy or further refining a pose estimation for an object is an important step to achieve high grasp rate during deployment. With accurate pose estimation associated with its confidence score which is indicated by an accuracy check and downstream refinement techniques, robotic machinery may be used to automatically pick up objects in the proper manner so that they can be efficiently used in further applications such as assembly.
[0035] FIG. 1 illustrates an example architecture of a system 100 for determining pose prediction accuracy including a computer system 102, one or more image sensors 104, a robotic arm 106, and a plurality of objects 108 within a bin 110 according to embodiments described herein. In embodiments, the one or more image sensors 104 (e.g., mono color cameras or color cameras) may capture one or more images of the objects 108 within bin 110. The dashed lines in FIG. 1 may represent the capture range or area for each image sensor of the image sensors 104. The image sensors 104 may be in communication with the computer system 102 via an Internet connection, BLUETOOTH, WIFI, near field communication (NFC), or other suitable communication methods for transmitting the captured images to the computer system 102. The image sensors 104 may be within a proximal range of the objects 108 and bin 110 that corresponds to the operating range of the sensors or cameras that can vary depending on the type of sensor or camera included in the system 100.
[0036] The image sensors 104 may capture a set of images that are related to real -world random bin scenes. In some embodiments, a bin 110 may include a collection of objects 108 such as industrial parts that are randomly distributed throughout the bin 110. The image sensors 104 may be used to photograph the bin 110. In some cases, the image sensors 104 may include a monochrome camera and / or a color camera for taking pictures of the bin 110 that is filled with randomly distributed objects 108. In some other cases, the image sensors 104 may include multiple cameras arranged at different angles to capture pictures of an object(s) 108 in the bin 110 from different angles. The computer system 102 may implement an object detection model which may be further implemented using a neural network for identifying individual objects 108 from the objects randomly distributed in the bin 110. Theobjects 108 that are identified in the bin 110 may then be provided as input pose estimation algorithm implemented by the computer system 102. The computer system 102 may implement a pose estimation algorithm for generating an initial pose estimation for the detected objects 108, which may include performing a 6DoF pose estimation process on the detected individual objects 108. In some embodiments, the 6DoF pose estimation process may be refined using deep learning methods, iteratively closet point (ICP) registration, and / or multi-view geometry algorithms. The output of the pose estimation algorithm may be in the form of a six degrees of freedom (6DoF) pose of a given object 108.
[0037] In embodiments, the computer system 102 may implement a pose estimation algorithm for generating an initial pose estimation for an object of the objects 108 within the bin 110 using the image(s) captured by the image sensors 104. In an embodiment, the computer system 102 may implement and utilize a deep learning -based six-degrees of freedom (6DoF) pose estimation model or algorithm to obtain an initial pose for all the objects that are detected in the image(s) captured by the sensors 104. In accordance with at least one embodiment, the computer system 102 may generate one or more initial pose estimations for the object(s) using one or more images captured by the image sensors 104 from multiple views of the object(s). For example, the image sensors 104 of system 100 may be configured to capture different views of the objects 108 and bin 110 within an environment or viewing space. In an embodiment, the computer system 102 may implement a triangulation algorithm (triangulation refinement algorithm) to refine the initial pose estimation.
[0038] The computer system 102 may utilize a stereo matching algorithm and the images captured by the image sensors 104 to generate a depth map (e.g., point cloud) for a respective object of the objects 108. In embodiments, the computer system may also utilize an artificial neural network (ANN) that is trained using images of objects to generate a depth map for the object. The ANN may use the images captured by image sensors 104 to generate a depth map for an object(s) 108. The computer system 102 may utilize stereo or multi -view matching algorithms such as block matching, semi-global matching, or Al deep learning based matching algorithms to generate the depth map. As used herein, a point cloud may include a set of data points in a space which represent the external surfaces of objects and a depth map can be converted to its corresponding point cloud representation. These points may collectively represent the shape of an object or scene in three dimensions. Block matching (BM) may describe a method for calculating disparity maps from stereo image pairs (e.g., the images captured by the image sensors 104). Block matching includes an algorithm whoseobjective is to estimate the depth information from at least two cameras that observe the same scene from slightly different perspectives or different perspectives. Block matching may include defining a range of possible disparities that is typically set based on the geometry of the stereo setup (e.g., image sensors 104) and an expected range of scene depths. A window selection process may select a small block or window around each pixel in a reference image that may be used to find a corresponding block in the other image. A matching cost computation may be used for each block in the reference image where the algorithm shifts the block across a corresponding line (or a small area around it) in the other image, within the defined disparity range. At each position it may calculate a matching cost which quantifies how similar the two blocks are. The disparity at which a cost is minimized is considered a best match and the disparity value is assigned to the original pixel in the reference image. This process is repeated for every pixel or subset of pixels in the reference image to construct a full disparity map.
[0039] Semi-global matching (SGM) algorithms may be implemented by the computer system 102 for generating the depth map in embodiments. The SGM algorithm may include a computer vision algorithm that estimates a dense disparity map from a rectified stereo image pair. In embodiments, the SGM algorithm may compute a matching cost along multiple one dimensional (ID) paths in the image. These costs reflect how well pixels in one image match corresponding pixels in another image. Common methods to calculate these costs include using measures like Census Transform, Mutual Information, or simpler intensity difference metrics. SGM algorithms may use 8 or more paths covering horizontal, vertical, diagonal directions. For each path, the algorithm accumulates the costs while regulated by a smoothness constraint. This constraint penalizes large changes in disparity between adjacent pixels, promoting smoother and more consistent disparity maps. After aggregating costs from all paths for each pixel, the algorithm selects the disparity value that minimizes the total aggregated cost. Optional post-processing steps such as median filtering or left-right consistency checks can be applied to refine the disparity map. These steps may help to reduce noise and improve the accuracy of the disparity map, particularly in occluded or low-texture areas.
[0040] Al-based stereo or multi-view matching algorithms may be implemented by the computer system 102 for generating the depth map in embodiments using rectified images. The Al-based stereo matching algorithm may extract the 2D and / or 3D features using deep convolutional neural networks, generate image pair correlations and geometry encodings, and iteratively refine a disparity map using the aggregated cost function. The depth map and pointcloud of a scene can be further obtained using the computed disparity map. The Al-based stereo matching algorithm is robust to various lighting conditions and offers improved local matching details.
[0041] In embodiments, the computer system 102 may implement a rendering engine or rendering application for generating a depth or depth image of an object using the initial pose estimation for an object and camera intrinsics of the image sensors 104. The pose prediction of the initial pose estimation for the object generated by the computer system 102 using, for example, the stereo matching algorithm, may be determined by comparing the depth map of the stereo matching algorithm to the depth image of the object generated using the rendering application. In embodiments, the computer system 102 may calculate or determine a value that represents an accuracy of the depth for a given object 108 or indicates the difference of the depth for a given object 108 in the bin 110 by comparing the depth map for the object and the rendered depth for the object. The determined value reflects the pose estimation accuracy along the depth direction. For example, the following formula may be used for comparing depth information between the depth map generated using a stereo matching algorithm and the images of the given object 108 and the depth information, map, or depth of the given object 108 from the rendered depth image generated by the rendering application:where Nnrepresents the total pixels of a segmented object 108, p represents one pixel within the segmented object 108, Dptereodenotes the measured / estimated depth of pixel p within part n from a depth sensor (described in more detail below with reference to FIG. 5) or stereo matching algorithm for the pixel p within the segmented object 108, and D^enderdenotes the rendered depth for the pixel p within the segmented object 108 based on the initial pose estimation from a pose predictor model such as the pose estimation algorithm. The computed difference between the two masked depth maps for each part indicates a confidence or accuracy of the predicted pose of that part. The smaller the computed depth difference, the higher the confidence or accuracy of predicted pose for that part. In embodiments, the camera intrinsics of the image sensors 104 may by represented be a camera intrinsics matrix that includes at least a focal length of the image sensors 104 and a camera principal point of the image sensors 104.
[0042] The computer system 102 may compare the value that represents an accuracy of the depth for the object 108 as originally calculated by the initial pose estimation to one ormore thresholds, described in greater detail below with reference to FIG. 3, to determine an action to be taken by the robotic arm 106. For example, the computer system 102 may proceed to a motion planning process for the robotic arm 106 that includes the robotic arm grasping or otherwise interacting with the detected object 108 in bin 110 for further processing such as moving the object 108 to a different conveyer belt or machine. The computer system 102 may compare the determined value to a first threshold and if the value is less than the first threshold the computer system 102 may instruct the robotic arm 106 to grasp the object 108. This may represent a determination by the computer system 102 that the depth information included in the initial pose estimation for the object is within a certain range that the robotic arm 106 may grasp the object 108 without causing harm or damage to the object 108, other objects 108, the bin 110, or robotic arm 106 itself.
[0043] The computer system 102 may proceed to other operations such as further refinement of the initial pose estimation for the object in response to the value being less than or greater than other thresholds as described herein. For example, the computer system may use an ICP registration algorithm (ICP algorithm) to refine the initial pose estimation for the object 108 in response to the value being greater than a certain threshold or less than another certain threshold as described in FIG. 3. This represents the determination by the computer system 102 that the initial pose estimation for the object includes a depth that is too inaccurate or inconsistent for appropriate grasping by the robotic arm 106. In embodiments, the ICP algorithm may iteratively refine a transformation (i.e., translation and rotation) needed to align one point cloud (a source depth map) with another point cloud (a target depth map).
[0044] FIG. 2 illustrates an example diagram 200 for determining pose prediction accuracy according to embodiments described herein. FIGs. 2, 4, and 5 depict processes implemented by a computer system that is not depicted but that implements the steps or features of each component or process illustrated in FIGs. 2, 4, and 5. The diagram 200 depicts a process for determining pose prediction accuracy for a configuration of a system as that depicted in FIG. 1 that includes multiple image sensors as opposed to the configuration described below with reference to FIG. 5 which includes a single image sensor and a depth sensor. The diagram 200 includes a computer system receiving images from the image sensors for an object in a bin from a first view at 202 (View 1) and from a second view at 204. In embodiments, the image sensors may include color or mono image sensors. The image sensors may capture a scene of the objects and / or bin in the images. As described above, the computer system implementing the pose prediction accuracy features may utilizean algorithm, such as a stereo matching algorithm 206 that uses the rectified images from the image sensors 202 and 204 to generate a depth map 208 for the objects included in the images.
[0045] The diagram 200 also includes a pose estimation 210 for an object in an image from 202 using a pose estimation algorithm. The computer system may also generate or determine an initial pose estimation for the object from a different view point, such as from a second image sensor, at 204 (View 2). In some embodiments, the computer system may include generating a combined pose estimation at 210 which may combine or otherwise aggregate the images from different viewpoints 202 and 204 for the object. The pose estimation 210 may be implemented using a deep learning -based artificial neural network model or a conventional feature-based pose estimation model. The pose estimation 210 may contain an aggregated pose refinement module to refine the pose of the object using a refiner neural network or multi-view triangulation technique. The pose prediction accuracy features described herein, however, do not require the combination of images as depicted by the dashed line going from 204 (View 2) to the pose estimation 210. The diagram 200 includes the computer system implementing a real-time rendering engine (application) at 212 to render a depth map (depth information) for the object using the initial pose estimation (210) for the object and camera intrinsics for the image sensors.
[0046] In embodiments, the diagram 200 includes part selection for grasping determination at 214. The part selection for grasping includes comparing, by the computer system, the depth map 208 with the depth information or depth generated by the real-time rendering engine 212 to determine a value that represents an accuracy of the depth included in the initial pose estimation 210 for the object. As described herein, the value may be compared to one or more thresholds to determine whether the depth for the object is accurate and therefore suitable for motion planning of the robotic arm to interact with the object (e.g., grasp the object). In embodiments, the computer system may take other actions based on a comparison of the value to one or more thresholds. For example, the computer system may utilize one or more refinement algorithms or processes to further refine the initial pose estimation 210 or move on to analyzing a different object in the bin all together to avoid spending more time on a particular object. The diagram 200 includes motion planning 216 for the robotic arm by the computer system. For example, the computer system may instruct the robotic arm to grasp the object using the initial 6DoF pose estimation or refined 6DoF pose estimation for the object in the bin. The computer system may instruct the robotic arm to grasp the object, move the object, turn the object, and transport the object to anotherdestination such as a conveyer belt or to another machine for further processing such as assembly, blow off, wash, deburring, gauging, inspection, etc.
[0047] FIG. 3 illustrates examples of thresholds for comparison to determined depths of an object for motion planning according to embodiments described herein. For example, the value determined by the computer system implementing the pose prediction accuracy features described herein may represent an accuracy of the depth for the object in the bin included in the initial pose estimation for the object. This value may be associated with a group priority 300 that further represents an average depth difference with respect to the robot base of the robotic arm. The thresholds 302, 304, and 306 include values in millimeters (mm) that are example values used for comparison to the value generated by the computer system upon comparing the depth map generated by the stereo matching algorithm or multiview matching algorithm to the depth information of the rendered application from the initial pose estimation for the object and the camera intrinsics of the image sensors. In embodiments, the thresholds 302, 304, and 306 may be specified by a user through a configuration file or a graphical user interface.
[0048] In embodiments, the computer system may determine different motion planning scenarios or instructions for the robotic arm or other components of a system based on a comparison of the value to the thresholds 302, 304, and 306. For example, a value that is less than the first threshold 302 may represent that the computer system is confident in the depth included in the initial pose estimation for the object. Confidence in depth of the initial pose estimation for the object is used to generate instructions to use the robotic arm to grasp the object and perform some other action such as moving the object from the bin to another location. To continue the example, a value that is greater than the first threshold 302 but less than a third threshold 306 (e.g., falls within a range of the second threshold 304) may be determined to be less accurate and thus further refinement may be required. For example, the computer system may utilize ICP registration refinement to further refine the initial pose estimation before determining to grasp the object. In cases where the value exceeds the second threshold 304 and fails in the third category 306, the computer system may move on to determining a 6DoF refined pose estimation and planning grasping for another object in the bin. The thresholds may be dynamically adjusted, for example, when there are a few parts left in the bin, the thresholds are changed to allow the pose of more or all the parts in the bin to be refined to empty the container. In embodiments, the system may utilize two thresholds, e.g., first threshold 302 and a second threshold that corresponds to 306 in this use case, that is separated into three categories (e.g., in the first category that corresponds to the firstthreshold 302, in the second category between the first threshold and the second threshold 306, and the third category which falls only in the third threshold 306).
[0049] FIG. 4 illustrates an example diagram 400 for determining pose prediction accuracy according to embodiments described herein. The diagram 400 depicts a process for determining pose prediction accuracy for a configuration of a system as that depicted in FIG. 1 that includes multiple image sensors as opposed to the configuration described below with reference to FIG. 5 which includes a single image sensor and a depth sensor. The diagram 400 includes a computer system receiving images from the image sensors (color or mono image sensors 402) placed at different viewpoints and generating an initial pose estimation for an object in a bin from one or more views (404 - 408) using a pose estimation algorithm. As depicted in FIG. 4, the computer system may generate a number of initial pose estimations for an object in a bin from multiple viewpoints (404 - 408) using color or mono color image sensors or cameras 402. As described above, the computer system implementing the pose prediction accuracy features may utilize an algorithm, such as a stereo matching algorithm or a multi -view stereo (MVS) algorithm 410 that uses the images from the image sensors 402 to generate a depth map 412 for the objects included in the images. In an embodiment, a point cloud may be obtained from an estimated depth map 412 and may be used for ICP registration refinement of the 6DoF pose of a detected object.
[0050] In embodiments, the diagram 400 includes part selection for grasping determination at 414. The part selection for grasping includes comparing, by the computer system, the depth map 412 with the depth information or depth generated by a real-time rendering application to determine a value that represents an accuracy of the depth included in the initial pose estimation (404 for example) for an object in the bin. In embodiments, the computer system may implement a real-time rendering engine (application) to render a depth map (depth information) for the object using the initial pose estimation (404) for the object and camera intrinsics for the image sensors 402. FIG. 4 depicts optional inclusion of different initial pose estimations from different views, 406 and 408, to combine into one initial pose estimation for an object or for use in generating the depth map (depth information) by the rendering application.
[0051] As described herein, the value may be compared to one or more thresholds to determine whether the depth for the object is accurate and therefore suitable for motion planning 416 of the robotic arm to interact with the object (e.g., grasp the object). In embodiments, the computer system may take other actions based on a comparison of the value to one or more thresholds. For example, the computer system may utilize one or morerefinement algorithms or processes, such as an ICP registration refinement 418, to further refine the initial pose estimation or move on to analyzing a different object in the bin all together to avoid spending more time on a particular object. In scenarios where the computer system determines that ICP registration refinement 418 should be utilized to update the initial pose estimation 404, information such as images and initial pose estimations for certain objects may be passed 420 to the ICP refinement 418 process implemented by the computer system. In scenarios where the depth for the object included in the initial pose estimation is determined to be accurate, in response to comparing the value to a threshold, the computer system may instruct the robotic arm to grasp the object using the initial estimated 6DoF for the object in the bin.
[0052] The computer system may instruct the robotic arm to grasp the object, move the object, turn the object, and transport the object to another destination such as a conveyer belt or to another machine for further processing such as assembly, blow off, wash, deburring, gauging, inspection, etc. In an embodiment, once an initial pose estimation has been refined via the ICP refinement 418, the computer system may generate new depth information (depth) or updated depth information for an object using the refined pose estimation and the rendering application. The new or updated depth information may be compared to the depth map 412 and a new value may be determined that represents an accuracy of the depth included in the refined pose estimation. The computer system may compare the new value to one or more thresholds and continue to motion planning 416 as described herein.
[0053] FIG. 5 illustrates an example diagram 500 for determining pose prediction accuracy according to embodiments described herein. The diagram 500 depicts a process for determining pose prediction accuracy for a configuration of a system that includes a depth sensor 502 and a single image sensor 504, which may be a color or mono image sensor, as opposed to the configuration described above with reference to FIG. 4 which includes multiple image sensors. The diagram 500 includes a computer system receiving an image from the image sensor 504 and generating an initial pose estimation 506 for an object in a bin using a pose estimation algorithm. As depicted in FIG. 5, the computer system may receive depth information for a plurality of objects captured by the depth sensor 502 and generate a depth map 508. In an embodiment, a point cloud may be obtained from an estimated depth map 508 and may be used for ICP registration refinement of the 6DoF pose of a detected object.
[0054] In embodiments, the diagram 500 includes part selection for grasping determination at 510. The part selection for grasping 510 includes comparing, by thecomputer system, the depth map 508 with the depth information or depth generated by a realtime rendering application to determine a value that represents an accuracy of the depth included in the initial pose estimation (506 for example) for an object in the bin. In embodiments, the computer system may implement a real-time rendering engine (application) to render depth information (depth) for the object using the initial pose estimation (506) for the object and camera intrinsics for the image sensor 504.
[0055] As described herein, the value may be compared to one or more thresholds to determine whether the depth for the object is accurate and therefore suitable for motion planning 512 of the robotic arm to interact with the object (e.g., grasp the object). In embodiments, the computer system may take other actions based on a comparison of the value to one or more thresholds. For example, the computer system may utilize one or more refinement algorithms or processes, such as an ICP registration refinement 514, to further refine the initial pose estimation 506 or move on to analyzing a different object in the bin all together to avoid spending more time on a particular object. In scenarios where the computer system determines that ICP registration refinement 514 should be utilized to update the initial pose estimation 506, information such as images, depth maps, and initial pose estimations for certain objects may be passed 516 to the ICP registration refinement 514 process implemented by the computer system. In scenarios where the depth for the object included in the initial pose estimation 506 is determined to be accurate, in response to comparing the value to a threshold, the computer system may instruct the robotic arm to grasp the object using the initial estimated 6DoF pose estimation 506 for the object in the bin.
[0056] The computer system may instruct the robotic arm to grasp the object, move the object, turn the object, and transport the object to another destination such as a conveyer belt or to another machine for further processing such as assembly, blow off, wash, deburring, gauging, inspection, etc. In an embodiment, once an initial pose estimation has been refined via the ICP registration refinement 514, the computer system may generate new rendered depth information (depth / depth map) or updated depth information for an object using the refined pose estimation and the rendering application. The new or updated depth information may be compared to the measured depth map 508 and a new value may be determined that represents an accuracy of the depth included in the refined pose estimation. The computer system may compare the new value to one or more thresholds and continue to motion planning 512 as described herein.
[0057] FIG. 6 illustrates a flow chart for determining pose prediction accuracy according to embodiments described herein. FIG. 6 includes exemplary process 600 which may beperformed by an environment or architecture such as depicted in FIGs. 1, 2, 4, 5, and 7 and by systems and components of FIGs. 1, 2, 4, 5, and 7. However, it will be recognized that any of the following blocks may be performed in any suitable order and that the process 600 may be performed in any environment or architecture and by any suitable computing device and / or controller of robotic arms, sensors, image capture sensors, and systems.
[0058] At step 602, the process 600 includes receiving images of a plurality of objects in a bin from multiple views captured by the two image sensors, the plurality of objects and the bin being a proximal distance from the two image sensors and the robotic arm. In embodiments, the proximal distance between the plurality of objects and the bin from the two image sensors and the robotic arm may refer to a distance between the operating range of the robotic arm as well as the sensor range or communication range between the image sensors, robotic arm, the objects and the bin. This may be different depending on the type of image sensors as well as communication types used by the system. For example, a computer in communication with the image sensors and robotic arm may communicate with the image sensors and robotic arm via WIFI and thus the proximal range between the components may be limited by an optimal WIFI range of the components in the system.
[0059] The process 600 may include generating an initial pose estimation for an object of the plurality of objects using a pose estimation algorithm and an image of the images at 704. In embodiments, the pose estimation algorithm may use more than a single image to generate the initial pose estimation for the object. In embodiments, the process 600 may include generating a depth map for the object using a stereo matching algorithm and the images at 606. The process 600 may include rendering, via a rendering application, a depth map (depth) for the object using the initial pose estimation for the object and camera intrinsics of the two image sensor(s) at 608. The process 600 may include determining a value that represents an accuracy of the depth for the object in the bin based on comparing the depth map for the object and the rendered depth for the object at 610. The process 600 may include at 612 instructing the robot arm to grasp the object based on comparing the value with a first threshold. For example, the robot arm may be instructed to grasp the object based on the value being less than the first threshold. In embodiments, the initial pose estimation may be refined using a refinement algorithm such as multi-view triangulation. In embodiments, the robot arm may be instructed to suppress ICP registration refinement and grasping of a detected object when the depth difference exceeds a certain threshold.
[0060] In an embodiment, in order to implement multi-view triangulation (e.g., triangle refinement algorithm) to aggregate pose estimation from multiple viewpoints, a set of 3Dpoints are pre-defined or sampled from a CAD mesh model of each detected mechanical part (object) in the bin. The 3D points of the CAD mesh model are converted to 2-dimensional points in an image space. With the initial pose of the mechanical part in each view, 2- dimensional (2D) points (pixel coordinates) in an image space of the 3D points selected from the CAD model are computed via back re-projection. A multi-view triangulation is applied to each point in the set of sampled points. Using the multi-view triangulation on the 2D points, the 3D coordinates in the camera view of the selected points are computed. The refined pose of the target part is easily computed with the 3D coordinates in the camera view using a rigid body constraint. The camera view may include the view obtained from the captured images.
[0061] FIG. 7 illustrates a simplified block diagram of one or more devices or systems for determining pose prediction accuracy according to embodiments of the present disclosure. FIG. 7 is a block diagram of an exemplary system or device 700 including a computer, computer device, server, or controller for determining a pose prediction accuracy and communicating with a robotic arm for picking, moving, or otherwise interacting with objects in a bin as described herein with reference to FIGs. 1-6. The system 700 includes a processor 704, such as a central processing unit (CPU), and / or logic, that executes computer executable instructions for performing the functions, processes, and / or methods described herein. In some examples, the computer executable instructions are locally stored and accessed from a non-transitory computer readable medium, such as storage 710, which may be a hard drive or flash drive. Read Only Memory (ROM) 706 includes computer executable instructions for initializing the processor 704, while the random-access memory (RAM) 708 is the main memory for loading and processing instructions executed by the processor 704.
[0062] The network interface 712 may connect to a wired network or cellular network and to a local area network or wide area network or BUUETOOTH or other suitable communication methods such as those described herein with reference to the communication component or communication module. The system 700 may also include a bus 702 that connects the processor 704, ROM 706, RAM 708, storage 710, and / or the network interface 712. The components within the system 700 may use the bus 702 to communicate with each other. The components within the system 700 are merely exemplary and might not be inclusive of every component for embodiments described herein. For instance, in some examples, the system 700 might not include a network interface 712. In embodiments the system 700 may include one or more components for interacting with a machine or system such as image sensors (cameras), depth sensors, or robotic arms. The system 700 may implement a number of algorithms and / or applications described herein for determining poseprediction accuracy including a pose estimation algorithm, a rendering application, an ICP registration algorithm, multi-view triangulation refinement, multi-view matching, block matching, semi-global matching, artificial intelligence deep learning based stereo matching, or artificial neural network trained algorithms trained with images of objects for generating or refining an initial pose estimation for an object.
[0063] While the disclosure has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive. It will be understood that changes and modifications may be made by those of ordinary skill within the scope of the following claims. In particular, the present disclosure covers further embodiments with any combination of features from different embodiments described above and below. Additionally, statements made herein characterizing the disclosure refer to an embodiment of the disclosure and not necessarily all embodiments.
[0064] The terms used in the claims should be construed to have the broadest reasonable interpretation consistent with the foregoing description. For example, the use of the article “a” or “the” in introducing an element should not be interpreted as being exclusive of a plurality of elements. Likewise, the recitation of “or” should be interpreted as being inclusive, such that the recitation of “A or B” is not exclusive of “A and B,” unless it is clear from the context or the foregoing description that only one of A and B is intended. Further, the recitation of “at least one of A, B and C” should be interpreted as one or more of a group of elements consisting of A, B and C, and should not be interpreted as requiring at least one of each of the listed elements A, B and C, regardless of whether A, B and C are related as categories or otherwise. Moreover, the recitation of “A, B and / or C” or “at least one of A, B or C” should be interpreted as including any singular entity from the listed elements, e.g., A, any subset from the listed elements, e.g., A and B, or the entire list of elements A, B and C.
Claims
CLAIMSWhat is claimed is:
1. A system comprising: an image sensor; a depth sensor; a robotic arm; a processor; and a memory including instructions that, when executed with the processor, cause the system to, at least: receive an image of a plurality of objects in a bin captured by the image sensor, the plurality of objects and the bin being a proximal distance from the image sensor, the depth sensor, and the robotic arm; receive depth information for the plurality of objects captured by the depth sensor; generate an initial pose estimation for an object of the plurality of objects using a pose estimation algorithm and the image; generate a depth map for the object using the depth information; render a depth for the object using the initial pose estimation for the object and camera intrinsics of the image sensor, wherein the camera intrinsics indicate imaging parameters that are used to capture the image; determine a value that represents an accuracy of the depth for the object in the bin based on comparing the depth map for the object and the rendered depth for the object; and instruct the robotic arm to grasp the object based on comparing the value with a first threshold.
2. The system according to claim 1, wherein the instructions are configured to implement an iterative closest point (ICP) registration algorithm to generate an updated pose estimation for the object based on the value being between the first threshold and a second threshold.
3. The system according to claim 2, wherein the instructions are configured to: render an updated depth for the object using the updated pose estimation for the object and the camera intrinsics; and determine an updated value that represents the accuracy of the depth for the object in the bin based on comparing the depth map for the object and the updated rendered depth for the object.
4. The system according to claim 3, wherein the instructions are configured to suppress transmission of instructions to the robotic arm to grasp the object based on comparing the value or the updated value with the second threshold or a third threshold.
5. The system according to claim 1, wherein the instructing the robotic arm to grasp the object is further based on the initial pose estimation.
6. The system according to claim 1, wherein the camera intrinsics include a focal length and camera principal point.
7. The system according to claim 1, wherein the image is a color image or a mono color image.
8. The system according to claim 1, wherein the first threshold is specified by a user through a configuration file or a graphical user interface.
9. The system according to claim 1, wherein the first threshold is based at least in part on physical properties of the object or a grasping tolerance of the robotic arm.
10. The system according to claim 1, wherein generating the initial pose estimation for the object includes using an artificial neural network trained with images of objects.
11. A system comprising: two image sensors; a robotic arm; a processor; and a memory including instructions that, when executed with the processor, cause the system to, at least:receive images of a plurality of objects in a bin from multiple views captured by the two image sensors, the plurality of objects and the bin being a proximal distance from the two image sensors and the robotic arm; generate an initial pose estimation for an object of the plurality of objects using a pose estimation algorithm and an image of the images; generate a depth map for the object using a stereo matching algorithm and the images; render a depth for the object using the initial pose estimation for the object and camera intrinsics of the two image sensors, wherein the camera intrinsics indicate imaging parameters that are used to capture the image; determine a value that represents an accuracy of the depth for the object in the bin based on comparing the depth map for the object and the rendered depth for the object; and instruct the robotic arm to grasp the object based on comparing the value with a first threshold.
12. The system according to claim 11, wherein the instructions are configured to use a triangulation technique to refine the initial pose estimation.
13. The system according to claim 11, wherein generating the initial pose estimation for the object includes using the pose estimation algorithm and the images.
14. The system according to claim 11, wherein the depth map for the object is generated using block matching, semi-global matching, an artificial intelligence deep learning based stereo matching algorithm, or multi-view matching algorithm.
15. The system according to claim 11, wherein the instructions are configured to implement an iterative closest point (ICP) registration algorithm to generate an updated pose estimation for the object based on the value being between the first threshold and a second threshold.
16. The system according to claim 11, wherein the instructing the robotic arm to grasp the object is further based on the initial pose estimation.
17. A computer-implemented method for determining pose prediction accuracy comprising:receiving, from two image sensors, images of a plurality of objects in a bin from multiple views that are captured by the two image sensors; generating an initial pose estimation for an object of the plurality of objects using a pose estimation algorithm and an image of the images; generating a depth map for the object using a stereo matching algorithm and the images; rendering a depth for the object using the initial pose estimation for the object and camera intrinsics of the two image sensors, wherein the camera intrinsics indicate imaging parameters that are used to capture the image; determining a value that represents an accuracy of the depth for the object in the bin based on comparing the depth map for the object and the rendered depth for the object; and instructing a robotic arm to grasp the object based on comparing the value with a first threshold, the robotic arm being a proximal distance from the two image sensors, the plurality of objects, and the bin.
18. The computer-implemented method according to claim 17, further comprising implementing an iterative closest point (ICP) registration algorithm to generate an updated pose estimation for the object based on the value being between the first threshold and a second threshold.
19. The computer-implemented method according to claim 17, further comprising suppressing grasping of the object based on the value being higher than a second threshold and a number of objects in the bin.
20. The computer-implemented method according to claim 17, wherein instructing the robotic arm to grasp the object is further based on the initial pose estimation.
Citation Information
Patent Citations
Image indexing and retrieval using local image patches for object three-dimensional pose estimation
US20200013189A1
Systems and methods for generating and using visual datasets for training computer vision models
US20220414928A1
In-hand pose refinement for pick and place automation
US20230071384A1
Controlling a robotic manipulator for packing an object
WO2023187006A1