3D Segmentation-Based Robot Bin Picking Without 6-DoF Pose Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current robotic systems struggle to pick objects from disorganized bins due to their inability to generalize to different objects and environments, especially in irregular and unknown conditions, as they rely on 6-DoF pose estimation which is slow and inaccurate for flexible or deformable objects, and lacks effectiveness in cluttered and irregular arrangements.

Innovation Solution

A method and system that uses 3-D geometry and segmentation from images to compute instance segmentation masks and pickability scores, allowing a robotic arm to select and pick objects without estimating 6-DoF poses, by capturing images, computing depth maps, and segmenting point clouds to determine the best grasping points and approach direction based on object protrusion, clutter, and distance from the end effector.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If 6-DoF pose estimation is used to guide bin picking, then the robotic system can identify object locations and orientations, but the system becomes slow and inaccurate for flexible or deformable objects and struggles with cluttered irregular arrangements

Engineering Contradiction:
Improvepose estimation accuracyVSAvoidpicking speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent extracts only the essential information needed for picking (instance segmentation masks identifying object boundaries and pickability scores indicating grasp quality) rather than computing full 6-DoF poses. This selective extraction of critical data eliminates unnecessary computational steps while maintaining picking effectiveness.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces the traditional mechanics-based 6-DoF pose estimation pipeline with a direct vision-to-grasp approach using instance segmentation and pickability scoring. This substitution eliminates complex mathematical transformations and directly maps visual features to grasping decisions, significantly reducing computation time.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If traditional vision systems are used to capture and analyze images for pose estimation, then object location can be determined, but the system fails to generalize to unknown objects and irregular environments

Engineering Contradiction:
Improvegeneralization capabilityVSAvoidenvironment adaptability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent employs instance segmentation masks that serve multiple functions: identifying object boundaries, separating individual objects from clutter, and providing spatial information for grasp planning. This multi-functional approach enables the system to handle diverse objects and environments with a single unified method.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent changes the fundamental parameters used for object identification from rigid 6-DoF pose estimates to flexible instance segmentation masks and pickability scores. This parameter transformation allows the system to adapt to unknown objects and irregular arrangements by focusing on boundary detection and grasp quality rather than precise geometric modeling.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If complex pose estimation algorithms are implemented to handle cluttered bins, then measurement accuracy improves, but processing time increases and system complexity grows

Engineering Contradiction:
Improveobject detection accuracyVSAvoidprocessing latency
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies instance segmentation to divide the cluttered bin scene into individual object instances, each with its own boundary mask. This segmentation approach directly addresses the complexity of cluttered environments by isolating objects of interest without requiring full scene understanding or complex pose estimation for each object.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs partial action by computing only the specific information needed for successful picking (instance masks and pickability scores) rather than complete 6-DoF poses for all objects. This selective computation reduces processing time while maintaining sufficient accuracy for the picking task.

Inventive Principle:
Principle #16Partial or excessive action

4Manufacturing precision

If the robotic system uses detailed 6-DoF pose information, then grasping precision can be improved, but the system complexity and computational requirements increase significantly

Engineering Contradiction:
Improvegrasping precisionVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent applies local quality by computing pickability scores that specifically evaluate grasp quality at potential contact points on each object, rather than requiring global 6-DoF pose information. This localized assessment focuses computational resources on the specific regions relevant for grasping, reducing overall system complexity while maintaining grasping precision.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12172310B2Systems and methods for picking objects using 3-D geometry and segmentation
Publication Date: 2024.12.24 INTRINSIC INNOVATION LLC
  • US12172310B2 patent drawing
  • US12172310B2 patent drawing
  • US12172310B2 patent drawing

AI summary

A method for controlling a robotic system includes: capturing, by an imaging system, one or more images of a scene; computing, by a processing circuit including a processor and memory, one or more instance segmentation masks based on the one or more images, the one or more instance segmentation masks detecting one or more objects in the scene; computing, by the processing circuit, one or more pickability scores for the one or more objects; selecting, by the processing circuit, an object among the one or more objects based on the one or more pickability scores; computing, by the processing circuit, an object picking plan for the selected object; and outputting, by the processing circuit, the object picking plan to a controller configured to control an end effector of a robotic arm to pick the selected object.