Robotic Grasp Point Detection in Occluded Dense Object Scenes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer vision systems for robotic grasping lack accuracy and efficiency, particularly in dense object scenes with occlusions, and require extensive real-world training data and processing resources.
Innovation Solution
A method that involves labeling images based on attempted object grasps, generating a trained graspability network using synthetic and real-world data, determining grasp points, and executing object grasps, leveraging deep learning algorithms for rapid and accurate target selection, and utilizing a system comprising an end effector, robotic arm, sensor suite, and computing system to prioritize grasp points based on graspability scores and object poses.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional computer vision systems are used for robotic grasping, then the system can operate with simple hardware, but the grasping accuracy and efficiency deteriorate significantly in dense object scenes with occlusions
Solution Approach 1:
The patent uses synthetic data that copies the visual characteristics of real-world objects and scenes to train the neural network. By generating virtual images that replicate real object appearances, textures, and occlusion patterns, the system achieves high grasping accuracy without requiring complex hardware sensors beyond standard cameras.
Solution Approach 2:
The patent replaces traditional mechanical vision systems with a neural network-based computational approach. Instead of relying on complex mechanical sensors or multiple cameras, the system uses a single camera combined with deep learning algorithms to achieve superior grasping performance in dense scenes.
2Measurement precision
If extensive real-world training data is collected, then the grasping system can achieve high accuracy, but the time and computational resources required increase significantly
Solution Approach 1:
The patent creates synthetic training data that copies real-world object characteristics, including diverse object shapes, textures, lighting conditions, and occlusion patterns. This virtual data generation approach provides extensive training material without requiring physical data collection, reducing training time while maintaining high accuracy.
Solution Approach 2:
The patent performs preliminary training using synthetic data before deploying the model for real-world applications. This pre-training phase establishes the foundational grasping capabilities efficiently, and the model can then adapt to real-world variations with minimal additional training, significantly reducing total training time and computational resources.
3Loss of information
If explicit object detection is performed in dense scenes, then object identification can be achieved, but the processing time and computational load increase
Solution Approach 1:
The patent merges object detection and grasping prediction into a single unified neural network architecture. Instead of performing separate detection and grasping tasks, the model simultaneously identifies objects and predicts grasp points in one computational pass, significantly improving processing speed while maintaining identification accuracy in dense scenes.
Solution Approach 2:
The patent uses synthetic data that copies complex scene structures and occlusion patterns to train the unified network. By exposing the model to virtual dense scenes during training, it learns to efficiently identify objects and determine grasp points without requiring time-consuming explicit detection algorithms, achieving both accuracy and speed.
4Speed
If deep learning algorithms are used for rapid target selection, then grasping speed improves, but the computational resources and processing power requirements increase
Solution Approach 1:
The patent trains deep learning models using synthetic data that copies real-world object characteristics. This approach enables the models to achieve high grasping speed and accuracy without requiring extensive computational resources during inference, as the synthetic training data pre-conditioning the models to handle real-world variations efficiently with lower computational overhead.
Data Source
AI summary
The method for increasing the accuracy of grasping an object can include: labelling an image based on an attempted object grasp by a robot and generating a trained graspability network using the labelled images. The method can additionally or alternatively include determining a grasp point using the trained graspability network; executing an object grasp at the grasp point S400; and/or any other suitable elements.


