Virtual Camera Training Data for Occluded Robot Picking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Robotic arms face challenges in successfully picking up objects that are occluded, leading to issues such as object jamming and damage during the picking process.
Innovation Solution
A training data generation device and method that utilizes virtual scene generation, orthographic and perspective virtual cameras to capture and label occluded states of objects, generating training data that accurately reflects the actual occluded states, enabling a learning model to determine and control the robotic arm's actions for successful object pickup.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a robotic arm picks up an occluded object using conventional methods, then the picking process can be completed, but object jamming and damage occur due to inaccurate occlusion detection
Solution Approach 1:
The patent transitions from 2D image data to 3D point cloud data to represent object spatial relationships. By constructing a three-dimensional point cloud model from multi-view images and using depth information, the system accurately determines occlusion states and spatial positions of objects, enabling the robotic arm to successfully pick up occluded objects without jamming or damage.
Solution Approach 2:
The patent creates a virtual copy of the physical scene by constructing a three-dimensional point cloud model that replicates the spatial relationships and occlusion states of objects. This virtual model serves as a training dataset that teaches the robotic arm how to identify and pick up occluded objects without physically attempting unsuccessful picks that would cause jamming or damage.
2Measurement precision
If vertical projection images are used to determine occlusion state, then accurate occlusion labeling is achieved, but the system complexity increases due to multiple virtual cameras
Solution Approach 1:
The patent introduces a point cloud construction module as an intermediary that processes images from multiple virtual cameras (front, rear, left, right, top, bottom views) and synthesizes them into a three-dimensional point cloud model. This intermediary structure integrates the information from multiple views and automatically determines occlusion states, reducing the complexity of manually processing multiple images while maintaining high detection accuracy.
Data Source
AI summary
A training data generation device includes a virtual scene generation unit, an orthographic virtual camera, an object-occlusion determination unit, an object-occlusion determination unit and a perspective virtual camera. The virtual scene generation unit is configured for generating a virtual scene, wherein the virtual scene comprises a plurality of objects. The orthographic virtual camera is configured for capturing a vertical projection image of the virtual scene. The object-occlusion determination unit is configured for labeling an occluded-state of each object according to the vertical projection image. The perspective virtual camera is configured for capturing a perspective projection image of the virtual scene. The training data generation unit is configured for generating a training data of the virtual scene according to the perspective projection image and the occluded-state of each object.


