Virtual Camera Training Data for Occluded Robot Picking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Robotic arms face challenges in successfully picking up objects that are occluded, leading to issues such as object jamming and damage during the picking process.

Innovation Solution

A training data generation device and method that utilizes virtual scene generation, orthographic and perspective virtual cameras to capture and label occluded states of objects, generating training data that accurately reflects the actual occluded states, enabling a learning model to determine and control the robotic arm's actions for successful object pickup.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a robotic arm picks up an occluded object using conventional methods, then the picking process can be completed, but object jamming and damage occur due to inaccurate occlusion detection

Engineering Contradiction:
Improvepicking success rateVSAvoidobject jamming and damage
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent transitions from 2D image data to 3D point cloud data to represent object spatial relationships. By constructing a three-dimensional point cloud model from multi-view images and using depth information, the system accurately determines occlusion states and spatial positions of objects, enabling the robotic arm to successfully pick up occluded objects without jamming or damage.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent creates a virtual copy of the physical scene by constructing a three-dimensional point cloud model that replicates the spatial relationships and occlusion states of objects. This virtual model serves as a training dataset that teaches the robotic arm how to identify and pick up occluded objects without physically attempting unsuccessful picks that would cause jamming or damage.

Inventive Principle:
Principle #26Copying

2Measurement precision

If vertical projection images are used to determine occlusion state, then accurate occlusion labeling is achieved, but the system complexity increases due to multiple virtual cameras

Engineering Contradiction:
Improveocclusion state detection accuracyVSAvoidvirtual camera system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces a point cloud construction module as an intermediary that processes images from multiple virtual cameras (front, rear, left, right, top, bottom views) and synthesizes them into a three-dimensional point cloud model. This intermediary structure integrates the information from multiple views and automatically determines occlusion states, reducing the complexity of manually processing multiple images while maintaining high detection accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12409557B2Training data generation device, training data generation method using the same and robot arm system using the same
Publication Date: 2025.09.09 IND TECH RES INST
  • US12409557B2 patent drawing
  • US12409557B2 patent drawing
  • US12409557B2 patent drawing

AI summary

A training data generation device includes a virtual scene generation unit, an orthographic virtual camera, an object-occlusion determination unit, an object-occlusion determination unit and a perspective virtual camera. The virtual scene generation unit is configured for generating a virtual scene, wherein the virtual scene comprises a plurality of objects. The orthographic virtual camera is configured for capturing a vertical projection image of the virtual scene. The object-occlusion determination unit is configured for labeling an occluded-state of each object according to the vertical projection image. The perspective virtual camera is configured for capturing a perspective projection image of the virtual scene. The training data generation unit is configured for generating a training data of the virtual scene according to the perspective projection image and the occluded-state of each object.