3D Voxel GUI Labeling for Accurate Robot Object Handling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for training neural networks to handle 3D objects in robot automation lack efficiency and user-friendliness in labeling training sets, particularly in handling the 3D surface and orientation of objects, and do not effectively utilize 3D image reconstruction for accurate robot commands.

Innovation Solution

A computer-implemented method using a GUI for manual annotation of a 3D voxel representation of training objects, combined with 2D and 3D neural networks for segmentation and reconstruction, to enhance the labeling process and improve the accuracy of robot commands.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation is performed using traditional 2D image views only, then the labeling process becomes complex and time-consuming, but the 3D surface structure and orientation information may be lost or difficult to capture accurately

Engineering Contradiction:
Improve3D surface structure accuracyVSAvoidlabeling time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent introduces a 3D voxel representation view in the GUI that allows annotators to label 3D objects from multiple perspectives including top, front, and side views. This dimensional enhancement enables accurate capture of 3D surface structures and orientations while maintaining efficient labeling workflow, directly resolving the contradiction between measurement precision and time loss.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If complex labeling procedures are used to capture 3D object features, then data quality improves, but ease of operation deteriorates

Engineering Contradiction:
Improvedata qualityVSAvoiduser-friendliness
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent segments the labeling interface into multiple orthogonal views (top, front, side) within the GUI, allowing annotators to work with simplified 2D representations for each view while collectively capturing complete 3D object features. This segmentation maintains data quality by ensuring all 3D characteristics are represented across views while significantly improving ease of operation compared to complex single-view 3D labeling.

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If traditional 2D image processing methods are used, then the system complexity remains low, but the ability to handle 3D objects and orientations is insufficient

Engineering Contradiction:
Improve3D object handling capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces voxel representations as an intermediary data structure between traditional 2D images and 3D object handling. The GUI displays and enables labeling of these voxel-based 3D models, which serve as a mediator that captures complete 3D surface information and orientations. This approach enhances 3D object handling capability while keeping system complexity manageable by using a standardized intermediate representation format.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250249580A1Improved method for labelling a training set involving a GUI
Publication Date: 2025.08.07 ROBOVISION
  • US20250249580A1 patent drawing
  • US20250249580A1 patent drawing
  • US20250249580A1 patent drawing

AI summary

The present invention relates to a computer-implemented method for labelling a training set, preferably for training a neural network, with respect to a 3D physical object by means of a GUI, the method comprising the steps of: obtaining a training set relating to a plurality of training objects, each of the training objects comprising a 3D surface similar to the 3D surface of said object, the training set comprising at least two images for each training object; generating, for each training object, a respective 3D voxel representation based on the respective at least two images; receiving, via said GUI, manual annotations with respect to a plurality of segment classes from a user of said GUI for labelling each of the training objects; and preferably training, based on said manual annotations, at least one NN, for obtaining said at least one trained NN.