Vision-Based 3D Object Handling via Voxel Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for handling 3D physical objects in robot automation lack effective 3D surface analysis, leading to inaccurate robot commands and inefficient object handling, as they primarily rely on 2D image processing and do not adequately account for the structure and features of 3D objects.
Innovation Solution
A method that involves obtaining images from multiple cameras positioned at different angles, generating a voxel representation of the 3D surface, and using trained neural networks for segmentation and measurement to compute precise robot commands for handling, incorporating techniques like voxel carving and semantic/instance segmentation to improve accuracy and robustness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If 2D image processing is used for object analysis, then processing complexity is reduced, but measurement precision and handling accuracy deteriorate
Solution Approach 1:
The patent transforms 2D image data into 3D voxel representations, adding the third dimension to capture depth and spatial structure. This enables the system to analyze the complete 3D surface geometry of objects, providing accurate measurements for handling while maintaining manageable processing complexity through structured 3D data representation.
2Measurement precision
If 3D surface analysis is implemented, then handling accuracy is improved, but device complexity increases
Solution Approach 1:
The patent segments the 3D object into volumetric pixels (voxels) that can be individually processed and analyzed. This segmentation approach simplifies complex 3D surface analysis by breaking down the object into manageable units, enabling accurate handling precision while reducing the computational complexity required for processing.
Solution Approach 2:
The patent creates a digital copy of the 3D object as a voxel model, which can be analyzed, measured, and manipulated without physically handling the actual object. This virtual representation enables accurate handling analysis while reducing the complexity of physical manipulation systems.
3Measurement precision
If multiple cameras are used for 3D imaging, then object analysis accuracy is improved, but device complexity increases
Solution Approach 1:
The patent merges images from multiple cameras into a unified 3D voxel representation. By combining the information from multiple camera views, the system achieves accurate 3D object analysis while managing device complexity through integrated processing of multiple image sources into a single coherent 3D model.
Data Source
AI summary
A method for generating a technical instruction for handling a 3D physical object present within a reference volume and comprising a 3D surface, the method comprising: obtaining at least two images of the object from a plurality of cameras positioned at different respective angles with respect to the object; generating, with respect to the 3D surface, a voxel representation segmented based on the at least two images, said segmenting comprising identifying a first segment component corresponding to a plurality of first voxels and a second segment component corresponding to a plurality of second voxels different from the plurality of first voxels; performing a measurement with respect to the plurality of first voxels; and computing the technical instruction for the handling of the object based on the segmented voxel representation and the measurement, wherein said segmenting relates to at least one trained NN being trained with respect to the 3D surface.


