3D Neural Network Training for Robust Robotic Object Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems for training robotic devices to perform object detection are limited by variations in deformation, object articulation, viewing angles, and lighting, which prevent effective detection in real-world environments.
Innovation Solution
A method involving a 3D camera to capture images, generate manipulated images by adjusting parameters, and process pairs of images to create reference images with embedded descriptors, enabling the robotic device to identify objects across varying conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If real world training images are collected under actual conditions, then the training data reflects real environment variations, but the system cannot generalize to unseen variations in deformation, articulation, viewing angles, and lighting
Solution Approach 1:
The system performs preliminary actions by collecting training images under multiple controlled conditions (different lighting, angles, deformations) before deployment. This advance preparation creates a comprehensive training dataset that enables the neural network to handle real-world variations without requiring adaptive learning in the field.
Solution Approach 2:
The patent systematically varies key parameters including lighting conditions, viewing angles, object deformations, and articulations during training image collection. By deliberately changing these parameters across multiple training sessions, the system learns to recognize objects under diverse conditions, resolving the contradiction between reliability and adaptability.
2Ease of manufacture
If training is limited to actual conditions used to collect training images, then the training process is simple and controlled, but the system fails to account for variations in the environment during real world operation
Solution Approach 1:
The patent segments the training process into multiple controlled sessions, each focusing on specific environmental variations (lighting, angles, deformations). This segmentation allows systematic exploration of parameter space while maintaining experimental control, balancing simplicity with comprehensive training.
Solution Approach 2:
The training system is designed to serve multiple functions: it can train under controlled conditions, systematically vary parameters, and generate comprehensive training datasets. This multi-functionality allows the same system to maintain simplicity while producing robust models that generalize to real-world variations.
3Stability of the object's composition
If conventional systems use fixed training conditions, then the training data is consistent and easy to manage, but variations in deformation, object articulation, viewing angle, and lighting prevent effective object detection
Solution Approach 1:
The patent introduces dynamics into the training process by systematically varying parameters such as lighting conditions, viewing angles, object deformations, and articulations. This dynamic approach creates training data that reflects real-world variability while maintaining systematic control, enabling the model to generalize across conditions.
Solution Approach 2:
The training dataset is constructed as a composite of images collected under multiple different conditions (various lighting, angles, deformations). By combining these diverse yet systematically controlled images into a unified training set, the system achieves both consistency in methodology and diversity in content, resolving the contradiction between stability and adaptability.
Data Source
AI summary
A method for training a deep neural network of a robotic device is described. The method includes constructing a 3D model using images captured via a 3D camera of the robotic device in a training environment. The method also includes generating pairs of 3D images from the 3D model by artificially adjusting parameters of the training environment to form manipulated images using the deep neural network. The method further includes processing the pairs of 3D images to form a reference image including embedded descriptors of common objects between the pairs of 3D images. The method also includes using the reference image from training of the neural network to determine correlations to identify detected objects in future images.


