Descriptor Image Training for Multi-Object Self-Supervised Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning models, such as dense object nets, are limited in handling multiple objects in a scene, requiring extensive labeled data and object masks, which hinders their application in practical scenarios like robotic manipulation.
Innovation Solution
A method for training machine learning models that generates descriptor images by forming pairs of images from different perspectives, using image augmentation techniques like resizing, distortion, and noise addition, to reduce the loss function and enhance data efficiency, allowing self-supervised learning without object masks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing machine learning models like dense object nets are used for isolated objects, then training can be performed with self-supervised learning, but the models fail to handle multiple objects in practical scenarios
Solution Approach 1:
The patent segments the training process into distinct phases: pre-training on isolated objects to learn fundamental descriptors, then fine-tuning on multiple objects to adapt to complex scenarios. This segmentation allows the model to progressively build capability without being overwhelmed by the complexity of multiple objects from the start.
Solution Approach 2:
The patent applies preliminary action by pre-training the model on isolated objects before exposing it to multiple objects. This preliminary training establishes a solid foundation of object descriptors that can later be adapted to more complex scenes, improving both versatility and reliability systematically.
2Measurement precision
If extensive labeled data and object masks are used for training, then model accuracy improves, but data preparation complexity and time increase significantly
Solution Approach 1:
The patent implements self-service through self-supervised learning where the model learns to generate its own training signals from unlabelled data. By using consistency regularization across augmented views, the model creates its own supervision without requiring manual labeling or object masks, thereby maintaining accuracy while eliminating time-consuming data preparation.
Solution Approach 2:
The patent changes the parameter of supervision from labeled to unlabelled data by introducing consistency regularization. This parameter change allows the model to learn effective descriptors without extensive labeled data, significantly reducing data preparation time while maintaining or improving descriptor accuracy through the self-supervised objective.
3Manufacturing precision
If object masks are required for training, then training precision improves, but the ease of operation and automation decrease
Solution Approach 1:
The model performs self-service by automatically learning object boundaries and descriptors without requiring external object masks. The consistency regularization mechanism enables the model to self-supervise on unlabelled data, achieving high training precision while fully automating the training process without manual mask creation.
Solution Approach 2:
The patent extracts the requirement for object masks by formulating a self-supervised learning objective that derives all necessary supervision signals from the images themselves. This extraction eliminates the need for separate mask data while maintaining training precision through the consistency-based learning approach.
4Ease of manufacture
If models are trained only on isolated objects, then training data collection is simpler, but the models cannot handle objects in complex scenes with multiple items
Solution Approach 1:
The patent applies preliminary action by first collecting simple data of isolated objects, then progressively introducing complex multi-object scenes in a staged fine-tuning process. This preliminary simplicity in data collection is maintained while ultimately achieving versatility through the two-stage training approach.
Solution Approach 2:
The patent introduces dynamics by transitioning from static isolated-object training to dynamic multi-object scene training. The model adapts its learning behavior through consistency regularization that works effectively across different scene complexities, enabling the system to handle varying degrees of scene complexity without retraining from scratch.
Data Source
AI summary
A method for training a machine learning model for generating descriptor images for images of one or of multiple objects. The method includes: formation of pairs of images which show the one or the multiple objects from different perspectives; generation, for each image pair, using the machine learning model, of a first descriptor image for the first image, and of a second descriptor image for the second image, which assigns descriptors to points of the one or multiple objects shown in the second image; sampling, for each image pair, of descriptor pairs, which include in each case a first descriptor from the first descriptor image and a second descriptor from the second descriptor image, which are assigned to the same point, and the adaptation of the machine learning method for reducing a loss.


