Descriptor Image Training for Multi-Object Self-Supervised Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning models, such as dense object nets, are limited in handling multiple objects in a scene, requiring extensive labeled data and object masks, which hinders their application in practical scenarios like robotic manipulation.

Innovation Solution

A method for training machine learning models that generates descriptor images by forming pairs of images from different perspectives, using image augmentation techniques like resizing, distortion, and noise addition, to reduce the loss function and enhance data efficiency, allowing self-supervised learning without object masks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing machine learning models like dense object nets are used for isolated objects, then training can be performed with self-supervised learning, but the models fail to handle multiple objects in practical scenarios

Engineering Contradiction:
Improveability to handle multiple objectsVSAvoidperformance in practical scenarios
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments the training process into distinct phases: pre-training on isolated objects to learn fundamental descriptors, then fine-tuning on multiple objects to adapt to complex scenarios. This segmentation allows the model to progressively build capability without being overwhelmed by the complexity of multiple objects from the start.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by pre-training the model on isolated objects before exposing it to multiple objects. This preliminary training establishes a solid foundation of object descriptors that can later be adapted to more complex scenes, improving both versatility and reliability systematically.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If extensive labeled data and object masks are used for training, then model accuracy improves, but data preparation complexity and time increase significantly

Engineering Contradiction:
Improvedescriptor accuracyVSAvoiddata preparation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements self-service through self-supervised learning where the model learns to generate its own training signals from unlabelled data. By using consistency regularization across augmented views, the model creates its own supervision without requiring manual labeling or object masks, thereby maintaining accuracy while eliminating time-consuming data preparation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the parameter of supervision from labeled to unlabelled data by introducing consistency regularization. This parameter change allows the model to learn effective descriptors without extensive labeled data, significantly reducing data preparation time while maintaining or improving descriptor accuracy through the self-supervised objective.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If object masks are required for training, then training precision improves, but the ease of operation and automation decrease

Engineering Contradiction:
Improvetraining precisionVSAvoidautomatic training capability
Core Design Contradiction:
Manufacturing precisionVSExtent of automation

Solution Approach 1:

The model performs self-service by automatically learning object boundaries and descriptors without requiring external object masks. The consistency regularization mechanism enables the model to self-supervise on unlabelled data, achieving high training precision while fully automating the training process without manual mask creation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent extracts the requirement for object masks by formulating a self-supervised learning objective that derives all necessary supervision signals from the images themselves. This extraction eliminates the need for separate mask data while maintaining training precision through the consistency-based learning approach.

Inventive Principle:
Principle #2Taking out (Extraction)

4Ease of manufacture

If models are trained only on isolated objects, then training data collection is simpler, but the models cannot handle objects in complex scenes with multiple items

Engineering Contradiction:
Improvetraining data collection easeVSAvoidhandling of complex scenes
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent applies preliminary action by first collecting simple data of isolated objects, then progressively introducing complex multi-object scenes in a staged fine-tuning process. This preliminary simplicity in data collection is maintained while ultimately achieving versatility through the two-stage training approach.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces dynamics by transitioning from static isolated-object training to dynamic multi-object scene training. The model adapts its learning behavior through consistency regularization that works effectively across different scene complexities, enabling the system to handle varying degrees of scene complexity without retraining from scratch.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20230150142A1Device and method for training a machine learning model for generating descriptor images for images of objects
Publication Date: 2023.05.18 ROBERT BOSCH GMBH
  • US20230150142A1 patent drawing
  • US20230150142A1 patent drawing
  • US20230150142A1 patent drawing

AI summary

A method for training a machine learning model for generating descriptor images for images of one or of multiple objects. The method includes: formation of pairs of images which show the one or the multiple objects from different perspectives; generation, for each image pair, using the machine learning model, of a first descriptor image for the first image, and of a second descriptor image for the second image, which assigns descriptors to points of the one or multiple objects shown in the second image; sampling, for each image pair, of descriptor pairs, which include in each case a first descriptor from the first descriptor image and a second descriptor from the second descriptor image, which are assigned to the same point, and the adaptation of the machine learning method for reducing a loss.