Descriptor Image Training Using Geometric Augmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for training dense object nets require elaborate arrangements with movable cameras to capture images from different perspectives, making the data collection process time-consuming and complex.

Innovation Solution

A method that generates augmented versions of camera images using position changes and geometric transformations, allowing a static camera to record multiple views, simplifying data collection and training by using contrastive loss to determine corresponding pixels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple images are recorded from different perspectives using a movable camera arrangement, then the training data diversity is improved, but the device complexity and time consumption increase

Engineering Contradiction:
Improvetraining data diversityVSAvoidcamera arrangement complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates virtual copies of the original camera image through geometric transformations (rotations, reflections, scaling, shearing). These transformed images serve as synthetic training data without requiring physical movement of the camera or additional拍摄 sessions, thus maintaining data diversity while eliminating complex hardware arrangements

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent applies parameter changes to the original image by transforming geometric parameters (rotation angles, scale factors, reflection axes, shearing parameters). These parameter transformations generate diverse training samples from a single source image, achieving adaptability without increasing device complexity

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If multiple images are recorded from different perspectives, then the training data diversity is improved, but the time consumption increases

Engineering Contradiction:
Improvetraining data diversityVSAvoiddata collection time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent performs preliminary geometric transformations on the original image to pre-generate multiple training samples. This preliminary action creates a diverse training dataset from a single captured image, eliminating the need for time-consuming multi-perspective拍摄 sessions while ensuring sufficient training data diversity

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates virtual copies of the original image through geometric transformations. These synthetic copies provide diverse training data without requiring additional physical拍摄 time, thus resolving the contradiction between data diversity and time consumption

Inventive Principle:
Principle #26Copying

3Device complexity

If a static camera is used to record images, then the device complexity is reduced, but the training data diversity deteriorates

Engineering Contradiction:
Improvecamera arrangement complexityVSAvoidtraining data diversity
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent compensates for the static camera limitation by applying geometric parameter transformations (rotation, reflection, scaling, shearing) to the captured image. These parameter changes synthesize multiple viewing perspectives from a single static capture, maintaining data diversity without requiring complex movable camera arrangements

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates virtual copies of the static camera image through geometric transformations. These transformed copies simulate multiple perspectives that would otherwise require physical camera movement, thus achieving data diversity with a simple static camera setup

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20230267724A1Device and method for training a machine learning model for generating descriptor images for images of objects
Publication Date: 2023.08.24 ROBERT BOSCH GMBH
  • US20230267724A1 patent drawing
  • US20230267724A1 patent drawing
  • US20230267724A1 patent drawing

AI summary

A method for training a machine learning model for generating descriptor images for images of one or more objects. The method includes recording multiple camera images, each showing one or more objects, and, for each camera image, generating one or more augmented versions of the camera image by applying a respective augmentation to the camera image for each augmented version of the camera image, wherein the augmentation comprises a change in position of pixel values of the camera image, generating pairs of training images each including the camera image and an augmentation of the camera image or two augmented versions of the camera image; and training the machine learning model with contrastive loss using the pairs of training images.