Descriptor Image Training Using Geometric Augmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for training dense object nets require elaborate arrangements with movable cameras to capture images from different perspectives, making the data collection process time-consuming and complex.
Innovation Solution
A method that generates augmented versions of camera images using position changes and geometric transformations, allowing a static camera to record multiple views, simplifying data collection and training by using contrastive loss to determine corresponding pixels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple images are recorded from different perspectives using a movable camera arrangement, then the training data diversity is improved, but the device complexity and time consumption increase
Solution Approach 1:
The patent creates virtual copies of the original camera image through geometric transformations (rotations, reflections, scaling, shearing). These transformed images serve as synthetic training data without requiring physical movement of the camera or additional拍摄 sessions, thus maintaining data diversity while eliminating complex hardware arrangements
Solution Approach 2:
The patent applies parameter changes to the original image by transforming geometric parameters (rotation angles, scale factors, reflection axes, shearing parameters). These parameter transformations generate diverse training samples from a single source image, achieving adaptability without increasing device complexity
2Adaptability or versatility
If multiple images are recorded from different perspectives, then the training data diversity is improved, but the time consumption increases
Solution Approach 1:
The patent performs preliminary geometric transformations on the original image to pre-generate multiple training samples. This preliminary action creates a diverse training dataset from a single captured image, eliminating the need for time-consuming multi-perspective拍摄 sessions while ensuring sufficient training data diversity
Solution Approach 2:
The patent creates virtual copies of the original image through geometric transformations. These synthetic copies provide diverse training data without requiring additional physical拍摄 time, thus resolving the contradiction between data diversity and time consumption
3Device complexity
If a static camera is used to record images, then the device complexity is reduced, but the training data diversity deteriorates
Solution Approach 1:
The patent compensates for the static camera limitation by applying geometric parameter transformations (rotation, reflection, scaling, shearing) to the captured image. These parameter changes synthesize multiple viewing perspectives from a single static capture, maintaining data diversity without requiring complex movable camera arrangements
Solution Approach 2:
The patent creates virtual copies of the static camera image through geometric transformations. These transformed copies simulate multiple perspectives that would otherwise require physical camera movement, thus achieving data diversity with a simple static camera setup
Data Source
AI summary
A method for training a machine learning model for generating descriptor images for images of one or more objects. The method includes recording multiple camera images, each showing one or more objects, and, for each camera image, generating one or more augmented versions of the camera image by applying a respective augmentation to the camera image for each augmented version of the camera image, wherein the augmentation comprises a change in position of pixel values of the camera image, generating pairs of training images each including the camera image and an augmentation of the camera image or two augmented versions of the camera image; and training the machine learning model with contrastive loss using the pairs of training images.


