Geographic Image Augmentation for Accurate Remote Sensing Labels
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The inefficiency and resource-intensiveness of acquiring large amounts of training data for machine learning models, particularly in remote sensing imaging, due to the labor-intensive labeling process and limitations of conventional data augmentation techniques, which often produce unsuitable augmentations.
Innovation Solution
A method of natural augmentation that involves ascertaining geographic location information to associate images with the same label, forming a training dataset from coregistered images, and using this dataset to train machine learning devices, enhancing model performance through improved alignment and alignment of imagery.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional data augmentation techniques (flips, rotations) are used to generate training data, then the quantity of training data increases, but the quality and suitability of the augmented data deteriorates
Solution Approach 1:
The patent copies geographic location information from labeled images to create corresponding training entries for additional images of the same location. Instead of transforming existing images through flips or rotations, the system retrieves and associates other images captured at the same geographic coordinates, preserving the authenticity and suitability of the training data while increasing its quantity.
2Reliability
If human labeling is performed on all training images to ensure accuracy, then the quality of labels improves, but the time and resource requirements increase
Solution Approach 1:
The system performs preliminary labeling by automatically transferring labels from one image to another based on geographic location matching. Images are pre-labeled using labels from geographically matched images, eliminating the need for manual human labeling of each individual image while maintaining label accuracy through the spatial relationship between images.
3Reliability
If more training data is acquired to improve model performance, then the model accuracy improves, but the resource intensity and cost increase
Solution Approach 1:
The system creates a multi-functional training dataset where a single labeled image serves multiple purposes by generating training entries for multiple geographically matched images. One labeled image can augment multiple training examples, maximizing the utility of each labeled image and reducing the overall resources needed to achieve sufficient training data volume for good model performance.
Data Source
AI summary
In some embodiments, a method for training a machine learning device includes: ascertaining geographic location information of at least one portion of a first image associated with a label; associating with the label a second image including at least a portion having substantially the same geographic location information as the at least one portion of the first image; optional alignment or coregistration of the first and second image to maximize mutual information overlap; forming a training dataset comprising the first and second images as input images and the label that the first and second images are associated with as outputs; optional binary categorization and curation of the resulting training dataset to ensure accuracy; and training the machine learning model using the augmented dataset.


