Autoencoder Latent Variable Augmentation for Image Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current machine learning systems for image classification and semantic segmentation lack robustness and efficiency in generating diverse training datasets, particularly in varying specific features like pedestrian attributes in street scenes, which limits their effectiveness and coverage.
Innovation Solution
A method utilizing a decoder and encoder neural network system to generate augmented datasets by analyzing latent variables, allowing selective variation of image features through weighted differences in average latent values, ensuring a robust training process and improved convergence.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional generative methods are used for data augmentation, then training datasets can be generated, but the datasets lack diversity and coverage in specific features
Solution Approach 1:
The patent segments the feature space by identifying and isolating specific attributes (e.g., pedestrian attributes) from the overall image data. By separating features into discrete controllable dimensions, the system can selectively augment specific features while maintaining others, thereby improving feature coverage without compromising robustness.
Solution Approach 2:
The patent employs parameter changes by systematically varying specific latent variables corresponding to targeted features (such as pedestrian attributes) while keeping other parameters constant. This selective parameter modification enables precise control over which features are augmented, enhancing dataset diversity in specific dimensions while preserving overall data quality and robustness.
2Productivity
If machine learning systems are trained with limited diverse data, then training time is reduced, but classification and segmentation effectiveness deteriorates
Solution Approach 1:
The patent applies preliminary action by pre-processing images to identify and extract specific features of interest before the main training process. By preparing feature-augmented datasets in advance with targeted diversity, the system ensures that the training phase can proceed efficiently with pre-curated diverse data, thereby maintaining both training efficiency and classification accuracy.
3Quantity of substance
If conventional data augmentation is applied, then dataset size increases, but selective feature variation is not achieved
Solution Approach 1:
The patent implements local quality by applying different augmentation strategies to different features within the same dataset. Instead of uniformly transforming all images, the system selectively varies specific features (such as pedestrian attributes) while preserving other characteristics, thereby achieving both increased dataset size and feature-selective diversity.
Data Source
AI summary
A method for training a machine learning system. The method includes generating an augmented dataset including input images for training the machine learning system, which is for classification and/or semantic segmentation of input images, using a first machine learning system, which is embodied as a decoder of an autoencoder, and a second machine learning system, which is embodied as an encoder of the autoencoder. Latent variables are ascertained from the input images using the encoder. The input images are classified as a function of ascertained feature characteristics of their image data. An augmented input image of the augmented dataset is ascertained from at least one of the input images as a function of average values of the ascertained latent variables in at least two of the classes. The image classes are selected so that the input images classified therein agree in their characteristics in a predefinable set of other features.


