Hyperspectral Data Augmentation via Multi-Algorithm Pixel Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data augmentation techniques for hyperspectral data face challenges such as the lack of training data, high costs, and potential bias in machine learning models, leading to inaccurate classification results, especially in material or vegetation identification, where expert labeling is required.
Innovation Solution
A method involving the use of multiple augmentation algorithms to generate diverse training data for hyperspectral images, where each pixel is augmented using different techniques such as PCA and multivariate noise methods, and stochastic random averaging, to create a comprehensive set of training data that reduces overfitting and enhances classification accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If data augmentation techniques are used to increase training data volume, then the generalization ability of machine learning models is improved, but the risk of introducing bias and generating inaccurate results increases
Solution Approach 1:
The patent applies multiple different augmentation algorithms (PCA, multivariate noise, stochastic random averaging) that transform the data through different parameter changes. Each algorithm modifies the hyperspectral data in distinct ways, creating diverse augmented samples that improve generalization while the variety of transformation parameters helps prevent systematic bias introduction.
Solution Approach 2:
The patent combines multiple augmentation algorithms to create a composite augmentation approach. By integrating PCA, multivariate noise, and stochastic random averaging methods together, the system creates a more robust training dataset that benefits from the strengths of each individual method while compensating for their individual weaknesses, thereby improving both generalization and reliability.
2Manufacturing precision
If multiple augmentation algorithms are used to generate diverse training data, then overfitting is reduced and classification accuracy is improved, but the complexity of the data processing system increases
Solution Approach 1:
The patent segments the data augmentation process into distinct modular algorithms (PCA component, multivariate noise component, stochastic random averaging component). Each algorithm operates as an independent module that can be applied separately to the hyperspectral data, allowing the system to achieve high classification accuracy through diverse data generation while maintaining manageable complexity through modular design.
Data Source
AI summary
A method for creating training data for an artificial intelligence system to classify hyperspectral data. The method including receiving a hyperspectral image from a data source, wherein the hyperspectral image includes a first pixel group associated with a first classification, forming from the hyperspectral image a first augmented image using a first augmentation algorithm and a second augmented image using a second augmentation algorithm, selecting a first group of sample pixels from the hyperspectral image, the first augmented image and the second augmented image, wherein each pixel of the selected first group of sample pixels is having an association with the first classification or with a second classification and providing the selected first group of sample pixels and the associated classifications of each pixel for an artificial training system to be used as a training data.


