Ensemble Data Filters for Machine Learning Training Set Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning algorithms, especially those used in autonomous systems, face challenges in learning information not included in their training data sets, leading to safety-critical situations, and require large storage capacities for training and testing data sets.
Innovation Solution
A method utilizing an ensemble of data filters to selectively generate a data set for training and testing by filtering and classifying data based on the machine learning algorithm's requirements, reducing the amount of data needed and optimizing its properties.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If all information is included in the training data set, then the machine learning algorithm can learn more properties, but the storage requirement for storing the training data set increases
Solution Approach 1:
The patent extracts only the essential and relevant information from the raw data set using filtering criteria. Instead of storing all available data, the system identifies and retains only those data elements that meet predefined relevance thresholds, thereby reducing storage requirements while preserving the most valuable information for training the machine learning algorithm
Solution Approach 2:
The patent applies different filtering criteria and relevance thresholds to different portions or types of data within the data set. Rather than uniformly processing all data, the system tailors the filtering approach to specific data characteristics, ensuring that each portion is evaluated according to its own importance and relevance to the machine learning task
2Adaptability or versatility
If a large data set is used for training, then more properties can be learned, but the storage capacity required increases
Solution Approach 1:
The patent extracts only the essential and relevant information from the raw data set using filtering criteria. Instead of storing all available data, the system identifies and retains only those data elements that meet predefined relevance thresholds, thereby reducing storage requirements while preserving the most valuable information for training the machine learning algorithm
Solution Approach 2:
The patent dynamically adjusts filtering parameters and relevance thresholds based on the specific requirements of the machine learning algorithm and the characteristics of the data. By optimizing these parameters, the system achieves an optimal balance between the comprehensiveness of the training data and the storage resources required
3Quantity of substance
If the training data set is reduced, then storage requirements decrease, but information not included in the training data set cannot be learned
Solution Approach 1:
The patent applies different filtering criteria and relevance thresholds to different portions or types of data within the data set. Rather than uniformly processing all data, the system tailors the filtering approach to specific data characteristics, ensuring that each portion is evaluated according to its own importance and relevance to the machine learning task
Solution Approach 2:
The patent dynamically adjusts filtering parameters and relevance thresholds based on the specific requirements of the machine learning algorithm and the characteristics of the data. By optimizing these parameters, the system achieves an optimal balance between the comprehensiveness of the training data and the storage resources required
Data Source
AI summary
A method for generating a data set for training and/or testing a machine learning algorithm. The method includes: providing a first data set, wherein the first data set comprises data potentially relevant to the machine learning algorithm, providing an ensemble of data filters, configuring each data filter of the ensemble of data filters on the basis of requirements of the machine learning algorithm, and selecting the first data set by filtering the first data set by means of at least a part of the configured data filters of the ensemble of data filters in order to obtain data for training and/or testing the machine learning algorithm, wherein the data form the data set for training and/or testing the machine learning algorithm.


