Machine Learning Training Data Reduction via Symbolic Grouping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning models require large datasets for training, making it inefficient to parse and classify diverse program guide listings data from multiple sources with different formats, especially when data is non-structured or from printed media.
Innovation Solution
A media guidance application parses data records into subsets, assigns symbols based on character types, and uses user input to classify these symbols, grouping records with identical arrangements and classifications for efficient training with a reduced data set, incorporating visual clues and descriptive vectors for non-electronic sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If machine learning models are trained using large datasets to ensure accuracy, then model accuracy is improved, but training time and computational resources increase significantly
Solution Approach 1:
The patent segments the dataset into multiple groups based on data characteristics and formats. By training multiple specialized machine learning models on these segmented groups rather than one large model on the entire dataset, the system achieves comparable accuracy with reduced training time and computational resources for each individual model.
Solution Approach 2:
The patent applies partial action by selecting and processing only the most relevant features and data points from each dataset group. Rather than processing entire large datasets, the system identifies and uses critical subsets of data that are sufficient for training accurate models, thereby reducing training time while maintaining accuracy.
2Measurement precision
If data from multiple sources with different formats is processed individually, then data accuracy is improved, but processing complexity increases
Solution Approach 1:
The patent creates a universal data processing framework that can handle multiple data formats and sources through a single standardized interface. The system defines common data structures and processing pipelines that work across different data types, reducing processing complexity while maintaining the ability to accurately process diverse data from multiple sources.
Solution Approach 2:
The patent introduces intermediary components including data normalization layers and format conversion mechanisms that act as mediators between diverse data sources and the machine learning models. These intermediaries standardize data representations without losing critical information, simplifying the overall processing architecture while preserving data accuracy.
3Productivity
If diverse data formats from multiple sources are standardized, then system efficiency is improved, but information loss may occur
Solution Approach 1:
The patent applies local quality by preserving format-specific characteristics and data nuances within standardized structures. Rather than applying uniform standardization that loses information, the system maintains local data qualities and format-specific features within the standardized framework, ensuring efficiency gains without information loss.
Data Source
AI summary
Methods and systems are disclosed herein for accurately training a machine learning model with a reduced training data set. A large number of data records may be parsed. Each record may be reduced to a set of symbols representing the composition of each record. A user may assign a classification to each symbol within each record. Records with identical arrangements and classifications of symbols may be grouped together, and a representative sample of data records from each group may be fed into the model as the reduced training data set.


