Machine Learning Training Data Reduction via Symbolic Grouping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning models require large datasets for training, making it inefficient to parse and classify diverse program guide listings data from multiple sources with different formats, especially when data is non-structured or from printed media.

Innovation Solution

A media guidance application parses data records into subsets, assigns symbols based on character types, and uses user input to classify these symbols, grouping records with identical arrangements and classifications for efficient training with a reduced data set, incorporating visual clues and descriptive vectors for non-electronic sources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If machine learning models are trained using large datasets to ensure accuracy, then model accuracy is improved, but training time and computational resources increase significantly

Engineering Contradiction:
Improvemodel accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the dataset into multiple groups based on data characteristics and formats. By training multiple specialized machine learning models on these segmented groups rather than one large model on the entire dataset, the system achieves comparable accuracy with reduced training time and computational resources for each individual model.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by selecting and processing only the most relevant features and data points from each dataset group. Rather than processing entire large datasets, the system identifies and uses critical subsets of data that are sufficient for training accurate models, thereby reducing training time while maintaining accuracy.

Inventive Principle:
Principle #16Partial or excessive action

2Measurement precision

If data from multiple sources with different formats is processed individually, then data accuracy is improved, but processing complexity increases

Engineering Contradiction:
Improvedata accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates a universal data processing framework that can handle multiple data formats and sources through a single standardized interface. The system defines common data structures and processing pipelines that work across different data types, reducing processing complexity while maintaining the ability to accurately process diverse data from multiple sources.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces intermediary components including data normalization layers and format conversion mechanisms that act as mediators between diverse data sources and the machine learning models. These intermediaries standardize data representations without losing critical information, simplifying the overall processing architecture while preserving data accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If diverse data formats from multiple sources are standardized, then system efficiency is improved, but information loss may occur

Engineering Contradiction:
Improvesystem efficiencyVSAvoiddata information
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent applies local quality by preserving format-specific characteristics and data nuances within standardized structures. Rather than applying uniform standardization that loses information, the system maintains local data qualities and format-specific features within the standardized framework, ensuring efficiency gains without information loss.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11762947B2Methods and systems for training a machine learning system using a reduced data set
Publication Date: 2023.09.19 ROVI PRODUCT CORP
  • US11762947B2 patent drawing
  • US11762947B2 patent drawing
  • US11762947B2 patent drawing

AI summary

Methods and systems are disclosed herein for accurately training a machine learning model with a reduced training data set. A large number of data records may be parsed. Each record may be reduced to a set of symbols representing the composition of each record. A user may assign a classification to each symbol within each record. Records with identical arrangements and classifications of symbols may be grouped together, and a representative sample of data records from each group may be fed into the model as the reduced training data set.