Adaptive Oracle-Trained Learning System for Incremental Model Retraining

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems for building and maintaining machine learning models that process dynamic data face challenges due to data quality fluctuations, making it difficult and costly to obtain high-quality training data and adapt models to changing data distributions.

Innovation Solution

An adaptive oracle-trained learning framework that leverages a crowd or oracle for generating high-quality training data and uses active learning to monitor model performance and update training data sets, enabling incremental adaptation of models to maintain data quality and reduce replacement costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional methods are used to obtain high-quality training data, then data quality is improved, but cost and time consumption increase

Engineering Contradiction:
Improvedata qualityVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs self-service by automatically curating training data through a combination of automated filtering, active learning, and oracle feedback mechanisms. The system independently identifies high-quality data instances without requiring manual intervention, thereby maintaining high data quality while reducing time consumption.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements feedback loops where the oracle evaluates model predictions and provides feedback on data quality. This feedback is used to iteratively improve the training data curation process, allowing the system to automatically identify and select high-quality data instances more efficiently over time.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If traditional methods are used to obtain high-quality training data, then data quality is improved, but cost increases

Engineering Contradiction:
Improvedata qualityVSAvoidcost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The system reduces costs by performing self-service data curation through automated algorithms that filter and select training data. This eliminates the need for expensive manual annotation and verification processes, maintaining high data quality while significantly reducing the cost of data preparation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system uses copying by creating synthetic training data instances through active learning and data augmentation techniques. This allows the system to generate high-quality training data virtually, reducing the need for expensive physical data collection and manual verification processes.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If models are replaced to adapt to changing data distributions, then model effectiveness is improved, but replacement costs increase

Engineering Contradiction:
Improvemodel effectivenessVSAvoidreplacement costs
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The system implements dynamics by enabling incremental updates to the training data set rather than complete model replacements. The system dynamically adjusts the training data composition based on changing data distributions, allowing the model to adapt continuously while avoiding the costs associated with frequent full model replacements.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs preliminary action by proactively curating and preparing updated training data before complete model replacement becomes necessary. This allows the system to gradually transition to new data distributions, reducing the urgency and cost of full model replacements by preparing incremental updates in advance.

Inventive Principle:
Principle #10Preliminary action

4Reliability

If complete model replacement is used instead of incremental adaptation, then model effectiveness is ensured, but time and cost are consumed

Engineering Contradiction:
Improvemodel effectivenessVSAvoidtime consumption
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system implements dynamics by transitioning from static complete model replacement to dynamic incremental adaptation. The system continuously updates the training data set in response to changing conditions, allowing the model to maintain effectiveness through gradual evolution rather than periodic replacements, thereby reducing time consumption.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system ensures continuity of useful action by maintaining continuous incremental updates to the training data set rather than discontinuous complete replacements. This continuous adaptation process ensures model effectiveness is maintained at all times while eliminating the time losses associated with periodic full model replacement cycles.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS10614373B1Processing dynamic data within an adaptive oracle-trained learning system using curated training data for incremental re-training of a predictive model
Publication Date: 2020.04.07 BYTEDANCE INC
  • US10614373B1 patent drawing
  • US10614373B1 patent drawing
  • US10614373B1 patent drawing

AI summary

In general, embodiments of the present invention provide systems, methods and computer readable media for an adaptive oracle-trained learning framework for automatically building and maintaining models that are developed using machine learning algorithms. In embodiments, the framework leverages at least one oracle (e.g., a crowd) for automatic generation of high-quality training data to use in deriving a model. Once a model is trained, the framework monitors the performance of the model and, in embodiments, leverages active learning and the oracle to generate feedback about the changing data for modifying training data sets while maintaining data quality to enable incremental adaptation of the model.