Continuous ML Model Training on Ancillary Data Streams
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning model training methods require large dedicated datasets, which can be infeasible or costly in terms of resource usage and human effort, especially for large or 'deep' models.
Innovation Solution
The proposed method involves continuously training machine learning models on changing data by sampling from ancillary systems, training on the sampled data, evaluating model performance, comparing it with other models, and selecting the best model for deployment based on performance metrics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If large dedicated training datasets are collected and stored, then model training is enabled, but computer resource usage and storage costs increase significantly
Solution Approach 1:
The patent applies multi-functionality by using ancillary systems (which serve other primary purposes) as data sources for training machine learning models. Instead of dedicated training data collection, the system leverages data from systems that already exist for other functions, thereby enabling model training without requiring separate large-scale data storage infrastructure.
Solution Approach 2:
The patent introduces an intermediary mechanism where a data pooling system acts as a mediator between ancillary systems and the machine learning model training process. This intermediary collects and manages data from multiple ancillary systems, providing a unified data source that enables training without direct large-scale storage requirements.
2Quantity of substance
If large dedicated training datasets are collected and stored, then model training is enabled, but human time and effort for labeling increase significantly
Solution Approach 1:
The patent applies self-service by enabling ancillary systems to provide their own data for training purposes without requiring dedicated human labeling efforts. The ancillary systems automatically contribute data from their operational streams, eliminating the need for manual annotation of large training datasets.
Solution Approach 2:
The patent leverages the multi-functional nature of ancillary systems to serve dual purposes: their primary functions and simultaneously serving as data sources for machine learning training. This eliminates the need for separate dedicated training data collection and labeling processes.
3Adaptability or versatility
If models are trained on changing data continuously, then model performance adapts to current data, but data from ancillary systems must be constantly sampled and updated
Solution Approach 1:
The patent implements continuous training by maintaining an ongoing process where models are repeatedly trained on newly sampled data from ancillary systems. This continuous cycle of sampling, training, and evaluation enables the model to adapt to changing data distributions without requiring discrete batch processing.
Solution Approach 2:
The patent incorporates feedback mechanisms where model performance is evaluated on testing data, and this performance information feeds back into the selection process. The system uses performance comparisons to determine which models to deploy, creating a closed-loop feedback system that continuously improves model adaptability.
Data Source
AI summary
Provided are systems and methods for continuous training of machine learning (ML) models on changing data. In particular, the present disclosure provides example approaches to model training that take advantage of constantly evolving data that may be available in various ancillary systems that contain large amounts of data, but which are not specific to or dedicated for model training.


