Continuous ML Model Training on Ancillary Data Streams

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning model training methods require large dedicated datasets, which can be infeasible or costly in terms of resource usage and human effort, especially for large or 'deep' models.

Innovation Solution

The proposed method involves continuously training machine learning models on changing data by sampling from ancillary systems, training on the sampled data, evaluating model performance, comparing it with other models, and selecting the best model for deployment based on performance metrics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If large dedicated training datasets are collected and stored, then model training is enabled, but computer resource usage and storage costs increase significantly

Engineering Contradiction:
Improvetraining data amountVSAvoidstorage requirement
Core Design Contradiction:
Quantity of substanceVSVolume of stationary object

Solution Approach 1:

The patent applies multi-functionality by using ancillary systems (which serve other primary purposes) as data sources for training machine learning models. Instead of dedicated training data collection, the system leverages data from systems that already exist for other functions, thereby enabling model training without requiring separate large-scale data storage infrastructure.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces an intermediary mechanism where a data pooling system acts as a mediator between ancillary systems and the machine learning model training process. This intermediary collects and manages data from multiple ancillary systems, providing a unified data source that enables training without direct large-scale storage requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If large dedicated training datasets are collected and stored, then model training is enabled, but human time and effort for labeling increase significantly

Engineering Contradiction:
Improvetraining data amountVSAvoidlabeling time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent applies self-service by enabling ancillary systems to provide their own data for training purposes without requiring dedicated human labeling efforts. The ancillary systems automatically contribute data from their operational streams, eliminating the need for manual annotation of large training datasets.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent leverages the multi-functional nature of ancillary systems to serve dual purposes: their primary functions and simultaneously serving as data sources for machine learning training. This eliminates the need for separate dedicated training data collection and labeling processes.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If models are trained on changing data continuously, then model performance adapts to current data, but data from ancillary systems must be constantly sampled and updated

Engineering Contradiction:
Improvemodel adaptability to changing dataVSAvoiddata sampling and update process
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements continuous training by maintaining an ongoing process where models are repeatedly trained on newly sampled data from ancillary systems. This continuous cycle of sampling, training, and evaluation enables the model to adapt to changing data distributions without requiring discrete batch processing.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The patent incorporates feedback mechanisms where model performance is evaluated on testing data, and this performance information feeds back into the selection process. The system uses performance comparisons to determine which models to deploy, creating a closed-loop feedback system that continuously improves model adaptability.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250148365A1Continuous Training of Machine Learning Models on Changing Data
Publication Date: 2025.05.08 GOOGLE LLC
  • US20250148365A1 patent drawing
  • US20250148365A1 patent drawing
  • US20250148365A1 patent drawing

AI summary

Provided are systems and methods for continuous training of machine learning (ML) models on changing data. In particular, the present disclosure provides example approaches to model training that take advantage of constantly evolving data that may be available in various ancillary systems that contain large amounts of data, but which are not specific to or dedicated for model training.