Composite Machine Learning Model via Distributed Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning models require large and diverse training datasets to generalize well, leading to time and compute-intensive training processes that are often impractical on a single machine, especially with the constraints of data sovereignty and transfer regulations.
Innovation Solution
A cooperative and coordinated approach is adopted by sharing machine learning models among multiple entities, allowing local data to be trained on and then combining these models to form a composite model that can be used at a single location, thereby leveraging crowdsourced data and compute resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If machine learning models are trained on large and diverse training datasets to achieve good generalization, then model accuracy and robustness are improved, but training time and computational resources increase significantly
Solution Approach 1:
The patent divides the training process into multiple parallel segments across different computing entities. Each entity trains a local model on its own dataset simultaneously, rather than one entity training sequentially on all data. This segmentation of the training workload reduces the time each entity spends training while collectively achieving comprehensive model coverage through multiple diverse datasets.
2Reliability
If machine learning models are trained on large and diverse training datasets to achieve good generalization, then model accuracy and robustness are improved, but computational resources and processing power increase significantly
Solution Approach 1:
The patent combines multiple local models trained on different datasets into a single composite model. By merging the capabilities of multiple models that each processed smaller portions of data, the system achieves the benefits of training on large diverse datasets without requiring any single computing entity to allocate massive computational resources for training on all data simultaneously.
3Adaptability or versatility
If machine learning models are trained locally on separate datasets at different sites, then data sovereignty and transfer regulations are maintained, but model generalization capability is reduced due to lack of diverse training data
Solution Approach 1:
The patent uses locally-trained models as intermediaries that represent the diverse training data without requiring actual data transfer. Each site keeps its data locally for training, maintaining sovereignty, but the resulting models serve as intermediaries that capture the patterns and diversity of their respective datasets. These model intermediaries are then combined to achieve generalization capabilities equivalent to training on all diverse data simultaneously, without violating data transfer regulations.
Data Source
AI summary
Aspects of the subject disclosure may include, for example, combining a plurality of machine learning (ML) models to form a composite model, the plurality of ML models including a first ML model trained on first local data received at a first network location and added models including a second ML model through an nth ML model, respective added models of the added models each being respectively trained on respective training data at a respective network location remote from the first network location; receiving input data at the first network location; providing the input data to the composite model; receiving, from the composite model, a conclusion about a status of the input data; receiving an indication to update one or more models of the plurality of ML models; and updating the one or more models according to the indication. Other embodiments are disclosed.


