Synthetic Data Model Selection for Faster ML Deployment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing model selection processes for predicting events such as time series events are inefficient, leading to extensive lead times and poor performance, with issues often discovered only after deployment, resulting in software service downtimes.
Innovation Solution
A system and method utilizing synthetic datasets for training and testing predictive models, employing features for dataset selection, generation, and guided user interface for rapid model selection, including data augmentation and model testing to identify well-performing models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If predictive models are trained with large enterprise datasets, then model performance improves, but training time and computational resources increase significantly
Solution Approach 1:
The patent creates synthetic copies of enterprise data through data generation modules that simulate real data patterns and characteristics. These synthetic datasets serve as substitutes for large-scale enterprise data, enabling model training without requiring access to or processing of massive actual enterprise datasets, thus reducing training time while maintaining model performance.
Solution Approach 2:
The system performs preliminary data preparation by generating and curating synthetic datasets before the actual model training process. This pre-processing step creates ready-to-use training data that eliminates the need for time-consuming data collection, cleaning, and preparation from large enterprise sources during the modeling phase.
2Reliability
If models are extensively tested with multiple iterations, then model reliability improves, but deployment lead time increases
Solution Approach 1:
The system implements automated self-testing and self-evaluation mechanisms where the model testing module automatically evaluates model performance against synthetic test data and provides self-correction feedback. This automation reduces the need for manual iterative testing and human intervention, thereby maintaining model reliability while significantly reducing the overall deployment lead time.
Solution Approach 2:
The system incorporates feedback loops where test results from synthetic data evaluation are automatically fed back into the model training and refinement process. This continuous feedback mechanism enables rapid iteration and improvement without requiring extensive manual testing cycles, thus improving model reliability efficiently within compressed timelines.
3Productivity
If synthetic data is used instead of enterprise data, then training speed improves, but data fidelity to real scenarios may decrease
Solution Approach 1:
The data generation modules employ parameter adjustments and transformations to synthesize data that preserves critical statistical properties and patterns of real enterprise data. By carefully controlling generation parameters to match real-world distributions and relationships, the system maintains data fidelity while achieving rapid training speeds through synthetic data.
Solution Approach 2:
The system addresses data fidelity concerns by adding multiple dimensions to synthetic data generation, including temporal patterns, spatial relationships, and contextual features that mirror real enterprise scenarios. This multi-dimensional approach ensures that synthetic data captures complex real-world characteristics beyond simple statistical properties, maintaining fidelity while enabling fast training.
Data Source
AI summary
Methods, computer program products, and systems are presented. The method, computer program products, and systems can include, for instance: examining an enterprise dataset, the enterprise dataset defined by enterprise collected data; selecting one or more synthetic dataset in dependence on the examining, the one or more synthetic dataset including data other than data collected by the enterprise; training a set of predictive models using data of the one or more synthetic dataset to provide a set of trained predictive models; testing the set of trained predictive models with use of holdout data of the one or more synthetic dataset; and presenting prompting data on a displayed user interface of a developer user in dependence on result data resulting from the testing, the prompting data prompting the developer user to direct action with respect to one or more model of the set of predictive models.


