Synthetic Data Service Engine for Automated ML Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for developing machine-learning training datasets are limited by the complexity and cost of manual development, making high-quality datasets inaccessible and expensive, and existing solutions are inadequate for universal use across different domains.
Innovation Solution
A distributed computing system providing synthetic data as a service (SDaaS) that abstracts underlying operations, allowing customers to configure, generate, access, and manage synthetic data training datasets using engines like asset assembly, scene assembly, frameset package generation, and crowdsourcing, reducing the need for manual data labeling and refinement.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If manual methods are used to develop machine-learning training datasets, then data quality can be improved, but the complexity and cost increase significantly
Solution Approach 1:
The patent uses synthetic data generation to create artificial training datasets that replicate the statistical properties and characteristics of real-world data without requiring manual collection or labeling. The system generates synthetic images, audio, or text data through computational models, providing high-quality training data at reduced complexity and cost.
Solution Approach 2:
The system employs automated pipelines where machine-learning models generate their own training data through synthetic data generation. The feedback loop engine automatically identifies data deficiencies and triggers synthetic data creation without human intervention, enabling self-service data development.
2Measurement precision
If manual labeling is performed to create training datasets, then data accuracy improves, but the time and resources required increase
Solution Approach 1:
Instead of manually labeling real data, the system creates synthetic copies of data with known ground truth labels embedded during generation. This eliminates the time-consuming manual labeling process while maintaining high accuracy through controlled generation parameters and automated verification.
Solution Approach 2:
The system performs preliminary data generation and validation before actual training needs arise. Synthetic data is pre-generated with embedded labels and verified for quality, so when training is required, accurate data is already prepared and available immediately.
3Adaptability or versatility
If comprehensive training datasets are created for multiple domains, then versatility improves, but the cost and complexity of development increase
Solution Approach 1:
The synthetic data generation system is designed as a universal platform that can generate training data for multiple domains (images, audio, text, sensor data) using a common architecture. The feedback loop engine and parameter control mechanisms work across different data types, providing versatile domain coverage without proportionally increasing complexity.
Solution Approach 2:
The system achieves domain versatility by adjusting generation parameters rather than building separate systems for each domain. By modifying parameters such as data type, distribution characteristics, and domain-specific features, the same synthetic generation infrastructure can produce diverse training datasets across multiple application areas.
4Reliability
If high-quality training datasets are produced manually, then model performance improves, but computational overhead and memory requirements increase
Solution Approach 1:
The system generates synthetic training data that preserves the essential statistical properties and relationships needed for model training without requiring expensive manual curation processes. The synthetic copies provide sufficient training quality while reducing the computational overhead associated with manual data preparation, validation, and management.
Data Source
Figure 1A
Figure 1B
Figure 2A
AI summary
Various embodiments, methods and systems for implementing a distributed computing system feedback loop engine are provided. Initially, a training dataset report is accessed. The training dataset report identifies a synthetic data asset having values for asset- variation parameters. The synthetic data asset is associated with a frameset. Based on the training dataset report, the synthetic data asset with a synthetic data asset variation is updated. The frameset is updated using the updated synthetic data asset.