Synthetic Data Service Engine for Automated ML Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for developing machine-learning training datasets are limited by the complexity and cost of manual development, making high-quality datasets inaccessible and expensive, and existing solutions are inadequate for universal use across different domains.

Innovation Solution

A distributed computing system providing synthetic data as a service (SDaaS) that abstracts underlying operations, allowing customers to configure, generate, access, and manage synthetic data training datasets using engines like asset assembly, scene assembly, frameset package generation, and crowdsourcing, reducing the need for manual data labeling and refinement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual methods are used to develop machine-learning training datasets, then data quality can be improved, but the complexity and cost increase significantly

Engineering Contradiction:
Improvedata qualityVSAvoiddevelopment complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent uses synthetic data generation to create artificial training datasets that replicate the statistical properties and characteristics of real-world data without requiring manual collection or labeling. The system generates synthetic images, audio, or text data through computational models, providing high-quality training data at reduced complexity and cost.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system employs automated pipelines where machine-learning models generate their own training data through synthetic data generation. The feedback loop engine automatically identifies data deficiencies and triggers synthetic data creation without human intervention, enabling self-service data development.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If manual labeling is performed to create training datasets, then data accuracy improves, but the time and resources required increase

Engineering Contradiction:
Improvedata accuracyVSAvoiddevelopment time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

Instead of manually labeling real data, the system creates synthetic copies of data with known ground truth labels embedded during generation. This eliminates the time-consuming manual labeling process while maintaining high accuracy through controlled generation parameters and automated verification.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary data generation and validation before actual training needs arise. Synthetic data is pre-generated with embedded labels and verified for quality, so when training is required, accurate data is already prepared and available immediately.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If comprehensive training datasets are created for multiple domains, then versatility improves, but the cost and complexity of development increase

Engineering Contradiction:
Improvedomain coverageVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The synthetic data generation system is designed as a universal platform that can generate training data for multiple domains (images, audio, text, sensor data) using a common architecture. The feedback loop engine and parameter control mechanisms work across different data types, providing versatile domain coverage without proportionally increasing complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system achieves domain versatility by adjusting generation parameters rather than building separate systems for each domain. By modifying parameters such as data type, distribution characteristics, and domain-specific features, the same synthetic generation infrastructure can produce diverse training datasets across multiple application areas.

Inventive Principle:
Principle #35Parameter changes

4Reliability

If high-quality training datasets are produced manually, then model performance improves, but computational overhead and memory requirements increase

Engineering Contradiction:
Improvemodel performanceVSAvoidcomputational overhead
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system generates synthetic training data that preserves the essential statistical properties and relationships needed for model training without requiring expensive manual curation processes. The synthetic copies provide sufficient training quality while reducing the computational overhead associated with manual data preparation, validation, and management.

Inventive Principle:
Principle #26Copying

Data Source

PatentEP3803721B1Distributed computing system with a synthetic data as a service feedback loop engine
Publication Date: 2023.04.05 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3803721B1 patent drawingFigure 1A
  • EP3803721B1 patent drawingFigure 1B
  • EP3803721B1 patent drawingFigure 2A

AI summary

Various embodiments, methods and systems for implementing a distributed computing system feedback loop engine are provided. Initially, a training dataset report is accessed. The training dataset report identifies a synthetic data asset having values for asset- variation parameters. The synthetic data asset is associated with a frameset. Based on the training dataset report, the synthetic data asset with a synthetic data asset variation is updated. The frameset is updated using the updated synthetic data asset.