SDaaS Synthetic Data Engine for ML Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for developing machine-learning training datasets are inadequate, as they are time-consuming, expensive, and lack universal availability across domains, with limited comprehensive functionality, making high-quality datasets difficult to produce and inaccessible for widespread use.

Innovation Solution

A distributed computing system providing synthetic data as a service (SDaaS) that abstracts underlying operations, allowing customers to configure, generate, access, and manage synthetic data training datasets using engines like asset assembly, scene assembly, frameset package generation, and crowdsourcing, reducing manual development complexity and enabling mass production of training datasets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional methods are used to develop machine-learning training datasets, then manual labeling and curation can be performed, but the process is time-consuming and labor-intensive

Engineering Contradiction:
Improvedataset qualityVSAvoiddataset creation time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system generates synthetic data that replicates real-world patterns and characteristics without requiring actual real-world data collection. The synthetic data generator creates virtual instances that mimic real data distributions, allowing rapid dataset creation without manual labeling or physical data gathering, thus resolving the contradiction between data quality and creation time

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces manual mechanical labeling processes with automated computational generation. Instead of human annotators manually tagging data, the system uses algorithms to automatically generate and label synthetic data, substituting the mechanical manual process with automated computational operations that are much faster and more scalable

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If comprehensive training datasets are created manually, then high-quality labeled data can be obtained, but the cost and complexity increase significantly

Engineering Contradiction:
Improvedata labeling accuracyVSAvoiddataset development complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system enables self-service dataset generation where users can configure parameters and generate synthetic data without requiring complex manual curation processes. The synthetic data generator automatically handles data synthesis, labeling, and formatting based on user-defined parameters, eliminating the need for complex manual workflows while maintaining data quality

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent uses parameter variation to generate diverse synthetic data by adjusting control parameters such as data distribution characteristics, complexity levels, and domain-specific attributes. By programmatically varying these parameters, the system can generate comprehensive datasets across different domains without manually creating each data point, thus reducing complexity while maintaining quality

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If theoretical solutions for developing training datasets are implemented, then alternative techniques can be explored, but the infrastructure is inaccessible or too expensive

Engineering Contradiction:
Improvedataset generation flexibilityVSAvoidinfrastructure accessibility
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The system provides a universal synthetic data generation platform that can create datasets for multiple domains and applications through a single integrated system. The synthetic data generator is designed to handle various data types and scenarios without requiring separate specialized infrastructure for each domain, making the solution broadly accessible and adaptable while reducing overall infrastructure complexity and cost

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11615137B2Distributed computing system with a crowdsourcing engine
Publication Date: 2023.03.28 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11615137B2 patent drawing
  • US11615137B2 patent drawing
  • US11615137B2 patent drawing

AI summary

Various embodiments, methods and systems for implementing a distributed computing system crowdsourcing engine are provided. Initially, a source asset is received from a distributed synthetic data as a service (SDaaS) crowdsource interface. A crowdsource tag is received for the source asset via the distributed SDaaS crowdsource interface. Based in part on the crowdsource tag, the source asset is ingested. Ingesting the source asset comprises automatically computing values for asset-variation parameters of the source asset. The asset-variation parameters are programmable for machine-learning. A crowdsourced synthetic data asset comprising the values for asset-variation parameters is generated.