Synthetic Data Generation Service for ML Training Efficiency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for generating and sharing machine learning training datasets are limited by resource constraints, competition among providers, confidentiality, and privacy concerns, leading to restricted availability and high costs, which hinders the democratization of these datasets across different domains.

Innovation Solution

A distributed computing system providing synthetic data as a service (SDaaS) using a service-oriented architecture, which abstracts underlying operations and enables customers to configure, generate, access, and manage synthetic data training datasets, obviating the need for manual development and refinement through engines like asset assembly, scene assembly, and feedback loops.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual development and labeling of training datasets is performed, then data quality and accuracy can be ensured, but significant time and effort are required

Engineering Contradiction:
Improvedata qualityVSAvoidtime and effort
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent uses synthetic data generation to create copies of real-world training data through simulation. Instead of manually labeling real data, the system generates synthetic training datasets that replicate the characteristics and patterns of real data, thereby ensuring data quality while eliminating the time-consuming manual labeling process

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical process of manual data collection, labeling, and annotation with an automated computational system. The synthetic data generation system uses algorithms and simulations to automatically create training datasets, substituting human labor with automated computational processes that maintain data quality while dramatically reducing time and effort

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If more training datasets are generated and shared across domains, then democratization of machine learning is improved, but resource constraints and competition limit availability

Engineering Contradiction:
Improvedataset availabilityVSAvoidresource constraints
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent creates a synthetic data generation system that produces training datasets applicable across multiple domains and use cases. By generating synthetic data that can be adapted to various machine learning tasks and industries, the system maximizes the utility and versatility of the generated datasets, allowing one system to serve multiple purposes and reduce the need for domain-specific data collection resources

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent enables flexible adjustment of synthetic data generation parameters to create datasets tailored to different domains, applications, and resource constraints. By modifying generation parameters such as data volume, complexity, and domain-specific characteristics, the system can optimize dataset production to match available resources while maintaining broad adaptability across different machine learning scenarios

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If confidential and sensitive data is used for training, then real-world accuracy is improved, but privacy and security concerns arise

Engineering Contradiction:
Improvereal-world accuracyVSAvoidprivacy and security risks
Core Design Contradiction:
Manufacturing precisionVSObject-affected harmful factors

Solution Approach 1:

The patent creates synthetic copies of sensitive real-world data that preserve the statistical properties, patterns, and relationships necessary for accurate machine learning training while containing no actual personally identifiable or confidential information. These synthetic copies serve as safe substitutes that maintain real-world accuracy without exposing sensitive data

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces synthetic data as an intermediary between real-world sensitive data and machine learning models. Instead of directly using confidential data, the system generates synthetic representations that mediate the training process, preserving the beneficial learning signals while eliminating privacy and security risks associated with direct use of sensitive data

Inventive Principle:
Principle #24Intermediary (Mediator)

4Ease of operation

If synthetic data generation infrastructure is made accessible, then democratization is improved, but infrastructure costs and complexity increase

Engineering Contradiction:
ImproveaccessibilityVSAvoidinfrastructure complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent implements a self-service synthetic data generation platform where users can independently generate training datasets without requiring deep expertise in infrastructure management. The system provides automated workflows, pre-configured generation parameters, and user-friendly interfaces that enable end users to create synthetic datasets on-demand, eliminating the need for users to manage complex infrastructure themselves

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces a service layer that acts as an intermediary between users and the complex synthetic data generation infrastructure. This service layer abstracts away the underlying computational complexity, resource management, and technical details, presenting a simplified interface to users while handling the intricate infrastructure operations in the background

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240095080A1Automatic suggestion of variation parameters & pre-packaged synthetic datasets
Publication Date: 2024.03.21 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20240095080A1 patent drawing
  • US20240095080A1 patent drawing
  • US20240095080A1 patent drawing

AI summary

Various techniques are described for automatically suggesting variation parameters used to generate a tailored synthetic dataset to train a particular machine learning model. A seeding taxonomy associates a plurality of machine learning scenarios with corresponding subsets of variation parameters. A selected machine learning scenario is used to retrieve a corresponding subset of variation parameters associated with the selected machine learning scenario by the seeding taxonomy. The seeding taxonomy may be adaptable using a feedback loop that tracks selected variation parameters and updates the seeding taxonomy. The suggested variation parameters are presented as suggestions to assist users to identify and select relevant variation parameters faster and more efficiently. Further embodiments relate to pre-packaging synthetic datasets for common or anticipated machine learning scenarios. A user interface may present available packages of synthetic data for a selected industry sector and/or scenario, and a selected package may be made available for download.