Scene Assembly Engine for Programmable Synthetic Training Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for generating training datasets for machine-learning models are inefficient, expensive, and lack comprehensive functionality, particularly in democratizing their availability across different domains, and manual development leads to inaccuracies and high costs.

Innovation Solution

A distributed computing system provides synthetic data as a service (SDaaS) that automates the generation and refinement of training datasets using a service-oriented architecture, incorporating engines like asset assembly, scene assembly, frameset assembly, and crowdsourcing to create programmable data representations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual methods are used to create training datasets, then customization and quality control are improved, but time consumption and labor costs increase significantly

Engineering Contradiction:
Improvedata qualityVSAvoidtime consumption
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system generates synthetic data that copies and varies from real-world data patterns without requiring manual creation. The synthetic data generator creates realistic data representations through algorithmic processes, preserving the quality and diversity needed for training while eliminating manual labor and time consumption associated with collecting and annotating real data.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

Manual mechanical processes of data collection, cleaning, and annotation are replaced with automated computational systems. The synthetic data generator uses machine learning models and algorithms to automatically produce training data, substituting human expertise with automated processes that operate continuously without fatigue or time limitations.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If comprehensive training datasets are created, then machine learning model performance is improved, but computational resources and infrastructure costs increase

Engineering Contradiction:
Improvemodel training efficiencyVSAvoidcomputational resource consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The system enables self-service data generation where the synthetic data generator automatically creates training datasets based on user-defined parameters and existing data patterns. This eliminates the need for expensive infrastructure and manual intervention, allowing comprehensive dataset creation through automated processes that consume significantly fewer computational resources.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes parameters of data generation by using synthetic data that can be programmatically controlled. Instead of collecting real data and manually adjusting parameters, the system generates data with controlled characteristics through algorithmic parameter adjustment, achieving comprehensive datasets with lower computational overhead.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If diverse training datasets are generated across multiple domains, then model versatility is improved, but system complexity and development costs increase

Engineering Contradiction:
Improvedomain coverageVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The synthetic data generator is designed as a universal system that can generate diverse training data across multiple domains through a single unified platform. The system uses configurable templates and parameters to create domain-specific synthetic data without requiring separate complex systems for each domain, thereby achieving versatility while managing complexity through standardized architecture.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3803722B1Distributed computing system with a synthetic data as a service scene assembly engine
Publication Date: 2025.12.31 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3803722B1 patent drawingFigure 1A
  • EP3803722B1 patent drawingFigure 1B
  • EP3803722B1 patent drawingFigure 2A

AI summary

Various embodiments, methods and systems for implementing a distributed computing system scene assembly engine are provided. Initially, a selection of a first synthetic data asset and a selection of a second synthetic data asset are received from a distributed synthetic data as a service (SDaaS) integrated development environment (IDE). A synthetic data asset is associated with asset-variation parameters and scene-variation parameters, the asset-variation parameters and scene-variation parameters are programmable for machine-learning. Values for generating a synthetic data scene are received. The values correspond to asset-variation parameters or scene-variation parameters. Based on the values, the synthetic data scene is generated using the first synthetic data asset and the second synthetic data asset.