Personalized Training Data Synthesis for ML Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Obtaining and labeling training data for machine learning systems is labor-intensive and costly, especially due to the sheer volume of data required, and synthetic data alone can struggle to accurately represent real-world scenarios, known as the 'synth to real' gap.

Innovation Solution

A method for generating personalized training data using a computational model based on user image data, combining real-world 3D topology and synthetic elements, which allows for automatic labeling and generation of varied views under different conditions, reducing manual labor and improving data accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual labelling of training data is performed, then data accuracy is improved, but productivity deteriorates due to arduous and time-consuming human involvement

Engineering Contradiction:
Improvedata labelling accuracyVSAvoiddata generation speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent uses synthetic data generation to create copies of real-world training data scenarios. A computational model generates synthetic images, point clouds, and labels that replicate real driving conditions without requiring manual human labelling, thus maintaining accuracy while dramatically improving productivity

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs self-service by automatically generating labelled training data through computational models without human intervention. The model synthesizes data and assigns labels autonomously, eliminating the need for manual labelling while maintaining data quality

Inventive Principle:
Principle #25Self-service

2Productivity

If synthetic data is used for training, then productivity is improved due to easier and cheaper generation, but measurement precision deteriorates due to the 'synth to real' gap

Engineering Contradiction:
Improvedata generation efficiencyVSAvoidreal-world data accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent employs parameter changes by systematically varying environmental conditions (lighting, weather, time of day) and object parameters in synthetic data generation. This creates diverse, realistic training scenarios that bridge the synth-to-real gap while maintaining high productivity through automated generation

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system creates composite training datasets by combining synthetic generated data with real-world data characteristics. The computational model integrates multiple data types (images, point clouds, labels) with varying degrees of synthetic and real-world properties to achieve both efficiency and accuracy

Inventive Principle:
Principle #40Composite materials

3Reliability

If large volume of training data is collected, then machine learning model performance is improved, but loss of time increases due to the sheer amount of data required

Engineering Contradiction:
Improvemodel performanceVSAvoiddata collection duration
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-generating comprehensive training datasets using computational models before actual model training begins. Synthetic data covering various scenarios is created in advance, eliminating the time-consuming process of collecting and labelling large volumes of real-world data

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240249507A1Method and system for dataset synthesis
Publication Date: 2024.07.25 SONY INTERACTIVE ENTERTAINMENT LLC
  • US20240249507A1 patent drawing
  • US20240249507A1 patent drawing
  • US20240249507A1 patent drawing

AI summary

A method of dataset generation is described. The method comprises steps including receiving user image data and generating personalised training data based on the received user image data. Generating personalised training data comprises generating a computational model based at least in part on the received user data and generating the personalised training data based on the computational model.