Synthetic Image Data Generation With Automatic Labels for Dynamic Movement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The scarcity of labeled datasets for training machine learning models to detect dynamic movements in image sequences, particularly in automotive scenarios, poses challenges due to the difficulty in obtaining real-world footage of dangerous situations and the high production cost and inconsistency of manual labeling, which limits the generalizability and accuracy of these models.

Innovation Solution

A system generates synthetic image data based on user-defined parameters using physics and rendering engines, automatically labeling actions or movements to create datasets that can be used for training or evaluating machine learning models, such as CNNs, to improve generalizability and reduce manual effort.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual labeling of real-world image data is used, then labeled datasets can be obtained for training machine learning models, but production costs are high and labeling consistency is poor

Engineering Contradiction:
Improvelabeling consistencyVSAvoidproduction cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent creates synthetic copies of real-world driving scenarios through simulation environments. Instead of manually labeling actual footage, the system generates synthetic image sequences that replicate dangerous driving situations, allowing automated labeling with perfect consistency while eliminating the high costs and inconsistencies of manual human labeling.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The simulation environment automatically generates and labels its own training data without human intervention. The system self-services by creating synthetic datasets with ground truth labels embedded in the simulation parameters, eliminating the need for expensive manual labeling processes while maintaining perfect labeling accuracy.

Inventive Principle:
Principle #25Self-service

2Quantity of substance

If real-world footage of dangerous situations is obtained, then training data for detecting dynamic movements can be collected, but the data is scarce and difficult to obtain

Engineering Contradiction:
Improvedataset volumeVSAvoiddata acquisition difficulty
Core Design Contradiction:
Quantity of substanceVSDifficulty of detecting and measuring

Solution Approach 1:

The system performs preliminary action by pre-programming various dangerous driving scenarios and conditions in the simulation environment before generating data. By pre-defining edge cases, rare events, and hazardous situations in the simulation logic, the system can generate unlimited quantities of such difficult-to-obtain footage without needing to capture them from actual roads.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates synthetic copies of rare and dangerous driving scenarios that are difficult to capture in reality. By replicating these situations in a controlled simulation environment, the system can generate abundant training data for edge cases that would otherwise be scarce or impossible to obtain from real-world footage.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If more labeled data is collected to improve model generalizability, then model performance improves, but the cost and time for data collection and labeling increase

Engineering Contradiction:
Improvemodel generalizabilityVSAvoiddata production time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The simulation environment enables continuous generation of synthetic training data without interruption. Unlike manual collection and labeling processes that are discrete and time-consuming, the system can continuously generate diverse scenarios with automated labeling, providing an unlimited stream of training data that rapidly improves model generalizability without proportional increases in time or cost.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentEP3732618B1Method, apparatus, and system for generating synthetic image data for machine learning
Publication Date: 2025.07.16 HERE GLOBAL BV
  • EP3732618B1 patent drawingFigure 1
  • EP3732618B1 patent drawingFigure 2
  • EP3732618B1 patent drawingFigure 3

AI summary

An approach is provided for generating synthetic image data for machine learning. The approach, for instance,involvesdetermining, by a processor, a set of parameters for indicating an action by one or more objects. The action is a dynamic movement of the one or more objects through a geographic space over a period of time. The approach also involves processing the set of parameters to generate synthetic image data. The synthetic image data includes a computer- generated image sequence of the one or more objects performing the action in the geographic space over the period of time. The approach further involves automatically labeling the synthetic image data with at least one label representing the action, the set of parameters, or a combination thereof. The approach further involves providing the labeled synthetic image data for training or evaluating a machine learning model to detect the action.