Multi-Object Tracking Training With Synthetic Simulation Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Deep learning models for multi-object tracking face challenges in obtaining high-quality, properly-labeled training data, which is time-consuming and expensive to produce, and may raise privacy concerns when using real human data.

Innovation Solution

Generate synthetic training data using simulation environments that automatically label objects, including occlusion and anomalous conditions, to create preprocessed training data for multi-object tracking models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual labeling of real video frames is used to train multi-object tracking models, then the training data quality is improved, but the time and cost required for data preparation increases significantly

Engineering Contradiction:
Improvetraining data qualityVSAvoiddata preparation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent creates synthetic copies of training data by rendering virtual objects in simulated environments. Instead of manually labeling real video frames, the system generates synthetic images with automatically generated ground truth annotations, copying the essential training requirements without the manual labor overhead.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary actions by pre-rendering synthetic training data with known ground truth annotations before actual model training begins. The simulation environment pre-computes object positions, occlusions, and annotations in advance, eliminating the need for post-capture labeling work.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If more training data is collected from real videos to improve model performance, then the model accuracy is improved, but the complexity and cost of data annotation increases

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata annotation complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces complex real-world data collection and annotation processes with synthetic data generation. Virtual copies of objects are rendered in controlled environments where ground truth is automatically known, eliminating the need for complex manual annotation pipelines while maintaining model training effectiveness.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The simulation environment serves itself by automatically generating ground truth annotations during the rendering process. The system self-annotates synthetic images with precise object locations, identities, and occlusion states without requiring external annotation tools or human annotators.

Inventive Principle:
Principle #25Self-service

3Reliability

If real human data is used for training to improve tracking performance, then the model learns from authentic scenarios, but privacy concerns arise

Engineering Contradiction:
Improvetracking performanceVSAvoidprivacy concerns
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent creates synthetic copies of human subjects using virtual avatars in simulated environments. These synthetic representations mimic real human behavior and appearances sufficiently for training purposes while containing no actual personal information, thereby eliminating privacy risks associated with using real human data.

Inventive Principle:
Principle #26Copying

4Productivity

If frame-skipping and interpolation techniques are used to reduce labeling workload, then the annotation time is reduced, but the quality and sufficiency of training data decreases

Engineering Contradiction:
Improveannotation efficiencyVSAvoidtraining data sufficiency
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

Instead of reducing labeling effort through frame-skipping, the system copies complete training data for all frames by rendering every frame in the simulation sequence with full annotations. This generates sufficient training data without requiring any frame interpolation or skipping techniques.

Inventive Principle:
Principle #26Copying

Data Source

PatentEP4214677B1Training multi-object tracking models using simulation
Publication Date: 2025.09.24 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP4214677B1 patent drawingFigure 1A
  • EP4214677B1 patent drawingFigure 1B
  • EP4214677B1 patent drawingFigure 2

AI summary

Training a multi-object tracking model includes: generating a plurality of training images based at least on scene generation information, each training image comprising a plurality of objects to be tracked; generating, for each training image, original simulated data based at least on the scene generation information, the original simulated data comprising tag data for a first object; locating, within the original simulated data, tag data for the first object, based on at least an anomaly alert (e.g., occlusion alert, proximity alert, motion alert) associated with the first object in the first training image; based at least on locating the tag data for the first object, modifying at least a portion of the tag data for the first object from the original simulated data, thereby generating preprocessed training data from the original simulated data; and training a multi-object tracking model with the preprocessed training data to produce a trained multi-object tracker.