Synthetic Scene Labels for Multi-Object Tracking Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Deep learning models for multi-object tracking face challenges in obtaining high-quality, properly-labeled training data, which is time-consuming and expensive to produce, and may raise privacy concerns when using real human data.

Innovation Solution

Utilize simulation to generate synthetic training data with automatic labeling, detecting and modifying tag data for occlusions, proximity, and motion alerts to improve training efficiency and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If real human data is used for training, then tracking accuracy may be improved, but privacy concerns arise and data collection becomes complex

Engineering Contradiction:
Improvetracking accuracyVSAvoidprivacy concerns
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent creates synthetic copies of human data through simulation environments. Virtual humans with realistic appearances and behaviors are generated to replicate real human interactions and scenarios, providing training data that maintains accuracy while eliminating privacy risks associated with using actual human data

Inventive Principle:
Principle #26Copying

2Measurement precision

If manual labeling is performed to ensure data quality, then labeling precision improves, but time consumption and cost increase significantly

Engineering Contradiction:
Improvelabeling precisionVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The simulation environment automatically generates precisely labeled training data without human intervention. The system self-services by creating synthetic data with inherent ground truth labels through the simulation physics and object tracking, eliminating the need for manual labeling while maintaining high precision

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent performs data labeling in advance during the simulation data generation process. All annotations, bounding boxes, and tracking labels are pre-computed as part of the simulation output, so when the data is used for training, no additional manual labeling time is required

Inventive Principle:
Principle #10Preliminary action

3Productivity

If frame-skipping and interpolation techniques are used, then productivity improves, but measurement precision of tracking deteriorates

Engineering Contradiction:
Improvelabeling efficiencyVSAvoidtracking precision
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The simulation environment continuously tracks all virtual objects at every time step and automatically generates labels for all frames without skipping. The system self-services by maintaining precise tracking state throughout the simulation, providing complete frame-by-frame labeling data that preserves tracking precision while enabling efficient batch processing

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250278843A1Training multi-object tracking models using simulation
Publication Date: 2025.09.04 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20250278843A1 patent drawing
  • US20250278843A1 patent drawing
  • US20250278843A1 patent drawing

AI summary

Training a multi-object tracking model includes: generating a plurality of training images based at least on scene generation information, each training image comprising a plurality of objects to be tracked; generating, for each training image, original simulated data based at least on the scene generation information, the original simulated data comprising tag data for a first object; locating, within the original simulated data, tag data for the first object, based on at least an anomaly alert (e.g., occlusion alert, proximity alert, motion alert) associated with the first object in the first training image; based at least on locating the tag data for the first object, modifying at least a portion of the tag data for the first object from the original simulated data, thereby generating preprocessed training data from the original simulated data; and training a multi-object tracking model with the preprocessed training data to produce a trained multi-object tracker.