Automated Training Data Generation for Re-Identification Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for training machine-learning models for visual person re-identification require extensive manual annotation, which is time-consuming and inefficient, limiting the scalability and adaptability of re-identification systems.

Innovation Solution

A computer system and method that automatically generate training data by grouping samples of media data representing the same person, animal, or object into tuples using secondary information such as positional, temporal, or wireless identifiers, allowing for triplet loss-based training without relying on manual effort.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation methods are used to create training images, then training data can be generated, but the process becomes work-intensive and slow

Engineering Contradiction:
Improvetraining data qualityVSAvoidtraining efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system automatically generates training data by self-organizing media samples into tuples based on secondary information, eliminating the need for manual annotation. The computer system performs all grouping and training data generation operations autonomously, allowing the training process to serve itself without human intervention.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary grouping of media samples into tuples using secondary information before the actual training process. This pre-organization of training data into structured tuples (baseline, positive, negative) prepares everything in advance, so that when training begins, the data is already ready to use without requiring manual annotation during the training phase.

Inventive Principle:
Principle #10Preliminary action

2Quantity of substance

If large amounts of manual annotation work are performed, then comprehensive training data can be created, but the process becomes time-consuming

Engineering Contradiction:
Improvetraining data volumeVSAvoidannotation time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The system replaces the mechanical manual annotation process with an automated computer-based system that uses secondary information (metadata, timestamps, GPS coordinates, device identifiers) to automatically group media samples. This substitution eliminates the need for human annotators to manually review and label each image, dramatically reducing annotation time while maintaining or increasing training data volume.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system introduces secondary information as an intermediary to automatically associate media samples with their corresponding tuples. Instead of manual annotation directly linking images to identities, the secondary information (such as metadata from mobile devices, timestamps, and location data) acts as a mediator that automatically groups samples, enabling rapid generation of large volumes of training data.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Extent of automation

If automated grouping using secondary information is implemented, then manual annotation effort is reduced, but system complexity increases

Engineering Contradiction:
Improvetraining data generation automationVSAvoidsystem complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The system uses secondary information that is already universally generated by mobile devices and media capture systems (metadata, timestamps, GPS, device identifiers). By leveraging this existing multi-functional data that serves multiple purposes (device operation, location tracking, timing), the system avoids adding dedicated complex annotation infrastructure, achieving automation while keeping complexity manageable.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12032656B2Concept for generating training data and training a machine-learning model for use in re-identification
Publication Date: 2024.07.09 GRAZPER TECH APS
  • US12032656B2 patent drawing
  • US12032656B2 patent drawing
  • US12032656B2 patent drawing

AI summary

Examples relate to a concept for generating training data and training a machine-learning model for use in re-identification. A computer system for generating training data for training a machine-learning model for use in re-identification comprising processing circuitry configured to obtain media data, the media data comprising a plurality of samples representing a person, an animal or an object. The processing circuitry is configured to process the media data to identify tuples of samples that represent the same person, animal or object. The processing circuitry is configured to generate the training data based on the identified tuples of samples that represent the same person, animal or object.