Synthetic Training Data Generation for Crowd State Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for determining the state of a crowd in images struggle with low frame rates, leading to reduced accuracy in person trajectory tracking and crowd state determination, especially when dealing with still images.
Innovation Solution
A data generating system that synthesizes images of people in specific states and combines them with background images to create training data, allowing for the generation of a large amount of training data used for machine-learning algorithms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If person regions or head positions are associated between frames to acquire person trajectory or tracking result, then person behavior determination or headcount counting can be performed, but at low frame rates the motion amount of persons increases making it difficult to associate regions/positions between frames, leading to reduced accuracy
Solution Approach 1:
The patent applies preliminary action by pre-defining determination blocks (spatial regions) in the image before performing crowd state determination. Instead of tracking persons across frames, the system divides the image into predetermined determination blocks and aggregates optical flow attributes within each block, enabling accurate analysis even at low frame rates where motion between frames is large.
Solution Approach 2:
The patent segments the image into multiple determination blocks as determination units. By dividing the image space into discrete blocks and analyzing optical flow attributes within each block independently, the system can determine crowd states (steady/non-steady) without needing to associate person regions across frames, thus maintaining accuracy at low frame rates.
2Measurement precision
If optical flow attributes are aggregated for determination blocks to evaluate crowd state, then crowd state determination can be performed, but at low frame rates optical flow becomes difficult to find correctly, reducing accuracy of aggregated attributes and state determination performance
Solution Approach 1:
The patent applies preliminary action by pre-defining determination blocks (spatial regions) in the image before performing crowd state determination. Instead of tracking persons across frames, the system divides the image into predetermined determination blocks and aggregates optical flow attributes within each block, enabling accurate analysis even at low frame rates where motion between frames is large.
Solution Approach 2:
The patent uses determination blocks as disposable analysis units that are defined once and used for aggregating optical flow attributes. Each determination block serves as a temporary container for attribute aggregation, allowing the system to perform crowd state determination without requiring persistent object tracking across frames, thus maintaining performance at low frame rates.
3Measurement precision
If a large amount of training data is collected for machine-learning a dictionary of discriminator, then crowd state recognition accuracy can be improved, but working loads for collecting training data increase
Solution Approach 1:
The patent uses synthetic training data generated by computer graphics as copies of real-world crowd scenes. Instead of collecting large amounts of real training data, the system generates synthetic images with known crowd states (steady/non-steady) using controlled virtual environments, significantly reducing the workload of data collection while providing sufficient training samples for machine learning.
Solution Approach 2:
The patent applies preliminary action by pre-generating synthetic training data with known ground truth labels for crowd states. The training data is prepared in advance with controlled parameters (number of persons, arrangement, motion patterns), allowing the discriminator to be trained efficiently without the need for manual annotation of real-world data.
Data Source
AI summary
In the data generating system, at least one processor generates a synthesized image in which a person image corresponding to a person state is synthesized with a prepared image at a predetermined size, specifies a label for the synthesized image, and outputs a pair of synthesized image and the label.


