Multi-Part Video Analysis for Person Retrieval
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video tracking techniques focus primarily on facial recognition, making it difficult to identify individuals without clear face shots and fail to track people or baggage across multiple cameras, limiting their effectiveness in retrieval and monitoring applications.
Innovation Solution
A video analysis system that tracks multiple body parts of individuals across frames from multiple cameras, computes scores for each part, and compares feature values to determine identity, allowing for accurate retrieval and tracking of individuals and baggage by selecting the best shots representing each part's features.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If only facial recognition is used for person tracking, then the system is simple to operate, but retrieval accuracy deteriorates when no clear face shot is available
Solution Approach 1:
The patent segments the person tracking task into multiple independent part detection tasks (face, head, upper body, lower body). Each part is detected and tracked separately using dedicated detection models, and the results are integrated to achieve robust person retrieval even when facial features are unavailable or ambiguous.
2Speed
If only face detection is performed, then the detection process is fast, but tracking across multiple cameras fails when face is not captured
Solution Approach 1:
The system employs multiple detection models with different detection ranges and characteristics (face detection model, head detection model, upper body detection model, lower body detection model). These models serve universal purposes by detecting various body parts that can all contribute to person identification and tracking, ensuring reliability across different camera views and conditions.
3Quantity of substance
If only face images are saved as best shots, then storage space is minimized, but person identification becomes difficult when face shots are unavailable
Solution Approach 1:
The patent applies local quality by selecting and saving best shots for each body part (face, head, upper body, lower body) based on detection scores and quality metrics. This allows the system to store multiple types of identifying information with different local qualities, ensuring that at least some identifiable features are always available even when facial shots are poor or unavailable.
4Measurement precision
If multiple body parts are tracked and analyzed, then person retrieval accuracy is improved, but system complexity increases
Solution Approach 1:
The patent merges the detection results, tracking information, and best shot selections from multiple body parts (face, head, upper body, lower body) into a unified person identification system. The scoring mechanism combines evidence from different parts, and the best shot selection integrates multiple part images to achieve accurate person retrieval while managing system complexity through coordinated processing.
Data Source
AI summary
A video analysis apparatus is connected to an image capturing apparatus including a plurality of cameras, in which the video analysis apparatus, for a plurality of persons in images captured by the cameras, tracks at least one of the plurality of persons, detects a plurality of parts of the tracked person, and on the basis of information defining scores used to determine a best shot of each of the parts, computes a score for the part for each of frames of videos from security cameras, compares, for each of the parts, the scores for each part computed for each of the frames to determine a best shot of each of the parts, stores, for each of the parts, a feature value in association with the part, and compares the feature values of each of the parts of the plurality of persons to determine identity of the persons.


