Pedestrian Re-Identification via Video Segmentation and Attention
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current pedestrian re-identification methods encode entire videos, leading to poor accuracy due to high variability in pedestrian surface information across frames, which reduces the effectiveness of similarity measurement between videos.
Innovation Solution
The method segments videos into fixed-length segments, encodes each segment separately, and calculates similarity scores between target and candidate segments using a deep network with a collaborative attention mechanism, improving the utilization of pedestrian surface information and accuracy of re-identification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If entire videos are encoded for pedestrian re-identification, then the completeness of pedestrian information is preserved, but the variability in pedestrian surface information across frames reduces measurement accuracy
Solution Approach 1:
The patent divides the video into multiple fixed-length segments, encoding each segment separately rather than encoding the entire video as one unit. This segmentation reduces the variability of pedestrian surface information within each segment while maintaining comprehensive coverage across the whole video through multiple encodings.
2Measurement precision
If multiple video segments are encoded separately, then the accuracy of similarity measurement is improved, but the computational complexity and processing time increase
Solution Approach 1:
The patent combines multiple encoding results from different video segments through a fusion mechanism that integrates the encoded features. This merging approach maintains the accuracy benefits of separate segment encoding while managing computational complexity through efficient feature fusion rather than processing each segment independently to completion.
Data Source
AI summary
A pedestrian re-identification method includes: obtaining a target video containing a target pedestrian and at least one candidate video; encoding each target video segment in the target video and each candidate video segment in the at least one candidate segment separately; determining a score of similarity between the each target video segment and the each candidate video segment according to encoding results, the score of similarity being used for representing a degree of similarity between pedestrian features in the target video segment and the candidate video segment; and performing pedestrian re-identification on the at least one candidate video according to the score of similarity.


