Generic Mapping for Video Object Tracking Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional tracking models fail to effectively account for future appearance changes of target objects in video sequences, leading to drift and incorrect tracking due to background inclusion, and require complex computational processes.
Innovation Solution
A method and apparatus that utilize a generic mapping mechanism learned from external data to adapt to appearance variations, applying it to novel tracking settings without further adaptation, using a deep convolutional network to match target objects across frames by comparing candidate patches with a reference frame.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional tracking models are used to track target objects in video sequences, then the tracking process can be performed with simple models, but the models fail to account for appearance changes leading to drift and incorrect tracking
Solution Approach 1:
The system performs preliminary learning of appearance variations offline before actual tracking. A generic mapping is pre-computed from training data that captures how objects may appear different under various conditions. This preliminary action enables the tracker to handle appearance changes without requiring complex real-time adaptation, thus improving reliability while maintaining simplicity.
Solution Approach 2:
The system changes the parameter representation by using a generic mapping that transforms image patches into a feature space where appearance variations are normalized. Instead of tracking raw pixel values that change with appearance, the system tracks transformed features that remain stable, thereby maintaining tracking accuracy despite appearance changes.
2Reliability
If complex computational processes are used to account for appearance changes, then tracking accuracy can be improved, but computational complexity increases
Solution Approach 1:
Complex computations are moved to an offline preliminary stage where a generic mapping is learned from training data. Once computed, this mapping is stored and reused during actual tracking operations. This eliminates the need for complex real-time computations while maintaining high tracking accuracy, thus improving reliability without increasing operational computational complexity.
Solution Approach 2:
The system creates a copy or representation of appearance variation patterns through the generic mapping. Instead of performing complex comparisons of raw image patches, the system uses the pre-computed mapping to generate simplified feature representations that can be efficiently compared, reducing computational complexity while preserving tracking accuracy.
3Adaptability or versatility
If generic mapping is learned offline from external data, then adaptability to appearance variations is improved, but training data requirements increase
Solution Approach 1:
The generic mapping is designed to be universal and applicable to multiple objects and scenarios. By learning general appearance variation patterns that can apply to different target objects, the system achieves high adaptability without needing extensive training data for each specific object. This universal approach allows the same mapping to handle diverse appearance changes across different tracking scenarios.
Solution Approach 2:
The system transforms the training problem by changing parameters to focus on learning generic appearance variation patterns rather than object-specific features. This parameter transformation allows the system to learn from smaller, more diverse datasets rather than requiring large volumes of data for each specific object class, thereby improving adaptability with reduced data requirements.
Data Source
AI summary
A method of tracking a position of a target object in a video sequence includes identifying the target object in a reference frame. A generic mapping is applied to the target object being tracked. The generic mapping is generated by learning possible appearance variations of a generic object. The method also includes tracking the position of the target object in subsequent frames of the video sequence by determining whether an output of the generic mapping of the target object matches an output of the generic mapping of a candidate object.


