Visual Tracking Using Distracter and Supporter Context
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Visual tracking in unconstrained environments is challenging due to variations in appearance, cluttered backgrounds, and the emergence of regions with similar appearance to the target, leading to tracking failures, especially when the target leaves the field of view and reappears.
Innovation Solution
The method exploits context information by identifying 'distracters' and 'supporters' using a sequential randomized forest and online template-based appearance models, where distracters are regions with similar appearance and supporters are local key-points with motion correlation, to prevent drift and reacquire the target effectively.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional visual tracking methods are used in unconstrained environments, then the tracking process is simple, but tracking reliability deteriorates due to appearance variations, cluttered backgrounds, and similar-looking regions
Solution Approach 1:
The patent segments the visual tracking problem into multiple components: target appearance modeling, background segmentation, distracter identification, and supporter detection. By dividing the complex tracking task into these manageable segments, the system achieves higher reliability without overwhelming complexity, as each segment can be optimized independently
Solution Approach 2:
The patent introduces intermediate structures such as appearance models and context descriptors that act as mediators between the raw video data and the tracking decision. These intermediaries process and filter information, improving tracking reliability by providing a structured representation that is more robust to appearance variations and cluttered backgrounds
2Measurement precision
If the tracker follows regions with similar appearance to the target, then detection sensitivity increases, but tracking precision deteriorates due to drift and loss of the correct target
Solution Approach 1:
The patent applies local quality by analyzing specific local features and descriptors within the target region and its surroundings. Instead of relying solely on global appearance matching, the system examines local patterns, textures, and contextual relationships that are unique to the target, enabling precise identification even when overall appearance is similar to distracters
Solution Approach 2:
The patent inverts the traditional approach by not only searching for regions similar to the target but also actively identifying and excluding distracters. By flipping the problem from 'find what matches' to 'find what doesn't match,' the system improves precision by eliminating false positives through distracter detection and supporter verification
3Productivity
If the target leaves the field of view, then the tracker can reduce computational load, but tracking continuity deteriorates and reacquisition becomes difficult
Solution Approach 1:
The patent performs preliminary actions by maintaining appearance models and contextual information about the target even when it leaves the field of view. The system prepares for potential reacquisition by keeping the target's visual characteristics stored and updated, so when the target reappears, the tracker can quickly resume tracking without needing to relearn the target's appearance
Solution Approach 2:
The patent implements feedback mechanisms that continuously update the appearance model and contextual descriptors based on observed target behavior and environmental patterns. This feedback loop ensures that even during periods when the target is absent, the system maintains an accurate representation of the target, improving both processing efficiency during absence and reacquisition capability upon return
Data Source
AI summary
The present disclosure describes systems and techniques relating to identifying and tracking objects in images, such as visual tracking in video images in unconstrained environments. According to an aspect, a system includes one or more processors, and computer-readable media configured and arranged to cause the one or more processors to: identify an object in a first image of a sequence of images, identifying one or more regions similar to the object in the first image of the sequence of images, identifying one or more features around the object in the first image of the sequence of images, preventing drift in detection of the object in a second image of the sequence of images based on the one or more regions similar to the object, and verifying the object in the second image of the sequence of images based on the one or more features.


