Remote Camera Object Tracking Using Pre-Captured Visual Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems struggle to accurately select and track a specific object within a video frame when multiple objects are present, often leading to degraded image quality due to depth of field changes or incorrect object selection.
Innovation Solution
A tracking application on a device compares objects in the video frame with previously captured visual content items, using techniques like edge detection and face identification to identify and match objects, and employs signature generation to speed up processing, enabling accurate tracking and zooming on the desired object.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the device selects objects for tracking using known techniques, then multiple objects can be identified in the frame, but the system cannot determine which specific object to track, leading to incorrect selection and degraded image quality
Solution Approach 1:
The system performs preliminary actions by capturing visual content items (photos, videos) of the user before the live viewing session. These pre-captured items serve as reference data that is processed to generate signatures or feature representations. When tracking is needed, the system compares these pre-generated references with objects detected in the live frame, enabling rapid and accurate identification of the user's child without real-time delays or manual selection.
Solution Approach 2:
The system creates copies or representations of the user's child from pre-captured visual content items. These copies (in the form of processed image data, feature vectors, or signatures) are stored and used as templates for comparison. During live tracking, the system compares detected objects against these copied representations to identify matches, eliminating the need to analyze and select from all objects in real-time.
2Ease of operation
If the system tracks the wrong object due to incorrect selection, then depth of field changes occur, but this leads to degraded image quality and poor user experience
Solution Approach 1:
The system implements feedback by continuously comparing detected objects in the live frame against pre-captured reference content. The comparison process provides feedback on whether the detected object matches the user's child, allowing the system to confirm correct tracking or adjust if a mismatch is detected. This feedback loop ensures that tracking decisions are based on verified matches rather than incorrect assumptions.
3Reliability
If manual selection is used to identify the correct object, then tracking accuracy improves, but this requires user intervention and reduces automation
Solution Approach 1:
The system performs preliminary actions by capturing visual content items (photos, videos) of the user before the live viewing session. These pre-captured items serve as reference data that is processed to generate signatures or feature representations. When tracking is needed, the system compares these pre-generated references with objects detected in the live frame, enabling rapid and accurate identification of the user's child without real-time delays or manual selection.
Solution Approach 2:
The system enables self-service by automatically performing the entire tracking identification process without requiring user intervention. The pre-captured visual content items serve as the user's digital representation, and the system autonomously compares detected objects against these references to identify and track the correct subject. This eliminates the need for manual object selection while maintaining high tracking accuracy.
Data Source
AI summary
Systems and methods are disclosed for determining which of the multitude of objects within a feed being received from a remote camera to track. Specifically, objects within an image feed received from a remote camera are detected and compared with objects in visual content items captured by the user's device (e.g., pictures/videos captured by the smart phone or the electronic tablet). If a match is found between an object within the feed of the video (e.g., a person) and an object within visual content items captured on the user's device (e.g., the same person), the system will proceed to track the identified object.


