Egocentric Tracking for Scalable User Localization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for user localization in smart spaces face scalability issues when tracking a large number of participants, often requiring exponentially more sensors and experiencing occlusion problems.
Innovation Solution
The use of egocentric tracking with camera-equipped head-mounted devices (HMDs) and external visual markers allows for scalable user localization, leveraging passive identifiers and inertial measurement units to capture visual information and determine user attention, even in large spaces and crowded environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional sensor arrays are used for user localization, then localization accuracy is achieved, but system scalability deteriorates and occlusion problems increase
Solution Approach 1:
Instead of using active sensors to track users, the system inverts the approach by using passive visual markers that users wear to be tracked by fixed cameras. This inversion eliminates the need for exponential sensor proliferation and resolves scalability issues while maintaining localization accuracy.
Solution Approach 2:
The system creates visual copies of users through markers that replicate user identity and position information. These markers serve as information carriers that can be detected by cameras from a distance, enabling scalable tracking without requiring direct line-of-sight or complex sensor arrays.
2Measurement precision
If active sensors are used for tracking users, then user localization is achieved, but device complexity increases exponentially
Solution Approach 1:
The system extracts the tracking function from complex active sensors and relocates it to simple passive visual markers. This extraction separates the identification function from the sensing function, reducing device complexity while maintaining localization capability through distributed visual markers instead of centralized sensor arrays.
Solution Approach 2:
The system replaces expensive, complex active sensors with inexpensive passive visual markers that can be easily attached to users. These simple markers serve as disposable tracking elements that reduce overall system complexity while enabling continuous user localization throughout the smart space.
3Reliability
If extensive sensor arrays are deployed, then tracking capability is improved, but occlusion problems worsen
Solution Approach 1:
The system transitions from three-dimensional active sensor tracking to two-dimensional visual marker detection on user surfaces. This dimensional change allows markers to be detected from various angles and distances, reducing occlusion effects by providing multiple detection perspectives through the visual field rather than requiring omnidirectional sensor coverage.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enables efficient and scalable user localization, reducing the need for extensive sensor instrumentation and minimizing occlusion, allowing for accurate tracking of multiple users and measurement of meeting events, conversation metrics, and personalized viewing activities.
Implementation Method 1
capturing the visual information from a viewer's field-of-view
Implementation Method 2
combining the data from accelerometer and gyroscope
Implementation Method 3
Visual markers, also referred to as markers herein, are widely used in Augmented Reality (AR) systems
Data Source
AI summary
A system to determine viewer attention to presented content. Markers are applied to a presentation. Head orientation information is received from a body mounted camera worn by a viewer and is used to determine a sequence of head orientations of the viewer. The sequence of head orientations is associated with an identifier of the viewer and a corresponding sequence of first time stamps. A sequence of images is captured by the body mounted camera worn by the viewer. The sequence of images is associated with an identifier of the viewer and the body mounted camera and a corresponding sequence of second time stamps. Respective members of the sequences of images and head orientations are associated, and the presentation content is identified by evaluating information from visible markers in the captured sequence of images. Viewer attention to different elements can then be determined.


