Head-Mounted Device Pointer Tracking for Low-Latency Content Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems for content access using head-mounted devices (HMDs) are limited by high latency and power consumption, requiring repetitive processes for text recognition and interaction, which are cumbersome and impractical for real-time interactions with visible content.
Innovation Solution
A head-mounted device configured to perform real-time head-tracking and pointer-tracking, capturing images of content and using machine learning techniques to map pointer locations within images, allowing for low-latency and low-power interactions with content items, even when the pointer occludes them.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional OCR processes are used for content recognition, then text recognition capability is provided, but high latency and power consumption occur
Solution Approach 1:
The system performs OCR processing on the first image (captured without pointer occlusion) to obtain recognition results in advance, before the pointer is introduced. This preliminary recognition allows the system to quickly map pointer locations to pre-recognized text, avoiding the need to perform OCR again when the pointer is present, thus reducing latency while maintaining recognition accuracy
Solution Approach 2:
The system captures a first image without the pointer and generates recognition results from it, then uses these results to identify content at pointer locations in a second image. This copying approach allows the system to reuse recognition data across multiple frames, reducing repeated processing and lowering latency
2Measurement precision
If conventional OCR processes are used for content recognition, then text recognition capability is provided, but high power consumption occurs
Solution Approach 1:
The system performs OCR processing on the first image to obtain recognition results in advance, before the pointer is introduced. This preliminary recognition allows the system to quickly map pointer locations to pre-recognized text, avoiding the need to perform OCR again when the pointer is present, thus reducing latency while maintaining recognition accuracy
Solution Approach 2:
The system captures a first image without the pointer and generates recognition results from it, then uses these results to identify content at pointer locations in a second image. This copying approach allows the system to reuse recognition data across multiple frames, reducing repeated processing and lowering latency
3Measurement precision
If repetitive image capture and OCR processes are performed, then content recognition is achieved, but the process becomes cumbersome and impractical
Solution Approach 1:
The system maintains continuous head-tracking and pointer-tracking processes that operate in parallel, allowing the user to interact with content naturally by simply pointing at it. The tracking processes continue ongoing without interruption, and the system continuously maps pointer locations to recognized text, providing a smooth and intuitive interaction experience without requiring repetitive capture or processing actions
Solution Approach 2:
The system introduces a coordinate system as an intermediary between the captured images and the recognized text. This coordinate system allows the system to map pointer locations in the second image to corresponding positions in the first image, enabling seamless content identification without requiring direct re-processing of the original image
Data Source
AI summary
A head-mounted device (HMD) can be configured to determine a request for recognizing at least one content item included within content framed within a display of the HMD. The HMD can be configured to initiate a head-tracking process that maintains a coordinate system with respect to the content, and a pointer-tracking process that tracks a pointer that is visible together with the content within the display. The HMD can be configured to capture a first image of the content and a second image of the content, the second image including the pointer. The HMD can be configured to map a location of the pointer within the second image to a corresponding image location within the first image, using the coordinate system, and provide the at least one content item from the corresponding image location.


