Head-Mounted Camera ROI Switching for Low-Power Facial Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing head-mounted systems for tracking facial expressions, such as AR/VR headsets, face challenges in efficiently collecting and processing data over extended periods due to power constraints from battery-operated devices, particularly when untethered.
Innovation Solution
Implementing a system with an inward-facing head-mounted camera that dynamically adjusts its region of interest (ROI) based on facial movements, using discrete photosensors to optimize power usage by changing binning values and framerate, and incorporating machine learning models for accurate facial expression detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the camera captures images of the entire face at high resolution continuously, then the accuracy of facial expression tracking is improved, but the power consumption increases
Solution Approach 1:
The face is divided into multiple regions of interest (ROIs) based on facial landmarks and expression dynamics. The camera only captures images of these specific ROIs rather than the entire face, significantly reducing the amount of data processed while maintaining tracking accuracy for relevant facial movements.
Solution Approach 2:
Different regions of the face are assigned different levels of imaging quality and capture frequency based on their importance for expression detection. High-movement regions receive higher quality imaging while stable regions use lower resolution, optimizing power consumption across different facial zones.
2Speed
If the camera framerate is increased to capture rapid facial movements, then the detection speed of facial expressions is improved, but the power consumption increases
Solution Approach 1:
The camera framerate is dynamically adjusted based on detected facial movement intensity. When rapid movements are detected, the framerate increases to capture the dynamics accurately. When movements are slow or static, the framerate decreases to conserve power, creating an adaptive sampling rate that responds to actual expression dynamics.
Solution Approach 2:
Instead of continuous high-rate sampling, the system uses periodic imaging at variable intervals. The imaging frequency is modulated based on detected facial activity levels, capturing images at higher rates during active expression periods and at lower rates during stable periods, reducing overall power consumption while maintaining detection capability.
3Measurement precision
If the binning value is reduced to increase image resolution, then the accuracy of facial feature detection is improved, but the computational load increases
Solution Approach 1:
The binning value is dynamically changed based on the detected facial expression type and movement characteristics. For subtle expressions requiring high precision, the binning value is reduced to increase resolution. For more obvious expressions, a higher binning value suffices, reducing computational load while maintaining adequate detection accuracy.
Data Source
AI summary
Utilization of windowing to set a region of interest (ROI) of a camera used for tracking facial expressions. In one embodiment, a system includes an inward-facing head-mounted camera that captures images of a region on a user's head utilizing a sensor that supports changing of its ROI. The system also includes a computer that detects, in a first subset of the images, a first sub-region in which changes due to a first facial movement reach a first threshold and reads from the camera a first ROI that covers at least a portion of the first sub-region. The computer detects, in a second subset of the images, a second sub-region in which changes due to a second facial movement reach a second threshold, and then reads from the camera a second ROI that covers at least a portion of the second sub-region, with the first and second ROIs being different.


