Hybrid Vehicle Face Tracking for Low-Latency AR HUD Pose Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional face tracking systems for vehicles face challenges in achieving accurate and low-latency face pose determination for augmented reality and autostereoscopic heads-up displays, while also dealing with high cost, complexity, and large footprint due to the use of multiple high-speed cameras and rigid stereo rigs.
Innovation Solution
A hybrid face tracking system that correlates facial landmark features identified in images captured at different frame rates by an in-cabin camera and a monocular camera, using scale factor adjustment, correlation during pose stability, and compensation for camera shifts, to determine accurate face poses with reduced complexity and cost.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If stereo disparity techniques with multiple high-speed cameras are used, then measurement precision of face pose is improved, but device complexity increases
Solution Approach 1:
The system divides the face tracking task into two segments: low-frame-rate stereoscopic capture for accurate depth measurement and high-frame-rate monocular capture for temporal resolution. By processing these segments differently and fusing their results, the system achieves both measurement precision and temporal responsiveness without requiring all cameras to operate at high speeds simultaneously.
Solution Approach 2:
The patent introduces an intermediary processing stage that fuses data from the stereoscopic camera pair and the monocular camera. This intermediary fusion process combines the accurate but low-temporal-resolution stereoscopic measurements with the high-temporal-resolution monocular measurements, producing a final face pose estimate that benefits from both data sources without directly requiring complex real-time stereo processing at high frame rates.
2Measurement precision
If multiple high-speed cameras are used, then measurement precision is improved, but bill of materials cost increases
Solution Approach 1:
The system segments the camera roles by function and speed requirements: stereoscopic cameras operate at lower frame rates where cost-effective sensors suffice, while the monocular camera operates at high frame rates for temporal tracking. This segmentation allows using fewer expensive high-speed cameras while maintaining overall system performance.
Solution Approach 2:
The patent changes the operational parameters of different camera types: stereoscopic cameras use lower frame rates (reducing cost requirements) while the monocular camera uses higher frame rates. This parameter differentiation allows the system to achieve high temporal resolution without requiring all cameras to be expensive high-speed models, thereby reducing the overall bill of materials cost.
3Measurement precision
If stereo camera rig is used, then measurement precision is improved, but area occupied increases
Solution Approach 1:
The patent implements a nested camera configuration where the monocular camera is positioned within or alongside the stereoscopic camera rig structure. This nesting allows the high-frame-rate monocular sensor to share the physical mounting space and optical path infrastructure of the stereoscopic system, reducing the overall footprint while maintaining both depth measurement capability and high temporal resolution.
Data Source
AI summary
At least one monocular camera is installed in a vehicle in addition to in-cabin camera(s). Images captured by the in-cabin camera(s) at a first frame rate are processed to determine first positions of facial landmark features of a user and first poses of the user's face. Images captured by the at least one monocular camera at a second frame rate, higher than the first frame rate, are processed to identify facial landmark features of the user in the second images. Second positions of the facial landmark features are determined at the second frame rate, by correlating the facial landmark features identified in the second images with the first positions of the facial landmark features determined from the first images. Second poses of the user's face are determined at the second frame rate based on the second positions of the facial landmark features and the first poses of the user's face.

