Emotion-Aware Ear Tracking for Spatial Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing camera-based head and ear tracking systems face challenges in accurately scaling images due to the lack of physical scale in camera images and the inability to account for user emotions, leading to variability in sound pressure levels and errors in spatial audio rendering.
Innovation Solution
The method involves acquiring images of a user, processing them to generate emotion-specific three-dimensional ear positions based on a 3D head geometry and identified emotions, and processing audio signals to generate processed audio signals based on these ear positions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If camera-based head tracking is used to correct audio for personal and/or near-field audio systems, then spatial audio rendering is improved, but measurement precision of ear positions deteriorates due to lack of physical scale in camera images
Solution Approach 1:
The patent introduces facial landmarks as intermediary reference points between the camera image and the ear position. By detecting landmarks such as eyes, nose, and mouth, and using their relative positions to infer head geometry, the system creates a chain of reference that bridges the gap between 2D camera images and 3D ear positions, improving measurement precision without requiring direct physical scale measurement
Solution Approach 2:
The patent transitions from 2D camera images to 3D ear position estimation by incorporating head geometry models and facial landmark relationships. This dimensional transformation allows the system to recover depth information and physical scale that is lost in 2D images, enabling accurate spatial audio rendering
2Device complexity
If presumptive dimensions such as average inter-ocular distance are used for image scaling, then device complexity is reduced, but manufacturing precision of scaling accuracy deteriorates due to significant variation between individuals
Solution Approach 1:
The system performs self-calibration by using the user's own facial landmarks as reference. Instead of relying on population averages, the system detects the individual user's facial features and uses their specific inter-landmark distances for scaling, making the scaling accuracy self-determined and user-specific without adding external calibration equipment
Solution Approach 2:
The patent performs facial landmark detection and head geometry estimation as preliminary steps before audio processing. By pre-establishing the user-specific scaling relationships based on detected facial features, the system prepares accurate reference data in advance, eliminating the need for complex real-time scaling calculations during audio rendering
3Speed
If traditional head tracking systems ignore user emotion, then processing speed is improved, but measurement precision of ear positions deteriorates due to facial gestures and manipulations
Solution Approach 1:
The system incorporates emotion detection as a feedback mechanism that monitors facial muscle movements and gestures. By detecting emotions such as smiling, frowning, or talking, the system can dynamically adjust the facial landmark relationships and compensate for pose changes, maintaining accurate ear position estimation even during emotional expressions
Solution Approach 2:
The patent makes the head geometry model dynamic by incorporating emotion states. Instead of using a static facial structure, the system adapts the facial landmark relationships based on detected emotions, allowing the model to account for real-time facial muscle movements and maintain measurement precision during dynamic expressions
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Techniques for image scaling based on emotion detection are described. In some embodiments, the techniques include acquiring one or more images of a user, processing the one more images to generate emotion-specific three-dimensional (3D) positions of ears of the user based on a 3D head geometry and an emotion of the user, where the emotion is identified based on the one or more images of the user, and processing one or more audio signals to generate one or more processed audio signals based on the three-dimensional positions of the ears. Further embodiments include systems and non-transitory computer-readable media that perform the steps of the method.