Eyewear Audio Source Separation Using Pose Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic eyewear devices struggle with distinguishing and separating audio signals from multiple users in an environment, leading to confusion and poor user experience.
Innovation Solution
The electronic eyewear device employs a microphone array and aligns its trajectories with remote electronic eyewear devices to simplify audio source separation. By tracking the location of moving remote devices or objects, such as a user's face, the device uses the known location of remote users to facilitate effective audio source separation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If audio signals from multiple users are captured in a multi-user environment, then the device can record comprehensive environmental audio, but the audio signals from different sources become confused and difficult to distinguish
Solution Approach 1:
The patent segments the mixed audio signal into distinct sources by using a microphone array to capture spatial information and applying beamforming techniques. The audio processing system divides the composite audio stream into separate channels corresponding to different users based on their spatial locations, allowing each audio source to be independently processed and identified.
Solution Approach 2:
The patent introduces pose trackers and location data as intermediary elements that bridge the gap between mixed audio signals and source identification. By incorporating visual tracking data from cameras and pose estimation algorithms, the system creates an intermediary mapping between spatial positions and audio sources, enabling accurate attribution of each audio signal to its corresponding user.
2Measurement precision
If the device uses traditional audio processing methods, then the system complexity remains low, but the ability to separate and identify audio sources from multiple users is insufficient
Solution Approach 1:
The patent merges multiple data streams including audio signals from the microphone array, visual data from cameras, pose tracking information, and location data into a unified processing framework. By combining these diverse inputs, the system achieves accurate audio source separation through multi-modal fusion, where each data type compensates for the limitations of others and collectively enables precise source identification.
Solution Approach 2:
The patent transitions from traditional single-dimensional audio processing to multi-dimensional processing by incorporating spatial coordinates, orientation data, and temporal information. The audio source separation is achieved by analyzing signals across multiple dimensions including horizontal and vertical angles, distance, and time, thereby creating a comprehensive spatial audio model that accurately distinguishes between multiple sources.
3Measurement precision
If the device tracks and processes location data of remote users, then audio source separation is facilitated, but the computational requirements and processing time increase
Solution Approach 1:
The patent performs preliminary tracking of user locations and poses continuously in the background using cameras and pose estimation algorithms, maintaining an up-to-date spatial map of all users in the environment. By pre-processing and continuously updating location data before audio separation is needed, the system reduces the computational burden during actual audio processing, as the spatial configuration is already established and ready for rapid audio source attribution.
Data Source
AI summary
Electronic eyewear device providing simplified audio source separation, also referred to as voice/sound unmixing, using alignment between respective device trajectories. Multiple users of electronic eyewear devices in an environment may simultaneously generate audio signals (e.g., voices/sounds) that are difficult to distinguish from one another. The electronic eyewear device tracks the location of moving remote electronic eyewear devices of other users, or an object of the other users, such as the remote user's face, to provide audio source separation using location of the sound sources. The simplified voice unmixing uses a microphone array of the electronic eyewear device and the known location of the remote user's electronic eyewear device with respect to the user's electronic eyewear device to facilitate audio source separation.


