Eyewear Audio Source Separation Using Pose Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing electronic eyewear devices struggle with distinguishing and separating audio signals from multiple users in an environment, leading to confusion and poor user experience.

Innovation Solution

The electronic eyewear device employs a microphone array and aligns its trajectories with remote electronic eyewear devices to simplify audio source separation. By tracking the location of moving remote devices or objects, such as a user's face, the device uses the known location of remote users to facilitate effective audio source separation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If audio signals from multiple users are captured in a multi-user environment, then the device can record comprehensive environmental audio, but the audio signals from different sources become confused and difficult to distinguish

Engineering Contradiction:
Improveaudio signal coverageVSAvoidaudio source distinguishability
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent segments the mixed audio signal into distinct sources by using a microphone array to capture spatial information and applying beamforming techniques. The audio processing system divides the composite audio stream into separate channels corresponding to different users based on their spatial locations, allowing each audio source to be independently processed and identified.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces pose trackers and location data as intermediary elements that bridge the gap between mixed audio signals and source identification. By incorporating visual tracking data from cameras and pose estimation algorithms, the system creates an intermediary mapping between spatial positions and audio sources, enabling accurate attribution of each audio signal to its corresponding user.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If the device uses traditional audio processing methods, then the system complexity remains low, but the ability to separate and identify audio sources from multiple users is insufficient

Engineering Contradiction:
Improveaudio source separation accuracyVSAvoidaudio processing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges multiple data streams including audio signals from the microphone array, visual data from cameras, pose tracking information, and location data into a unified processing framework. By combining these diverse inputs, the system achieves accurate audio source separation through multi-modal fusion, where each data type compensates for the limitations of others and collectively enables precise source identification.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transitions from traditional single-dimensional audio processing to multi-dimensional processing by incorporating spatial coordinates, orientation data, and temporal information. The audio source separation is achieved by analyzing signals across multiple dimensions including horizontal and vertical angles, distance, and time, thereby creating a comprehensive spatial audio model that accurately distinguishes between multiple sources.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If the device tracks and processes location data of remote users, then audio source separation is facilitated, but the computational requirements and processing time increase

Engineering Contradiction:
Improveaudio source localization accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary tracking of user locations and poses continuously in the background using cameras and pose estimation algorithms, maintaining an up-to-date spatial map of all users in the environment. By pre-processing and continuously updating location data before audio separation is needed, the system reduces the computational burden during actual audio processing, as the spatial configuration is already established and ready for rapid audio source attribution.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250175741A1Eyewear with audio source separation using pose trackers
Publication Date: 2025.05.29 SNAP INC
  • US20250175741A1 patent drawing
  • US20250175741A1 patent drawing
  • US20250175741A1 patent drawing

AI summary

Electronic eyewear device providing simplified audio source separation, also referred to as voice/sound unmixing, using alignment between respective device trajectories. Multiple users of electronic eyewear devices in an environment may simultaneously generate audio signals (e.g., voices/sounds) that are difficult to distinguish from one another. The electronic eyewear device tracks the location of moving remote electronic eyewear devices of other users, or an object of the other users, such as the remote user's face, to provide audio source separation using location of the sound sources. The simplified voice unmixing uses a microphone array of the electronic eyewear device and the known location of the remote user's electronic eyewear device with respect to the user's electronic eyewear device to facilitate audio source separation.