Audio Capture Device Gaze-Based Speech Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio systems struggle to intelligibly capture and playback speech in a user's environment, particularly when the user is wearing a playback device and the sound source is not directly facing them, as they fail to effectively filter out background noise and focus on the intended sound source based on the user's head pose.

Innovation Solution

An audio capture device, separate from the playback device, uses a machine learning model to process microphone signals and determine the user's gaze direction, extracting speech from the intended area of interest while filtering out other sounds, and sends this extracted speech to the playback device for enhanced audio playback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the playback device alone is used to capture audio, then the device structure is simple, but the speech intelligibility is poor due to inability to filter background noise effectively

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidsystem structure
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines the playback device worn by the user with a separate audio capture device (companion device) to create a collaborative audio system. The companion device captures environmental audio with its microphones while the playback device tracks head pose, and their data are merged through wireless communication to achieve both good speech intelligibility and reasonable system structure.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces a machine learning model as an intermediary component that processes the raw microphone signals from the companion device and the head pose data from the playback device. This intermediary processes the combined data to extract and enhance speech from the user's field of view, significantly improving speech intelligibility by filtering out background noise.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If the audio capture device processes all microphone signals equally, then the processing is simple, but the ability to focus on the user's intended sound source is poor

Engineering Contradiction:
Improvesound source localization accuracyVSAvoidsignal processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies local quality by directing the audio processing focus specifically to the region of the user's field of view rather than treating all directions equally. The machine learning model uses head pose information to identify the user's gaze direction and prioritizes extracting speech from that specific spatial region, improving sound source localization accuracy.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system performs preliminary action by tracking the user's head pose and determining the field of view before processing the audio signals. This pre-processing step of identifying the region of interest allows the machine learning model to focus computational resources on extracting speech from the relevant directional region, improving localization accuracy.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If the system captures all sounds in the environment, then no information is lost, but the intelligibility of the intended speech is reduced due to background noise

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidfiltering out non-speech sounds
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent extracts the useful speech signal from the mixed environmental audio by using the machine learning model to separate speech from background noise. The system takes out only the speech components that are spatially consistent with the user's field of view, improving speech intelligibility while discarding irrelevant background sounds.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system uses feedback by continuously monitoring the user's head pose and adjusting the audio extraction focus accordingly. As the user moves their head and changes their field of view, the machine learning model adapts to extract speech from the new directional region, maintaining high speech intelligibility dynamically.

Inventive Principle:
Principle #23Feedback

4Measurement precision

If the playback device is worn on the user's head, then spatial audio rendering is improved, but the device cannot independently determine the user's gaze direction

Engineering Contradiction:
Improvegaze direction determinationVSAvoiddevice configuration
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies universality by making the companion device serve multiple functions: it acts as an audio capture device with microphones, a gaze estimation device using its cameras to track the user's eyes, and a communication hub. This multi-functional design enables precise gaze direction determination without adding complexity to the wearable playback device.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12141347B1Machine learning and user driven selective hearing
Publication Date: 2024.11.12 APPLE INC
  • US12141347B1 patent drawing
  • US12141347B1 patent drawing
  • US12141347B1 patent drawing

AI summary

An audio processing device may generate a plurality of microphone signals from a plurality of microphones of the audio processing device. The audio processing device may determine a gaze of a user who is wearing a playback device that is separate from the audio processing device, the gaze of the user being determined relative to the audio processing device. The audio processing device may extract speech that correlates to the gaze of the user, from the plurality of microphone signals of the audio processing device by applying the plurality of microphone signals of the audio processing device and the gaze of the user to a machine learning model. The extracted speech may be played to the user through the playback device.