Audio Capture Device Gaze-Based Speech Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio systems struggle to intelligibly capture and playback speech in a user's environment, particularly when the user is wearing a playback device and the sound source is not directly facing them, as they fail to effectively filter out background noise and focus on the intended sound source based on the user's head pose.
Innovation Solution
An audio capture device, separate from the playback device, uses a machine learning model to process microphone signals and determine the user's gaze direction, extracting speech from the intended area of interest while filtering out other sounds, and sends this extracted speech to the playback device for enhanced audio playback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the playback device alone is used to capture audio, then the device structure is simple, but the speech intelligibility is poor due to inability to filter background noise effectively
Solution Approach 1:
The patent combines the playback device worn by the user with a separate audio capture device (companion device) to create a collaborative audio system. The companion device captures environmental audio with its microphones while the playback device tracks head pose, and their data are merged through wireless communication to achieve both good speech intelligibility and reasonable system structure.
Solution Approach 2:
The patent introduces a machine learning model as an intermediary component that processes the raw microphone signals from the companion device and the head pose data from the playback device. This intermediary processes the combined data to extract and enhance speech from the user's field of view, significantly improving speech intelligibility by filtering out background noise.
2Measurement precision
If the audio capture device processes all microphone signals equally, then the processing is simple, but the ability to focus on the user's intended sound source is poor
Solution Approach 1:
The patent applies local quality by directing the audio processing focus specifically to the region of the user's field of view rather than treating all directions equally. The machine learning model uses head pose information to identify the user's gaze direction and prioritizes extracting speech from that specific spatial region, improving sound source localization accuracy.
Solution Approach 2:
The system performs preliminary action by tracking the user's head pose and determining the field of view before processing the audio signals. This pre-processing step of identifying the region of interest allows the machine learning model to focus computational resources on extracting speech from the relevant directional region, improving localization accuracy.
3Reliability
If the system captures all sounds in the environment, then no information is lost, but the intelligibility of the intended speech is reduced due to background noise
Solution Approach 1:
The patent extracts the useful speech signal from the mixed environmental audio by using the machine learning model to separate speech from background noise. The system takes out only the speech components that are spatially consistent with the user's field of view, improving speech intelligibility while discarding irrelevant background sounds.
Solution Approach 2:
The system uses feedback by continuously monitoring the user's head pose and adjusting the audio extraction focus accordingly. As the user moves their head and changes their field of view, the machine learning model adapts to extract speech from the new directional region, maintaining high speech intelligibility dynamically.
4Measurement precision
If the playback device is worn on the user's head, then spatial audio rendering is improved, but the device cannot independently determine the user's gaze direction
Solution Approach 1:
The patent applies universality by making the companion device serve multiple functions: it acts as an audio capture device with microphones, a gaze estimation device using its cameras to track the user's eyes, and a communication hub. This multi-functional design enables precise gaze direction determination without adding complexity to the wearable playback device.
Data Source
AI summary
An audio processing device may generate a plurality of microphone signals from a plurality of microphones of the audio processing device. The audio processing device may determine a gaze of a user who is wearing a playback device that is separate from the audio processing device, the gaze of the user being determined relative to the audio processing device. The audio processing device may extract speech that correlates to the gaze of the user, from the plurality of microphone signals of the audio processing device by applying the plurality of microphone signals of the audio processing device and the gaze of the user to a machine learning model. The extracted speech may be played to the user through the playback device.


