Headset Audio System Using Eye Tracking for Persistent Sound Source Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio enhancement technologies fail to effectively isolate and enhance sound from a specific sound source of interest, as they require the user to continuously gaze or point at the source, which is not compatible with natural human behavior, leading to poor sound intelligibility in environments with multiple sound sources.
Innovation Solution
An audio system on a headset identifies and ranks sound sources based on eye tracking information, selectively applying filters to enhance or attenuate sound signals from these sources, allowing the primary sound source to remain of interest even if the user looks away, thereby providing a persistent ranking and improved listening experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If audio enhancement systems use gaze tracking to identify sound sources of interest, then sound intelligibility is improved, but the system requires continuous fixation on the sound source which conflicts with natural listener behavior
Solution Approach 1:
The system performs preliminary gaze tracking to identify and rank sound sources of interest before audio enhancement begins. Once a sound source is identified as interesting through gaze direction, the system maintains enhancement for that source even when the user looks away, based on the preliminary identification. This resolves the contradiction by establishing the sound source of interest in advance rather than requiring continuous verification through gaze.
Solution Approach 2:
The system dynamically adjusts the persistence of sound source ranking based on gaze behavior patterns. When the user's gaze indicates interest in a particular sound source, the system increases the persistence threshold, allowing that source to remain enhanced for longer periods even during gaze shifts. This dynamic adjustment allows the system to adapt to natural listening behaviors while maintaining sound intelligibility for identified sources of interest.
2Speed
If the audio system continuously updates sound source ranking based on eye tracking, then responsiveness to user interest is improved, but system complexity and processing requirements increase
Solution Approach 1:
Instead of continuously updating sound source rankings at every moment, the system implements periodic updates based on gaze dwell time and significant gaze direction changes. The ranking is updated periodically when the user's gaze indicates sustained interest in a particular sound source or when there is a significant change in gaze direction. This periodic action maintains responsiveness to user interest while reducing the processing burden and system complexity compared to continuous real-time updates.
3Quantity of substance
If the system enhances sound from multiple sound sources simultaneously, then audio content richness is improved, but difficulty in discerning particular speakers increases
Solution Approach 1:
The system applies different enhancement levels to different sound sources based on their ranked importance. The highest-ranked sound source receives the greatest enhancement, while lower-ranked sources receive progressively less enhancement or suppression. This local quality approach allows multiple sound sources to be present in the audio output, maintaining audio richness, while simultaneously ensuring that the primary sound source of interest remains clearly discernible through differential enhancement.
Data Source
AI summary
A system that uses persistent sound source selection to augment audio content. The system comprises one or more microphones coupled to a frame of a headset. The one or more microphones capture sound emitted by sound sources in a local area. The system further comprises an audio controller integrated into the headset. The audio controller receives sound signals corresponding to sounds emitted by sound sources in the local area. The audio controller further updates a ranking of the sound sources based on eye tracking information of the user. The audio controller further selectively applies one or more filters to the one or more of the sound signals according to the ranking to generate augmented audio data. The audio controller further provides the augmented audio data to a speaker assembly for presentation to the user.


