Near-Eye Audio Attenuation Using Speech Direction Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Near-eye display devices struggle to differentiate between ambient speech directed at the user and speech not intended for the user, leading to inappropriate volume attenuation in immersive audio-visual experiences.
Innovation Solution
The use of image sensor data, including depth and two-dimensional image data, combined with audio data and machine learning algorithms, to determine if speech is directed at the user, allowing for targeted attenuation of audio content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If ambient audio detection is used to trigger volume attenuation, then responsiveness to potential speech is improved, but false attenuation occurs when speech is not directed at the user
Solution Approach 1:
The patent combines multiple sensor types (microphones for audio detection, image sensors for visual detection, depth sensors for spatial mapping) into a unified sensor fusion system. This merging allows the system to cross-validate data from different modalities to accurately determine whether speech is directed at the user, resolving the contradiction between reliability and false attenuation by using redundant information sources to filter out false positives.
Solution Approach 2:
The patent introduces an intermediary processing layer that analyzes sensor data to determine speech direction. This intermediary system processes raw sensor inputs through machine learning models and signal processing algorithms to generate a reliable determination of whether speech is directed at the user, acting as a mediator between raw sensor data and the volume attenuation decision.
2Ease of operation
If volume attenuation is applied whenever speech is detected, then user awareness of ambient speech is improved, but immersion in the virtual experience is degraded
Solution Approach 1:
The patent applies local quality by making the volume attenuation selective rather than global. Instead of attenuating all audio content uniformly, the system uses sensor fusion to identify the direction and target of speech, then applies attenuation only when speech is determined to be directed at the user. This localized application of attenuation maintains immersion for non-directed speech while providing awareness for directed speech.
Solution Approach 2:
The patent implements dynamic adjustment of audio volume based on real-time sensor fusion analysis. The system continuously monitors audio and visual inputs, dynamically determining whether to attenuate volume based on the current context of speech direction. This dynamic approach allows the system to adapt to changing environmental conditions and maintain appropriate immersion levels.
3Measurement precision
If multiple sensors are integrated to determine speech direction, then detection accuracy is improved, but power consumption increases
Solution Approach 1:
The patent implements periodic action by activating the full sensor fusion system only when needed - specifically, when audio sensors detect potential speech events. During normal operation, the system uses lower-power modes with reduced sensor activation. When speech is detected, the system periodically activates image sensors, depth sensors, and processing units to perform direction determination, then returns to lower-power operation, thus balancing precision with power consumption.
Data Source
Figure 1
Figure 2
Figure 3A
AI summary
Examples disclosed herein relate to controlling volume on an immersive display device. One example provides a near-eye display device comprising a sensor subsystem, a logic subsystem, and a storage subsystem storing instructions executable by the logic subsystem to receive image sensor data from the sensor subsystem, present content comprising a visual component and an auditory component, while presenting the content, detect via the image sensor data that speech is likely being directed at a wearer of the near-eye display device, and in response to detecting that speech is likely being directed at the wearer, attenuate an aspect of the auditory component.