Near-Eye Audio Attenuation Using Speech Direction Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Near-eye display devices struggle to differentiate between ambient speech directed at the user and speech not intended for the user, leading to inappropriate volume attenuation in immersive audio-visual experiences.

Innovation Solution

The use of image sensor data, including depth and two-dimensional image data, combined with audio data and machine learning algorithms, to determine if speech is directed at the user, allowing for targeted attenuation of audio content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If ambient audio detection is used to trigger volume attenuation, then responsiveness to potential speech is improved, but false attenuation occurs when speech is not directed at the user

Engineering Contradiction:
Improveaccuracy of speech direction detectionVSAvoidcomplexity of sensor fusion system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines multiple sensor types (microphones for audio detection, image sensors for visual detection, depth sensors for spatial mapping) into a unified sensor fusion system. This merging allows the system to cross-validate data from different modalities to accurately determine whether speech is directed at the user, resolving the contradiction between reliability and false attenuation by using redundant information sources to filter out false positives.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediary processing layer that analyzes sensor data to determine speech direction. This intermediary system processes raw sensor inputs through machine learning models and signal processing algorithms to generate a reliable determination of whether speech is directed at the user, acting as a mediator between raw sensor data and the volume attenuation decision.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If volume attenuation is applied whenever speech is detected, then user awareness of ambient speech is improved, but immersion in the virtual experience is degraded

Engineering Contradiction:
Improveuser awareness of environmentVSAvoidappropriateness of attenuation
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent applies local quality by making the volume attenuation selective rather than global. Instead of attenuating all audio content uniformly, the system uses sensor fusion to identify the direction and target of speech, then applies attenuation only when speech is determined to be directed at the user. This localized application of attenuation maintains immersion for non-directed speech while providing awareness for directed speech.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements dynamic adjustment of audio volume based on real-time sensor fusion analysis. The system continuously monitors audio and visual inputs, dynamically determining whether to attenuate volume based on the current context of speech direction. This dynamic approach allows the system to adapt to changing environmental conditions and maintain appropriate immersion levels.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If multiple sensors are integrated to determine speech direction, then detection accuracy is improved, but power consumption increases

Engineering Contradiction:
Improvespeech direction detection precisionVSAvoidpower consumption of sensor subsystem
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent implements periodic action by activating the full sensor fusion system only when needed - specifically, when audio sensors detect potential speech events. During normal operation, the system uses lower-power modes with reduced sensor activation. When speech is detected, the system periodically activates image sensors, depth sensors, and processing units to perform direction determination, then returns to lower-power operation, thus balancing precision with power consumption.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentEP3465680B1Automatic audio attenuation on immersive display devices
Publication Date: 2020.09.09 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3465680B1 patent drawingFigure 1
  • EP3465680B1 patent drawingFigure 2
  • EP3465680B1 patent drawingFigure 3A

AI summary

Examples disclosed herein relate to controlling volume on an immersive display device. One example provides a near-eye display device comprising a sensor subsystem, a logic subsystem, and a storage subsystem storing instructions executable by the logic subsystem to receive image sensor data from the sensor subsystem, present content comprising a visual component and an auditory component, while presenting the content, detect via the image sensor data that speech is likely being directed at a wearer of the near-eye display device, and in response to detecting that speech is likely being directed at the wearer, attenuate an aspect of the auditory component.