Extended Reality Audio Rendering Interface for Listening Position Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users in augmented, virtual, or mixed reality systems struggle to select a desired listening position, leading to suboptimal audio experiences in immersive environments.

Innovation Solution

A user interface that allows users to indicate a desired listening position, enabling the device to select and render audio streams accordingly, using ambisonic coefficients for accurate 3D sound localization and dynamic adaptation to user movements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If automatic audio stream selection is implemented without user input, then device complexity is reduced, but audio immersion and user satisfaction deteriorate

Engineering Contradiction:
Improveaudio processing complexityVSAvoidaudio immersion quality
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The system automatically detects user presence in the visual scene and determines listening position without requiring explicit user input. The audio processing unit monitors video feed, identifies when the user appears in the scene, and automatically selects appropriate audio streams based on the determined listening position, allowing the system to serve itself rather than requiring continuous user commands

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system continuously monitors the visual scene to detect user presence and adjusts audio stream selection based on real-time detection results. This feedback loop ensures that audio immersion quality is maintained by adapting to user position and presence dynamically, resolving the contradiction between automatic operation and audio quality

Inventive Principle:
Principle #23Feedback

2Reliability

If user interface for listening position selection is added, then audio immersion is improved, but ease of operation deteriorates

Engineering Contradiction:
Improveaudio immersion qualityVSAvoiduser interaction complexity
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system automatically determines listening position by detecting user presence in the visual scene rather than requiring users to manually select positions. The audio processing unit autonomously identifies when the user appears in the scene and selects appropriate audio streams, eliminating the need for complex user interfaces while maintaining high audio immersion quality

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Instead of requiring users to actively select listening positions through interfaces, the system inverts the approach by automatically detecting user position from video input and configuring audio accordingly. This passive detection approach improves ease of operation while maintaining audio immersion

Inventive Principle:
Principle #13The other way round (Inversion)

3Reliability

If multiple audio streams are processed and rendered dynamically, then audio immersion is improved, but use of energy increases

Engineering Contradiction:
Improveaudio immersion qualityVSAvoidpower consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system dynamically adjusts audio stream rendering based on detected user presence and position. When the user is detected in the visual scene, the audio processing unit activates dynamic multi-stream rendering with spatial audio effects. When the user is not detected or conditions change, the system reduces processing to essential audio streams, optimizing energy consumption while maintaining immersion quality when needed

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes processing parameters based on user presence detection. When the user is present in the visual scene, the system enables full multi-stream processing with spatial effects. When conditions change, processing intensity is reduced, allowing energy consumption to adapt to actual immersion requirements rather than running at maximum capacity continuously

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP3994565B1User interface for controlling audio rendering for extended reality experiences
Publication Date: 2025.08.13 QUALCOMM INC
  • EP3994565B1 patent drawingFigure 1A
  • EP3994565B1 patent drawingFigure 1B
  • EP3994565B1 patent drawingFigure 1C

AI summary

A device may be configured to play one or more of a plurality of audio streams. The device may include a memory configured to store the plurality of audio streams, each of the audio streams representative of a soundfield. The device also may include one or more processors coupled to the memory, and configured to present a user interface to a user, obtain an indication from a user via the user interface representing a desired listening position; and select, based on the indication, at least one audio stream of the plurality of audio streams.