Spatialized Audio Rendering in Mixed Reality Headsets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies for presenting immersive audio in mixed reality environments often fail to accurately account for a user's surroundings and spatial movements, leading to inconsistencies that can detract from the immersive experience and even cause discomfort such as motion sickness.

Innovation Solution

A method that involves receiving an encoded audio stream, generating a decoded audio stream, and using sensors and application program inputs to create a spatialized audio stream based on the user's position and the position of virtual speakers, which is then presented through speakers on a wearable head device.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If virtual audio is presented without accounting for user surroundings and spatial movements, then device complexity is reduced, but auditory realism and user immersion deteriorate

Engineering Contradiction:
Improveaudio processing complexityVSAvoidauditory realism
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The system continuously receives sensor data about user head position, orientation, and movements, and uses this feedback to dynamically adjust the spatialization of virtual audio sources. The audio render service constantly updates audio playback parameters based on real-time sensor inputs, creating a closed-loop system that maintains auditory realism as the user moves through the mixed reality environment.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces an audio render service as an intermediary component between the application program and the audio output hardware. This service acts as a mediator that processes virtual audio sources, applies spatialization algorithms, incorporates sensor data, and generates the final spatialized audio stream for playback through the wearable device speakers.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If spatialized audio is generated based on sensor data and virtual speaker positions, then auditory immersion is improved, but processing time and computational resources increase

Engineering Contradiction:
Improveauditory immersionVSAvoidaudio processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary spatialization calculations by pre-determining the positions of virtual speakers in the mixed reality environment and establishing their spatial relationships. The audio render service is pre-configured with the spatial arrangement of virtual audio sources, allowing for faster real-time processing when sensor data arrives during audio playback.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If encoded audio streams are decoded and spatialized in real-time, then audio quality is improved, but processing complexity and energy consumption increase

Engineering Contradiction:
Improveaudio qualityVSAvoidprocessing energy
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The audio processing pipeline is segmented into distinct functional stages: an audio decode service that handles decoding of encoded audio streams, and an audio render service that handles spatialization and playback. This segmentation allows each service to specialize in its specific task, optimizing processing efficiency and energy usage while maintaining high audio quality through proper separation of decoding and spatialization operations.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250071502A1Immersive audio platform
Publication Date: 2025.02.27 MAGIC LEAP INC
  • US20250071502A1 patent drawing
  • US20250071502A1 patent drawing
  • US20250071502A1 patent drawing

AI summary

Disclosed herein are systems and methods for presenting audio content in mixed reality environments. A method may include receiving a first input from an application program; in response to receiving the first input, receiving, via a first service, an encoded audio stream; generating, via the first service, a decoded audio stream based on the encoded audio stream; receiving, via a second service, the decoded audio stream; receiving a second input from one or more sensors of a wearable head device; receiving, via the second service, a third input from the application program, wherein the third input corresponds to a position of one or more virtual speakers; generating, via the second service, a spatialized audio stream based on the decoded audio stream, the second input, and the third input; presenting, via one or more speakers of the wearable head device, the spatialized audio stream.