Spatial Audio Rendering via Visual Object Size Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio processing technologies fail to provide a seamless and immersive audio experience that dynamically adjusts to the size and position of visual objects, leading to disorienting transitions and inconsistent audio feedback in immersive environments.

Innovation Solution

A computer-implemented method that determines virtual placements for virtual speakers based on the size and position of visual objects, using binaural audio and vector-base amplitude panning to spatially render audio channels, allowing for dynamic changes in speaker placement and format translation while maintaining acoustic energy consistency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If audio is presented through fixed surround sound loudspeakers, then audio format consistency is maintained, but adaptability to visual object changes is reduced

Engineering Contradiction:
Improveadaptability to visual object changesVSAvoidaudio processing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic speaker placement where virtual speakers are repositioned in the audio field based on the size and position of visual objects. Instead of fixed speaker positions, the system continuously adjusts speaker coordinates to match visual object characteristics, enabling audio to adapt dynamically to changing visual content while maintaining processing through standardized algorithms.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes audio parameters (speaker position, spacing, distribution) based on visual object parameters (size, position). When visual objects change size or position, the audio system modifies corresponding speaker parameters to maintain coherence between visual and audio elements, achieving adaptability through parameter mapping.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If virtual speakers are moved closer together for smaller visual objects, then audio-visual coherence is improved, but spatial accuracy may be compromised

Engineering Contradiction:
Improveaudio-visual coherenceVSAvoidspatial accuracy
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent applies different speaker spacing rules based on local conditions (visual object size). For small visual objects, speakers are positioned closer together to match the compact visual presentation. For large visual objects, speakers are distributed wider apart to match the expansive visual content. This local adaptation maintains audio-visual coherence without compromising overall spatial accuracy.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The speaker configuration dynamically adjusts to match visual object characteristics. The system transitions between different speaker arrangements based on real-time visual object size, maintaining coherence while preserving spatial relationships through continuous repositioning rather than fixed configurations.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If multiple audio formats are supported, then versatility is improved, but processing complexity increases

Engineering Contradiction:
Improveaudio format compatibilityVSAvoidformat translation complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal audio rendering system that can handle multiple audio formats (5.1, 6.1, 7.1, stereo, object-based audio) through a single virtual speaker framework. The system translates different source formats into a common virtual speaker representation, enabling one system to perform multiple format conversions without requiring separate processing paths for each format type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240098442A1Spatial Blending of Audio
Publication Date: 2024.03.21 APPLE INC
  • US20240098442A1 patent drawing
  • US20240098442A1 patent drawing
  • US20240098442A1 patent drawing

AI summary

An audio processing system may obtain a size of a visual object to present to a display. The audio processing system may determine a virtual placement for each of a plurality of virtual speakers at least based on the size of the visual object. Each of the plurality of virtual speakers may be spatially rendered at each virtual placement through binaural audio, for playback through head-worn speakers. Other aspects are also described and claimed.