3D Audio Synthesis from Limited-Channel Signals

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio playback systems, such as headphones, fail to replicate the immersive surround sound experience of multi-channel audio mixes when limited to stereo channels, as they lack the motion-related information and directional cues present in original audio content, leading to an incomplete psycho-acoustic experience.

Innovation Solution

A method that identifies spectral components undergoing panning effects in multi-channel audio signals, generates virtual channels to mimic intermediate audio channels, and applies directional filtration to recreate the panning effect in a reduced set of output audio signals, effectively up-mixing and down-mixing audio to preserve spatial motion cues in a limited-channel setup.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multi-channel audio signals are played back on stereo devices, then the number of output channels is reduced, but the immersive sensation and spatial motion cues are lost

Engineering Contradiction:
Improvecompatibility with stereo devicesVSAvoidspatial motion cues
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

Virtual channels are introduced as intermediary elements between the original multi-channel audio and the final stereo output. These virtual channels carry intermediate panning information that is gradually mixed down to stereo, preserving spatial cues that would otherwise be lost in direct down-mixing

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms the audio signal from a reduced dimensional stereo output back into a higher dimensional representation by generating virtual channels. This creates an extended channel space that preserves 3D spatial information even when the final output is limited to two channels

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If virtual channels are generated to preserve panning effects, then the spatial motion cues are improved, but the processing complexity increases

Engineering Contradiction:
Improvepanning informationVSAvoidprocessing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The audio processing is segmented into distinct spectral bands using spectrogram analysis. By dividing the frequency spectrum into multiple bands and processing panning information separately in each band, the system manages complexity through modular frequency-domain processing rather than time-domain complexity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces traditional time-domain audio processing with frequency-domain processing using spectrograms and amplitude functions. This substitution enables more efficient detection of panning effects and generation of virtual channels through mathematical operations on spectral components

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11503419B2Detection of audio panning and synthesis of 3D audio from limited-channel surround sound
Publication Date: 2022.11.15 SPHEREO SOUND LTD
  • US11503419B2 patent drawing
  • US11503419B2 patent drawing
  • US11503419B2 patent drawing

AI summary

A method includes receiving a multi-channel audio signal (101) including multiple input audio channels (102, 104, 106, 108) that are configured to play audio from multiple respective locations relative to a listener. One or more spectral components that undergo a panning effect (1001, 1002, 1003) are identified in the multi-channel audio signal among at least some of the input audio channels. One or more virtual channels (1100, 1200, 1300) are generated, which together with the input audio channels form an extended set (111) of audio channels that retain the identified panning effect. A reduced set (222) of output audio signals, fewer in number than the input audio signals, is generated from the extended set, including recreating the panning effect in the output audio signals. The reduced set of output audio signals is outputted to a user.