2D to 3D Audio Conversion via HOA and Channel Objects

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

There is no simple and effective method to convert existing 2D audio content into high-quality 3D audio, particularly in formats like MPEG-H 3D Audio, which limits the creation of immersive audio experiences without remixing sound objects and lacks control over spatial distribution.

Innovation Solution

A method and apparatus that generate 3D sound representations by converting multi-channel 2D audio inputs into Higher Order Ambisonics (HOA) representations and channel object signals, allowing for improved spatial distribution and playback on various loudspeaker setups, using techniques like scaling, decorrelating, and converting signals to HOA format with predetermined spatial positions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing 2D audio content is converted to 3D audio format, then spatial distribution and height information are improved, but the conversion process lacks control and quality is insufficient

Engineering Contradiction:
Improvespatial distribution capabilityVSAvoidcontrol over conversion process
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent segments the 2D audio signal into multiple channel components and processes them separately through different transfer functions. Each channel can be independently positioned in 3D space and assigned to specific loudspeaker positions, providing both spatial distribution and control. The segmentation allows selective manipulation of individual channels while maintaining overall coherence.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds the vertical dimension (height information) to traditional 2D audio by introducing elevated loudspeaker positions and using transfer functions that account for three-dimensional spatial relationships. This transforms planar audio distribution into volumetric sound fields, enabling height perception while maintaining control through predefined spatial positions.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If 2D audio is converted to 3D audio with HOA representation, then spatial impression and height information are improved, but device complexity increases

Engineering Contradiction:
Improvespatial impression qualityVSAvoidconversion process complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent employs pre-calculated transfer functions that encode the spatial transformation from 2D to 3D audio. These transfer functions are prepared in advance for various loudspeaker configurations, eliminating the need for complex real-time calculations during playback. The preliminary preparation of spatial transformation data simplifies the actual conversion process while maintaining high spatial impression quality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transforms the audio signal by changing spatial parameters through transfer functions that map 2D channel positions to 3D loudspeaker positions. By parameterizing the spatial transformation through predefined transfer functions, the system achieves complex 3D spatial effects without requiring complex processing algorithms, thus reducing device complexity while improving spatial impression.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If channel-based methods are used for 3D audio, then specific loudspeaker setups are optimized, but adaptability to different loudspeaker configurations is limited

Engineering Contradiction:
Improvespatial positioning accuracyVSAvoidloudspeaker setup compatibility
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal 3D audio system that can adapt to various loudspeaker configurations through parameterizable transfer functions. The same processing framework works with different numbers and arrangements of loudspeakers by adjusting the transfer function parameters, making the system multi-functional and broadly applicable while maintaining precise spatial positioning through the mathematical relationships in the transfer functions.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3375208B1Method and apparatus for generating from a multi-channel 2d audio input signal a 3D sound representation signal
Publication Date: 2019.11.06 DOLBY INTERNATIONAL AB
  • EP3375208B1 patent drawingFigure 1
  • EP3375208B1 patent drawingFigure 2~3
  • EP3375208B1 patent drawingFigure 4

AI summary

Currently there is no simple and satisfying way to create 3D audio from existing 2D content. The conversion from 2D to 3D sound should spatially redistribute the sound from existing channels. From a multi-channel 2D audio input signal (x(k)(t)) a 3D sound representation is generated which includes an HOA representation Formula (I) and channel object signals Formula (II) scaled from channels of the 2D audio input signal. Additional signals Formula (III) placed in the 3D space are generated by scaling (21, 222; 41, 422; Formula (IV)) channels from the 2D audio input signal and by decorrelating (24, 25; 44, 45, 451; Formula (V)) a scaled version of a mix of channels from the 2D audio input signal, whereby spatial positions for the additional signals are predetermined. The additional signals Formula (III) are converted (27; 47) to a HOA representation Formula (I).