Audio Signal Upmixing via Direct Diffuse Decomposition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional upmixing algorithms for audio signals often fail to create a strong enough spatial immersive experience, particularly when upmixing from stereo to surround 5.1 or 7.1 formats, as they primarily upmix diffuse signals, leaving direct signals like rain or bird chirps in floor speakers, which are not naturally suited for these formats, leading to audible artifacts.
Innovation Solution
The method involves decomposing audio signals into diffuse and direct components, extracting audio objects from the direct signals, estimating their metadata, including height information, and rendering these objects and audio beds with height channels to predefined speaker positions, adjusting complexity-based gain and metadata to enhance spatial immersion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If conventional upmixing algorithms only upmix diffuse signals to height speakers, then the upmixing process is simple, but the spatial immersive experience is insufficient and audible artifacts occur
Solution Approach 1:
The audio signal is segmented into direct signals and diffuse signals through decomposition. Direct signals (e.g., rain, thunder, bird chirps) are extracted and processed separately from diffuse signals, allowing each component to be upmixed appropriately. This segmentation enables direct signals to be rendered to height speakers while maintaining their spatial characteristics, thereby improving spatial immersion without causing artifacts.
Solution Approach 2:
Different processing strategies are applied to different components of the audio signal. Direct signals receive object-based processing with metadata estimation for precise spatial placement, while diffuse signals receive traditional channel-based upmixing. This local quality approach ensures that each signal type is handled according to its characteristics, improving overall spatial immersion.
2Ease of manufacture
If direct signals are rendered to floor speakers only, then the processing is straightforward, but natural overhead sounds lose their spatial accuracy
Solution Approach 1:
Direct signals are extracted from the mixed audio signal using direct/diffuse decomposition. This extraction isolates overhead sounds (rain, thunder, bird chirps) from the general audio mixture, enabling them to be independently processed and rendered to appropriate speakers including height speakers, thereby preserving their spatial accuracy.
Solution Approach 2:
Metadata estimation acts as an intermediary process between direct signal extraction and final rendering. The metadata (including spatial position, height, and movement information) serves as a mediator that guides the rendering of direct signals to the correct speakers, ensuring spatial accuracy while maintaining processing efficiency.
3Measurement precision
If all audio objects are extracted and rendered individually, then spatial precision is maximized, but system complexity increases
Solution Approach 1:
The system dynamically adapts its processing based on the characteristics of the input signal. The direct/diffuse decomposition ratio, object extraction intensity, and metadata estimation are adjusted according to the signal content and listening environment. This dynamic approach maintains high spatial precision while managing system complexity through adaptive processing.
Solution Approach 2:
The system changes key parameters such as the decomposition threshold, object extraction sensitivity, and rendering gain based on signal characteristics. By adjusting these parameters dynamically, the system achieves high spatial precision for complex signals while reducing processing complexity for simpler signals.
Data Source
Figure 1~2
Figure 3~4
Figure 5~6
AI summary
Example embodiments disclosed herein relates to upmixing of audio signals. A method of upmixing an audio signal is described. The method includes decomposing the audio signal into a diffuse signal and a direct signal, generating an audio bed at least in part based on the diffuse signal, the audio bed including a height channel, extracting an audio object from the direct signal, estimating metadata of the audio object, the metadata including height information of the audio object; and rendering the audio bed and the audio object as an upmixed audio signal, wherein the audio bed is rendered to a predefined position and the audio object is rendered according to the metadata. Corresponding system and computer program product are described as well.