Audio Upmixing via Frequency Domain Prototype Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio upmixing technologies face challenges in accurately rendering spatially separated audio channels from a multichannel source while minimizing sonic artifacts and processing latency.
Innovation Solution
The method involves synthesizing prototype signals based on statistical characteristics of input signals, forming output signals as weighted combinations of these prototypes, and using nonlinear processing techniques to estimate the output signals, which allows for flexible temporal and frequency processing with reduced artifacts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If time domain upmixing is used to create weighted combinations of stereo input channels, then a single source can be rendered in a desired location, but multiple simultaneous sources cannot be isolated and weaker panned content is mixed into the center output channel
Solution Approach 1:
The patent segments the audio signal processing into frequency-specific channels using a filter bank. Each frequency channel processes sources independently, allowing multiple simultaneous sources to be isolated without mixing them into the center channel. This segmentation resolves the limitation of time domain upmixing where all channels are processed uniformly.
Solution Approach 2:
The patent applies different processing characteristics to different frequency regions. Each frequency channel can have its own weighting and processing parameters, allowing optimal separation of sources at different frequencies. This local quality approach enables better isolation of multiple sources compared to uniform time domain processing.
2Measurement precision
If complex upmixing algorithms are used to separate center channels, side-only channels, and surround channels, then spatial separation is improved, but processing latency increases
Solution Approach 1:
The patent performs preliminary frequency decomposition using a filter bank before channel separation processing. By organizing signals into frequency channels in advance, the subsequent separation operations become simpler and faster, reducing overall processing latency while maintaining high separation accuracy.
3Device complexity
If time domain upmixing creates weighted combinations of input channels, then processing is simple, but sonic artifacts are introduced and spatial separation is poor
Solution Approach 1:
The patent transitions from time domain processing to frequency domain processing by introducing a frequency dimension through filter bank decomposition. This dimensional change enables better spatial separation and artifact reduction while maintaining computational efficiency, as frequency channel operations are independent and can be processed in parallel.
Data Source
AI summary
An approach to forming output signals both permits flexible and temporally and/or frequency local processing of input signals while limiting or mitigating artifacts in such output signals. Generally, the approach involves first synthesizing prototype signals for the output signals, or equivalently characterizing such prototypes, for example, according to their statistical characteristics, and then forming the output signals as estimates of the prototype signals, for example, as weighted combinations of the input signals.


