Audio Upmixing via Downmix Modification and Mixing Matrix
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing parametric stereo and multichannel coding methods are inefficient in terms of computational resources and bandwidth, particularly in low-bitrate applications, and are prone to error propagation due to high structural complexity and limited processing power in portable devices.
Innovation Solution
The proposed method involves a spatial synthesis system with a downmix modifying processor and a mixing matrix that operates directly on the downmix signal, performing cross mixing and non-linear processing, allowing for parallel downmix and parameter extraction processes without intermediate information exchange, and uses low-degree polynomial gains to reduce error propagation and optimize resource usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If parametric coding methods are used to achieve excellent coding efficiency, then bandwidth efficiency is improved, but computational complexity increases
Solution Approach 1:
The audio signal is segmented into multiple channels (front left, front right, back left, back right, center, LFE) that are independently processed and then combined. This segmentation allows parallel processing of channel information, reducing the computational burden on sequential operations while maintaining efficient bandwidth utilization through independent channel encoding.
Solution Approach 2:
The mixing matrix is designed to be multi-functional, serving both as a downmixing matrix (combining multiple channels into fewer) and as an upmixing matrix (separating channels during decoding). This universal matrix structure eliminates the need for separate complex processing paths, reducing overall computational complexity while maintaining coding efficiency.
2Productivity
If intermediate buffers are used in parametric coding methods, then processing capability is improved, but device complexity increases
Solution Approach 1:
The patent extracts and removes the unnecessary intermediate buffering stage from the traditional parametric coding structure. By directly transforming channel information through the mixing matrix without storing intermediate results in buffers, the system maintains processing capability while significantly reducing structural complexity and memory requirements.
Solution Approach 2:
The mixing matrix is pre-configured with optimized transformation coefficients that perform multiple processing functions in a single operation. This preliminary preparation of the transformation matrix allows the system to achieve high processing capability without requiring complex intermediate buffering and multiple processing stages.
3Reliability
If error correction mechanisms are added to prevent error propagation, then reliability is improved, but computational complexity increases
Solution Approach 1:
The system applies error mitigation by design through the mixing matrix structure, which inherently distributes and balances channel information. This beforehand cushioning approach prevents error propagation without requiring additional complex error correction algorithms, maintaining reliability while minimizing computational overhead.
4Manufacturing precision
If more processing resources are allocated to maintain high listening quality, then sound quality is improved, but energy consumption increases
Solution Approach 1:
The system optimizes the mixing matrix parameters to achieve the best possible listening quality with minimal processing resources. By carefully selecting and tuning the transformation coefficients in the mixing matrix, the patent maintains high sound quality while reducing the computational energy required for processing, making it suitable for battery-powered portable devices.
Data Source
AI summary
An audio processing system (100) for spatial synthesis comprises an upmix stage (110) receiving a decoded m-channel downmix signal (X) and outputting, based thereon, an n-channel upmix signal (Y), wherein 2≦m<n. The upmix stage comprises a downmix modifying processor (120), which receives the m-channel downmix signal and outputting a modified downmix signal (d1, d2) obtained by cross mixing and non-linear processing of the downmix signal, and further comprises a first mixing matrix (130) receiving the downmix signal and the modified downmix signal, forming an n-channel linear combination of the downmix signal channels and modified downmix signal channels only and outputting this as the n-channel upmix signal. In an embodiment, the first mixing matrix accepts one or more mixing parameters (g, α1, . . . ) controlling at least one gain in the linear combination performed by the first mixing matrix. The gains are polynomials of degree ≦2.


