Single-Rate Audio Codec Architecture for Voice Upmix Compatibility
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio codec systems are often cumbersome due to disparate design paradigms and technological levels, with backward compatibility constraints leading to less coherent architectures, particularly in parametric multichannel systems, which can result in suboptimal performance for voice signals.
Innovation Solution
A versatile and architecturally uniform audio codec system is developed, featuring a single-rate architecture with a core coder and parametric upmix stage, capable of operating in both audio and voice modes, with adaptive sampling rates and optional spectral band replication for efficient voice signal processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If backward compatibility is preserved in parametric multichannel audio codec systems, then legacy equipment compatibility is maintained, but system architecture coherence deteriorates
Solution Approach 1:
The audio signal is segmented into different frequency bands (low-frequency baseband signal and high-frequency parametric signal), allowing independent processing of each segment. This segmentation enables the system to maintain compatibility with legacy equipment that processes only the baseband signal while incorporating advanced parametric processing for enhanced channels without compromising overall system coherence.
Solution Approach 2:
A downmix signal serves as an intermediary between the original multichannel audio and the legacy playback system. The downmix signal contains essential low-frequency information that legacy equipment can process, while the parametric upmix stage uses this intermediary to reconstruct high-frequency content, thus bridging the gap between old and new system requirements.
2Adaptability or versatility
If conventional audio coding formats are used, then broad audio signal compatibility is achieved, but voice signal processing performance deteriorates
Solution Approach 1:
The system dynamically adapts its processing parameters based on the input signal type. For voice signals, it employs reduced frame lengths and optimized parametric processing specifically tuned for speech characteristics, while maintaining compatibility with general audio signals through configurable processing modes that adjust to different signal types.
Solution Approach 2:
Different processing qualities are applied to different parts of the signal processing chain. The low-frequency baseband processing uses conventional robust methods ensuring broad compatibility, while the high-frequency parametric processing applies specialized algorithms optimized for voice signal characteristics, achieving local optimization without sacrificing overall system versatility.
Data Source
AI summary
An audio processing system (100) comprises a front-end component (102, 103), which receives quantized spectral components and performs an inverse quantization, yielding a time-domain representation of an intermediate signal. The audio processing system further comprises a frequency-domain processing stage (104, 105, 106, 107, 108), configured to provide a time-domain representation of a processed audio signal, and a sample rate converter (109), providing a reconstructed audio signal sampled at a target sampling frequency. The respective internal sampling rates of the time-domain representation of the intermediate audio signal and of the time-domain representation of the processed audio signal are equal. In particular embodiments, the processing stage comprises a parametric upmix stage which is operable in at least two different modes and is associated with a delay stage that ensures constant total delay.


