Single-Rate Audio Codec Architecture for Voice Upmix Compatibility

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio codec systems are often cumbersome due to disparate design paradigms and technological levels, with backward compatibility constraints leading to less coherent architectures, particularly in parametric multichannel systems, which can result in suboptimal performance for voice signals.

Innovation Solution

A versatile and architecturally uniform audio codec system is developed, featuring a single-rate architecture with a core coder and parametric upmix stage, capable of operating in both audio and voice modes, with adaptive sampling rates and optional spectral band replication for efficient voice signal processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If backward compatibility is preserved in parametric multichannel audio codec systems, then legacy equipment compatibility is maintained, but system architecture coherence deteriorates

Engineering Contradiction:
Improvebackward compatibilityVSAvoidsystem architecture coherence
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The audio signal is segmented into different frequency bands (low-frequency baseband signal and high-frequency parametric signal), allowing independent processing of each segment. This segmentation enables the system to maintain compatibility with legacy equipment that processes only the baseband signal while incorporating advanced parametric processing for enhanced channels without compromising overall system coherence.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A downmix signal serves as an intermediary between the original multichannel audio and the legacy playback system. The downmix signal contains essential low-frequency information that legacy equipment can process, while the parametric upmix stage uses this intermediary to reconstruct high-frequency content, thus bridging the gap between old and new system requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If conventional audio coding formats are used, then broad audio signal compatibility is achieved, but voice signal processing performance deteriorates

Engineering Contradiction:
Improveaudio signal compatibilityVSAvoidvoice signal processing performance
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The system dynamically adapts its processing parameters based on the input signal type. For voice signals, it employs reduced frame lengths and optimized parametric processing specifically tuned for speech characteristics, while maintaining compatibility with general audio signals through configurable processing modes that adjust to different signal types.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

Different processing qualities are applied to different parts of the signal processing chain. The low-frequency baseband processing uses conventional robust methods ensuring broad compatibility, while the high-frequency parametric processing applies specialized algorithms optimized for voice signal characteristics, achieving local optimization without sacrificing overall system versatility.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9812136B2Audio processing system
Publication Date: 2017.11.07 DOLBY INTERNATIONAL AB
  • US9812136B2 patent drawing
  • US9812136B2 patent drawing
  • US9812136B2 patent drawing

AI summary

An audio processing system (100) comprises a front-end component (102, 103), which receives quantized spectral components and performs an inverse quantization, yielding a time-domain representation of an intermediate signal. The audio processing system further comprises a frequency-domain processing stage (104, 105, 106, 107, 108), configured to provide a time-domain representation of a processed audio signal, and a sample rate converter (109), providing a reconstructed audio signal sampled at a target sampling frequency. The respective internal sampling rates of the time-domain representation of the intermediate audio signal and of the time-domain representation of the processed audio signal are equal. In particular embodiments, the processing stage comprises a parametric upmix stage which is operable in at least two different modes and is associated with a delay stage that ensures constant total delay.