Mixed Audio Coding for Low-Bitrate, Low-Delay Generic Signals

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conversational codecs struggle to maintain high-quality encoding of generic audio signals like music and reverberant speech at low bitrates without increasing processing delay, and existing switched codecs require longer delays due to speech-music classification and frequency domain transformation.

Innovation Solution

A mixed time-domain and frequency-domain coding model that dynamically allocates bits between adaptive and fixed codebooks, integrates frequency-domain coding in the LP residual domain, and uses a variable time support and cut-off frequency to enhance synthesis quality, while minimizing artifacts and delay.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If switched codecs are used to code generic audio signals in frequency domain, then synthesis quality is improved, but processing delay increases due to speech-music classification and transform requirements

Engineering Contradiction:
Improvesynthesis qualityVSAvoidprocessing delay
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent merges time-domain and frequency-domain coding techniques into a unified model. The frequency-domain coding mode is integrated within the time-domain CELP framework, allowing both approaches to coexist and work together seamlessly. This integration enables the system to achieve frequency-domain synthesis quality while maintaining time-domain processing speed, resolving the contradiction between quality improvement and delay increase.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements dynamic bit allocation between time-domain and frequency-domain coding modes based on signal characteristics. The system can adaptively switch between different coding modes and allocate bits dynamically to optimize performance for each specific signal type. This dynamic adaptation allows the system to maintain low delay while achieving high synthesis quality when needed.

Inventive Principle:
Principle #15Dynamics

2Manufacturing precision

If frequency-domain coding is integrated in time-domain CELP mode, then synthesis quality for generic audio is improved, but device complexity increases

Engineering Contradiction:
Improvesynthesis qualityVSAvoidcoding model complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent creates a universal coding model that can handle both speech and generic audio signals using the same integrated framework. The unified model performs multiple functions: it can operate in time-domain mode for speech, frequency-domain mode for generic audio, or mixed mode for optimal performance. This multi-functionality reduces overall system complexity compared to having separate specialized codecs, while still achieving high synthesis quality across different signal types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent changes key parameters such as bit allocation, coding mode selection, and transform characteristics to optimize performance. By dynamically adjusting these parameters based on signal analysis, the system achieves high synthesis quality without requiring complex fixed architectures. The parameter changes allow a relatively simple integrated model to perform as well as more complex specialized systems.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If bits are dynamically allocated among codebooks and frequency-domain modes, then coding efficiency is improved, but processing complexity increases

Engineering Contradiction:
Improvecoding efficiencyVSAvoidbit allocation complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements feedback mechanisms that continuously monitor signal characteristics and coding performance. Based on this feedback, the system dynamically adjusts bit allocation between different codebooks and frequency-domain modes. The feedback loop ensures that bits are allocated optimally to maximize coding efficiency for each specific signal type, achieving high productivity while keeping the complexity management through adaptive rather than fixed complex structures.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP4372747B1Coding generic audio signals at low bitrates and low delay
Publication Date: 2026.03.25 VOICEAGE EVS LLC
  • EP4372747B1 patent drawingFigure 1
  • EP4372747B1 patent drawingFigure 2
  • EP4372747B1 patent drawingFigure 3

AI summary

A mixed time-domain / frequency-domain coding device and method for coding an input sound signal, wherein a time-domain excitation contribution is calculated in response to the input sound signal. A cut-off frequency for the time-domain excitation contribution is also calculated in response to the input sound signal, and a frequency extent of the time-domain excitation contribution is adjusted in relation to this cut-off frequency. Following calculation of a frequency-domain excitation contribution in response to the input sound signal, the adjusted time-domain excitation contribution and the frequency-domain excitation contribution are added to form a mixed time-domain / frequency-domain excitation constituting a coded version of the input sound signal. In the calculation of the time-domain excitation contribution, the input sound signal may be processed in successive frames of the input sound signal and a number of sub-frames to be used in a current frame may be calculated. Corresponding encoder and decoder using the mixed time-domain / frequency-domain coding device are also described.