ToneFilling Audio Codec for Low Bit Rate Sinusoid Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern audio codecs face challenges in delivering high-quality audio at low bit rates with low latency, particularly in music coding, where they often introduce warbling and roughness artifacts due to suboptimal frequency response and sparse spectral coding, and fully parametric codecs sound artificial and lack scalability.

Innovation Solution

The ToneFilling technique generates synthetic tones by patching tone patterns into the MDCT spectrum, adapting them to their target location for seamless synthesis of high-quality sinusoidal tones and sweeps, integrating with existing transform coding schemes to minimize computational complexity and artifacts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If transform coding is used for music content, then compression efficiency is improved, but warbling and roughness artifacts are introduced at low bit rates

Engineering Contradiction:
Improvecompression efficiencyVSAvoidwarbling and roughness artifacts
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The patent segments the audio signal into sinusoidal components and noise components separately. By identifying and extracting sinusoidal tones using spectral patterns, the coding system can apply different coding strategies to different signal components, thereby reducing warbling artifacts in tonal regions while maintaining compression efficiency overall.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality enhancement by using spectral patterns to identify specific locations in the frequency spectrum where sinusoidal components exist. At these localized positions, the coding system can apply more precise quantization and reconstruction methods to eliminate warbling artifacts, while using coarser coding in non-tonal regions to maintain overall compression efficiency.

Inventive Principle:
Principle #3Local quality

2Loss of time

If low-delay transform window shape is used, then latency is reduced, but frequency response quality deteriorates

Engineering Contradiction:
ImprovelatencyVSAvoidfrequency response quality
Core Design Contradiction:
Loss of timeVSManufacturing precision

Solution Approach 1:

The patent uses spectral patterns as templates or copies of expected sinusoidal spectra. By comparing the actual spectrum against these pre-defined patterns, the system can accurately identify and reconstruct sinusoidal components even with low-delay windowing, thereby recovering frequency response quality without increasing latency.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent changes the approach from trying to optimize the transform window for both low delay and high quality simultaneously, to instead using spectral pattern matching to compensate for the suboptimal window characteristics. This parameter change in the coding strategy allows the system to achieve high frequency response quality despite using low-delay window shapes.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If spectral coefficients are sparsely coded to save bits, then bit rate is reduced, but temporal modulation artifacts and warbling are introduced

Engineering Contradiction:
Improvebit rateVSAvoidtemporal modulation artifacts and warbling
Core Design Contradiction:
Quantity of substanceVSObject-affected harmful factors

Solution Approach 1:

The patent performs preliminary action by pre-defining spectral patterns that represent expected sinusoidal components. During coding, instead of sparsely coding individual coefficients, the system matches the spectrum against these pre-defined patterns and codes the pattern parameters. This preliminary preparation allows for more efficient representation of tonal components, reducing bit rate while avoiding temporal modulation artifacts and warbling.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces spectral patterns as an intermediary between the raw spectral coefficients and the final coded representation. Rather than directly coding sparse coefficients or using complex parametric models, the system uses spectral patterns as a middle layer that bridges the gap, enabling efficient coding without introducing artifacts.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Adaptability or versatility

If fully parametric coding is used for transients and sinusoids, then scalability is improved, but sound quality becomes artificial

Engineering Contradiction:
ImprovescalabilityVSAvoidartificial sound quality
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The patent merges transform coding and parametric coding approaches by combining spectral pattern matching with the transform domain. Instead of using purely parametric modeling that sounds artificial, the system integrates pattern-based sinusoid identification within the transform coding framework, thereby achieving scalability while maintaining natural sound quality through the hybrid approach.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP2907132B1Apparatus and method for efficient synthesis of sinusoids and sweeps by employing spectral patterns
Publication Date: 2021.09.29 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • EP2907132B1 patent drawingFigure 1A
  • EP2907132B1 patent drawingFigure 1B
  • EP2907132B1 patent drawingFigure 1C

AI summary

An apparatus for generating an audio output signal based on an encoded audio signal spectrum is provided. The apparatus comprises a processing unit (1 15) for processing the encoded audio signal spectrum to obtain a decoded audio signal spectrum comprising a plurality of spectral coefficients, wherein each of the spectral coefficients has a spectral location within the encoded audio signal spectrum and a spectral value, wherein the spectral coefficients are sequentially ordered according to their spectral location within the encoded audio signal spectrum so that the spectral coefficients form a sequence of spectral coefficients. Moreover, the apparatus comprises a pseudo coefficients determiner (125) for determining one or more pseudo coefficients of the decoded audio signal spectrum, each of the pseudo coefficients having a spectral value. Furthermore, the apparatus comprises a replacement unit (135) for replacing at least one or more pseudo coefficients by a determined spectral pattern to obtain a modified audio signal spectrum, wherein the determined spectral pattern comprises at least two pattern coefficients, wherein each of the at least two pattern coefficients has a spectral value. Moreover, the apparatus comprises a spectrum-time-conversion unit (145) for converting the modified audio signal spectrum to a time-domain to obtain the audio output signal.