Audio Frame Reconstruction Using TCX LTP for Burst Packet Loss

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio coding systems face challenges in effectively managing signal fade-out during error concealment in switched audio coding systems, particularly in maintaining a comfortable noise level and spectral shape during burst packet losses, leading to suboptimal audio quality and increased computational complexity.

Innovation Solution

The proposed solution involves an apparatus and method for decoding audio signals that utilize a common comfort noise level tracing in the excitation domain, fading the TCX LTP gain to zero, and applying a consistent fade-out speed for both LTP and white noise, ensuring a smooth transition to a comfort noise-like signal during burst packet losses, while sharing noise level tracing functions to reduce complexity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If separate noise level tracing functions are used for ACELP and TCX modes, then each mode can maintain optimal noise characteristics, but the computational complexity increases

Engineering Contradiction:
Improvenoise level accuracyVSAvoidcomputational complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges the separate noise level tracing functions for ACELP and TCX modes into a single common function. This unified tracing mechanism operates in the excitation domain and serves both coding modes, thereby reducing computational complexity while maintaining the ability to trace noise levels accurately for each mode through shared processing logic.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The common noise level tracing function is designed to be universal, serving both ACELP and TCX coding modes simultaneously. By implementing a multi-functional tracing mechanism that adapts to different modes through parameter adjustments rather than separate implementations, the system achieves efficiency without sacrificing mode-specific optimization.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If TCX LTP gain is not faded during burst packet losses, then tonal components are preserved, but unwanted tonal artifacts are introduced in the comfort noise

Engineering Contradiction:
Improvetonal component preservationVSAvoidtonal artifacts
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent implements dynamic control of the TCX LTP gain during burst packet losses. The gain is faded out over time according to a controlled trajectory, allowing the system to adapt between preserving tonal components initially and eliminating them as comfort noise takes over, thereby avoiding persistent tonal artifacts while maintaining natural sound characteristics during the transition.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The fading of TCX LTP gain follows a periodic or staged approach during error concealment, where the gain is systematically reduced over multiple frames rather than abruptly removed. This controlled periodic reduction allows smooth transition from tonal preservation to comfort noise generation, preventing harmful tonal artifacts from persisting.

Inventive Principle:
Principle #19Periodic action

3Adaptability or versatility

If different fade-out speeds are used for LTP and white noise, then each component can be optimized independently, but the transition to comfort noise becomes less smooth

Engineering Contradiction:
Improveindependent optimizationVSAvoidtransition smoothness
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent combines the fade-out control of LTP and white noise into a synchronized process, where both components follow the same fade-out speed and trajectory. This unified approach ensures that the transition to comfort noise is smooth and coherent, avoiding discontinuities that would arise from independent fade-out rates while maintaining the ability to optimize the overall transition characteristics.

Inventive Principle:
Principle #5Merging (Combining)

4Measurement precision

If domain-specific noise level tracing is used for TCX, then tracing accuracy is improved, but complexity increases due to separate tracing mechanisms

Engineering Contradiction:
Improvenoise level tracing accuracyVSAvoidtracing mechanism complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces the excitation domain as an intermediary space for noise level tracing that serves both ACELP and TCX modes. By performing tracing operations in this intermediate domain rather than separately in each mode's native domain, the system achieves accurate noise level measurement for TCX while avoiding the complexity of entirely separate tracing mechanisms through a unified intermediate processing stage.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP3011563B1Audio decoding with reconstruction of corrupted or not received frames using TCX ltp
Publication Date: 2019.12.25 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • EP3011563B1 patent drawingFigure 1A
  • EP3011563B1 patent drawingFigure 1B
  • EP3011563B1 patent drawingFigure 1C

AI summary

An apparatus for decoding an encoded audio signal to obtain a reconstructed audio signal is provided. The apparatus comprises a receiving interface (1310) for receiving a plurality of frames, a delay buffer (1020; 1320) for storing audio signal samples of the decoded audio signal, a sample selector (1030; 1330) for selecting a plurality of selected audio signal samples from the audio signal samples being stored in the delay buffer (1020; 1320), and a sample processor (1040; 1340) for processing the selected audio signal samples to obtain reconstructed audio signal samples of the reconstructed audio signal. The sample selector (1030; 1330) is configured to select, if a current frame is received by the receiving interface (1310) and if the current frame being received by the receiving interface (1310) is not corrupted, the plurality of selected audio signal samples from the audio signal samples being stored in the delay buffer (1020; 1320) depending on a pitch lag information being comprised by the current frame. Moreover, the sample selector (1030; 1330) is configured to select, if the current frame is not received by the receiving interface (1310) or if the current frame being received by the receiving interface (1310) is corrupted, the plurality of selected audio signal samples from the audio signal samples being stored in the delay buffer (1020; 1320) depending on a pitch lag information being comprised by another frame being received previously by the receiving interface (1310).