Audio Frame Encoding with MDCT Overlap for Speech and Music

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio coding technologies face challenges in achieving efficient coding for both speech and music signals at low bitrates, as dedicated LPC-based speech coders perform poorly on music signals due to their inability to shape spectral distortion according to masking thresholds, while general audio coders fail to deliver convincing results for speech signals due to the lack of a speech source model.

Innovation Solution

The proposed solution combines the strengths of LPC-based coding and perceptual audio coding by using time-aliasing introducing transforms, such as the Modified Discrete Cosine Transform (MDCT), to enable critical sampling and cross-fading between frames, allowing for adaptive coding that can efficiently handle both speech and music signals by transforming overlapping frames into the frequency domain and using a redundancy reducing encoder.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If LPC-based speech coding is used, then speech coding efficiency is improved, but music coding quality deteriorates

Engineering Contradiction:
Improvespeech coding efficiencyVSAvoidmusic coding quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent implements dynamic mode switching between speech and music coding algorithms based on signal classification. The system analyzes input signal characteristics and adapts the coding approach in real-time, using LPC for speech segments and perceptual coding for music segments, thereby resolving the contradiction between speech efficiency and music quality

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes coding parameters dynamically based on signal type detection. When speech is detected, LPC parameters are activated; when music is detected, perceptual coding parameters are used. This parameter adaptation allows optimal performance for both speech and music without compromising either

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If general perceptual audio coding is used, then music coding quality is improved, but speech coding efficiency deteriorates

Engineering Contradiction:
Improvemusic coding qualityVSAvoidspeech coding efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system dynamically selects coding strategies based on real-time signal analysis. For speech signals, it switches to LPC-based coding to maintain efficiency; for music signals, it uses perceptual coding to ensure quality. This dynamic adaptation resolves the contradiction between music quality and speech efficiency

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent creates a universal audio coding framework that can handle both speech and music signals effectively. By integrating multiple coding algorithms (LPC and perceptual coding) into a single system with automatic signal classification, it achieves multi-functionality that resolves the contradiction between specialized performance for different signal types

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of substance

If time-aliasing introducing transforms with critical sampling are used, then overhead is reduced, but cross-fading capability is limited

Engineering Contradiction:
Improvecoding overheadVSAvoidcross-fading capability
Core Design Contradiction:
Loss of substanceVSAdaptability or versatility

Solution Approach 1:

The patent embeds overlapping frames within the critical sampling transform structure. By nesting the overlap-add technique inside the MDCT framework and carefully managing the time-aliasing properties, it achieves both reduced overhead and smooth cross-fading capability

Inventive Principle:
Principle #7Nested doll (Nesting)

Solution Approach 2:

The patent uses windowing functions as intermediaries to enable cross-fading between frames in the critical sampling transform. The window functions smoothly transition between adjacent frames, mediating the trade-off between overhead reduction and cross-fading capability

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP2144171B1Audio encoder and decoder for encoding and decoding frames of a sampled audio signal
Publication Date: 2018.05.16 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • EP2144171B1 patent drawingFigure 1
  • EP2144171B1 patent drawingFigure 2A~2J
  • EP2144171B1 patent drawingFigure 3A

AI summary

An audio encoder (10) adapted for encoding frames of a sampled audio signal to obtain encoded frames, wherein a frame comprises a number of time domain audio samples. The audio encoder (10) comprises a predictive coding analysis stage (12) for determining information on coefficients of a synthesis filter and an excitation frame based on a frame of audio samples, the excitation frame comprising samples of an excitation signal for the synthesis filter. The audio encoder (10) further comprises a time-aliasing introducing transformer (14) for transforming overlapping excitation frames to the frequency domain to obtain excitation frame spectra, wherein the time-aliasing introducing transformer (14) is adapted for transforming the overlapping excitation frames in a critically-sampled way. Moreover, the audio encoder (10) comprises a redundancy reducing encoder (16) for encoding the excitation frame spectra to obtain the encoded frames based on the coefficients and the encoded excitation frame spectra.