Audio Frame Encoding with MDCT Overlap for Speech and Music
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio coding technologies face challenges in achieving efficient coding for both speech and music signals at low bitrates, as dedicated LPC-based speech coders perform poorly on music signals due to their inability to shape spectral distortion according to masking thresholds, while general audio coders fail to deliver convincing results for speech signals due to the lack of a speech source model.
Innovation Solution
The proposed solution combines the strengths of LPC-based coding and perceptual audio coding by using time-aliasing introducing transforms, such as the Modified Discrete Cosine Transform (MDCT), to enable critical sampling and cross-fading between frames, allowing for adaptive coding that can efficiently handle both speech and music signals by transforming overlapping frames into the frequency domain and using a redundancy reducing encoder.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If LPC-based speech coding is used, then speech coding efficiency is improved, but music coding quality deteriorates
Solution Approach 1:
The patent implements dynamic mode switching between speech and music coding algorithms based on signal classification. The system analyzes input signal characteristics and adapts the coding approach in real-time, using LPC for speech segments and perceptual coding for music segments, thereby resolving the contradiction between speech efficiency and music quality
Solution Approach 2:
The patent changes coding parameters dynamically based on signal type detection. When speech is detected, LPC parameters are activated; when music is detected, perceptual coding parameters are used. This parameter adaptation allows optimal performance for both speech and music without compromising either
2Manufacturing precision
If general perceptual audio coding is used, then music coding quality is improved, but speech coding efficiency deteriorates
Solution Approach 1:
The system dynamically selects coding strategies based on real-time signal analysis. For speech signals, it switches to LPC-based coding to maintain efficiency; for music signals, it uses perceptual coding to ensure quality. This dynamic adaptation resolves the contradiction between music quality and speech efficiency
Solution Approach 2:
The patent creates a universal audio coding framework that can handle both speech and music signals effectively. By integrating multiple coding algorithms (LPC and perceptual coding) into a single system with automatic signal classification, it achieves multi-functionality that resolves the contradiction between specialized performance for different signal types
3Loss of substance
If time-aliasing introducing transforms with critical sampling are used, then overhead is reduced, but cross-fading capability is limited
Solution Approach 1:
The patent embeds overlapping frames within the critical sampling transform structure. By nesting the overlap-add technique inside the MDCT framework and carefully managing the time-aliasing properties, it achieves both reduced overhead and smooth cross-fading capability
Solution Approach 2:
The patent uses windowing functions as intermediaries to enable cross-fading between frames in the critical sampling transform. The window functions smoothly transition between adjacent frames, mediating the trade-off between overhead reduction and cross-fading capability
Data Source
Figure 1
Figure 2A~2J
Figure 3A
AI summary
An audio encoder (10) adapted for encoding frames of a sampled audio signal to obtain encoded frames, wherein a frame comprises a number of time domain audio samples. The audio encoder (10) comprises a predictive coding analysis stage (12) for determining information on coefficients of a synthesis filter and an excitation frame based on a frame of audio samples, the excitation frame comprising samples of an excitation signal for the synthesis filter. The audio encoder (10) further comprises a time-aliasing introducing transformer (14) for transforming overlapping excitation frames to the frequency domain to obtain excitation frame spectra, wherein the time-aliasing introducing transformer (14) is adapted for transforming the overlapping excitation frames in a critically-sampled way. Moreover, the audio encoder (10) comprises a redundancy reducing encoder (16) for encoding the excitation frame spectra to obtain the encoded frames based on the coefficients and the encoded excitation frame spectra.