MDCT Audio Encoder With Frequency-Domain Harmonic Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio coding methods face challenges in efficiently predicting periodic components of audio signals due to non-stationarity and high fundamental frequencies, leading to instability and increased bitrate, especially in transform coding with backward adaptation.
Innovation Solution
The implementation of Frequency Domain Least Mean Square Prediction (FDLMSP) directly in the Modified Discrete Cosine Transform (MDCT) domain, which models harmonic components using a real-valued linear equation system to predict audio frames based on the phase progression of harmonics, reducing bitrate and enhancing coding efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If Time Domain Long-term Prediction (TDLTP) is used with MDCT transform, then periodic components can be predicted, but the prediction stability deteriorates due to non-stationarity and high fundamental frequencies
Solution Approach 1:
The patent replaces the time-domain prediction mechanism with a frequency-domain prediction mechanism. Instead of transforming MDCT coefficients back to time domain for prediction and then transforming back, the invention performs prediction directly in the frequency domain by modeling harmonic components using spectral coefficients, thereby avoiding the instability caused by time-domain transformation and reconstruction.
Solution Approach 2:
The patent changes the domain parameter from time domain to frequency domain for the prediction process. By operating on spectral coefficients directly in the frequency domain and modeling harmonic components with parameters like fundamental frequency and harmonic numbers, the prediction becomes more stable for high fundamental frequencies and non-stationary signals.
2Productivity
If Frequency Domain Prediction (FDP) treating each harmonic component individually is used, then prediction can be performed in MDCT domain, but performance deteriorates when frequency resolution is low and harmonic components overlap heavily
Solution Approach 1:
The patent merges the modeling of multiple harmonic components into a unified real-valued linear equation system. Instead of treating each harmonic component separately as in traditional FDP, the invention combines all harmonic components into a single system of equations that can be solved simultaneously, thereby improving performance when harmonic components overlap heavily due to low frequency resolution.
3Ease of manufacture
If traditional TDLTP with backward adaptation is used, then audio coding can be performed, but bitrate increases due to the need for transformation and residual coding
Solution Approach 1:
The patent substitutes the multi-step time-domain prediction process with a direct frequency-domain prediction approach. By performing prediction on spectral coefficients without time-domain transformation, the invention eliminates redundant processing steps and reduces the amount of residual data that needs to be coded, thereby reducing bitrate while maintaining coding simplicity.
Data Source
AI summary
An encoder for encoding a current frame of an audio signal depending on one or more previous frames of the audio signal is provided. The previous frames precede the current frame, each of the current frame and the one or more previous frames having one or more harmonic components of the audio signal, each of the current frame and the one or more previous frames having a plurality of spectral coefficients in a frequency domain or in a transform domain. To generate an encoding of the current frame, the encoder is to determine an estimation of two harmonic parameters for each of the harmonic components of a most previous frame of the previous frames. Moreover, the encoder is to determine the estimation of the two harmonic parameters for each of the harmonic components of the most previous frame using a first group of three or more of the plurality of spectral coefficients of each of the previous frames of the audio signal.


