HFSC Tool for Sinusoidal Audio Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The MPEG-H 3D Audio Codec faces challenges in accurately representing and reconstructing high frequency tonal components, particularly for sounds with prominent high frequency partials, leading to distortion and artifacts due to the limitations of the eSBR tool, which struggles with highly pitched sounds and frequency variations.
Innovation Solution
The High Frequency Sinusoidal Coding (HFSC) tool is introduced to encode selected high frequency tonal components using sinusoidal modeling, representing them as sinusoidal trajectories with varying frequencies and amplitudes, and is activated only when strong tonal components are detected, thereby improving the representation and reducing distortion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If eSBR tool is used for high frequency coding, then bandwidth extension is achieved, but distortion and artifacts occur for highly pitched sounds with prominent high frequency partials
Solution Approach 1:
The patent changes the coding approach from spectral band replication (eSBR) to sinusoidal modeling (HFSC), fundamentally altering the representation parameters. By detecting prominent high frequency partials and modeling them as sinusoidal trajectories with explicit frequency, amplitude, and phase parameters, the system accurately represents highly pitched sounds without the distortion and artifacts that plague eSBR methods.
2Reliability
If sinusoidal modeling is applied to all frequency ranges, then tonal components are accurately represented, but computational complexity and memory requirements increase significantly
Solution Approach 1:
The patent applies sinusoidal modeling selectively only to frequency regions containing prominent partials, rather than uniformly across the entire spectrum. The encoder detects specific high frequency regions with strong tonal content and applies HFSC only there, while using conventional coding for other regions. This localized application maintains high accuracy for tonal components while significantly reducing overall computational complexity and memory requirements.
Solution Approach 2:
The system applies sinusoidal modeling partially - only to the extent necessary for representing prominent partials. By using a threshold-based detection mechanism, the encoder activates HFSC only when and where prominent high frequency partials are present, avoiding the excessive computational burden of applying sinusoidal modeling to all frequency ranges regardless of content.
3Reliability
If high frequency tonal components are encoded with high precision, then audio quality is maintained, but bit rate consumption increases
Solution Approach 1:
The patent transforms the coding parameters from spectral coefficients to sinusoidal trajectory parameters (frequency, amplitude, phase). This parameter transformation enables more efficient representation of tonal components by exploiting the structured nature of sinusoidal signals. The sinusoidal model captures the essential characteristics of prominent partials with fewer bits compared to traditional spectral coding, maintaining high audio quality while improving bit rate efficiency.
Data Source
AI summary
An audio signal encoding method is provided that comprises collecting audio signal samples, determining sinusoidal components in subsequent frames, estimating amplitudes and frequencies of the components for each frame, merging the obtained pairs into sinusoidal trajectories, splitting particular trajectories into segments, transforming particular trajectories to the frequency domain by way of a digital transform performed on segments longer than the frame duration, quantization and selection of transform coefficients in the segments, entropy encoding, outputting the quantized coefficients as output data, wherein segments of different trajectories starting within a particular time are grouped into Groups of Segments, and the partitioning of trajectories into segments is synchronized with the endpoints of a Group of Segments.


