HFSC Tool for Sinusoidal Audio Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The MPEG-H 3D Audio Codec faces challenges in accurately representing and reconstructing high frequency tonal components, particularly for sounds with prominent high frequency partials, leading to distortion and artifacts due to the limitations of the eSBR tool, which struggles with highly pitched sounds and frequency variations.

Innovation Solution

The High Frequency Sinusoidal Coding (HFSC) tool is introduced to encode selected high frequency tonal components using sinusoidal modeling, representing them as sinusoidal trajectories with varying frequencies and amplitudes, and is activated only when strong tonal components are detected, thereby improving the representation and reducing distortion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If eSBR tool is used for high frequency coding, then bandwidth extension is achieved, but distortion and artifacts occur for highly pitched sounds with prominent high frequency partials

Engineering Contradiction:
Improveaccuracy of high frequency representationVSAvoiddistortion and artifacts
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent changes the coding approach from spectral band replication (eSBR) to sinusoidal modeling (HFSC), fundamentally altering the representation parameters. By detecting prominent high frequency partials and modeling them as sinusoidal trajectories with explicit frequency, amplitude, and phase parameters, the system accurately represents highly pitched sounds without the distortion and artifacts that plague eSBR methods.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If sinusoidal modeling is applied to all frequency ranges, then tonal components are accurately represented, but computational complexity and memory requirements increase significantly

Engineering Contradiction:
Improveaccuracy of tonal component representationVSAvoidcomputational complexity and memory requirements
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies sinusoidal modeling selectively only to frequency regions containing prominent partials, rather than uniformly across the entire spectrum. The encoder detects specific high frequency regions with strong tonal content and applies HFSC only there, while using conventional coding for other regions. This localized application maintains high accuracy for tonal components while significantly reducing overall computational complexity and memory requirements.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system applies sinusoidal modeling partially - only to the extent necessary for representing prominent partials. By using a threshold-based detection mechanism, the encoder activates HFSC only when and where prominent high frequency partials are present, avoiding the excessive computational burden of applying sinusoidal modeling to all frequency ranges regardless of content.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If high frequency tonal components are encoded with high precision, then audio quality is maintained, but bit rate consumption increases

Engineering Contradiction:
Improveaudio qualityVSAvoidbit rate efficiency
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent transforms the coding parameters from spectral coefficients to sinusoidal trajectory parameters (frequency, amplitude, phase). This parameter transformation enables more efficient representation of tonal components by exploiting the structured nature of sinusoidal signals. The sinusoidal model captures the essential characteristics of prominent partials with fewer bits compared to traditional spectral coding, maintaining high audio quality while improving bit rate efficiency.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10971165B2Method and apparatus for sinusoidal encoding and decoding
Publication Date: 2021.04.06 HUAWEI TECH CO LTD
  • US10971165B2 patent drawing
  • US10971165B2 patent drawing
  • US10971165B2 patent drawing

AI summary

An audio signal encoding method is provided that comprises collecting audio signal samples, determining sinusoidal components in subsequent frames, estimating amplitudes and frequencies of the components for each frame, merging the obtained pairs into sinusoidal trajectories, splitting particular trajectories into segments, transforming particular trajectories to the frequency domain by way of a digital transform performed on segments longer than the frame duration, quantization and selection of transform coefficients in the segments, entropy encoding, outputting the quantized coefficients as output data, wherein segments of different trajectories starting within a particular time are grouped into Groups of Segments, and the partitioning of trajectories into segments is synchronized with the endpoints of a Group of Segments.