Adaptive Audio Companding for Dense Transient Noise Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio codecs face challenges in effectively reducing quantization noise, particularly pre-echo artifacts, due to the difficulty in predicting the type of companding needed for specific audio signals, and existing detectors are not accurate enough to distinguish between speech, applause, and tonal audio content.
Innovation Solution
A signal-dependent companding system that classifies audio signals as pure sinusoidal, hybrid, or pure transient using threshold values and applies selective companding operations in the QMF domain, using temporal and spectral sharpness measures to determine appropriate companding modes for different types of content, thereby reducing audio distortion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If companding is applied to reduce quantization noise, then temporal noise shaping is improved, but it is difficult to predict the type of companding needed for specific signal types
Solution Approach 1:
The system dynamically adapts the companding exponent based on the detected signal type. Different companding exponents are applied for speech, music, and transient signals, making the noise shaping process dynamic rather than static. This resolves the contradiction by enabling precise temporal noise shaping while adapting to different signal characteristics through real-time detection and adjustment.
Solution Approach 2:
The invention changes the companding exponent parameter based on signal type detection. By adjusting this key parameter according to whether the signal is speech, music, or transient, the system achieves optimal temporal noise shaping for each signal type without requiring complex prediction algorithms.
2Adaptability or versatility
If existing detectors are used to identify speech and applause, then detection capability is provided, but accuracy is insufficient to distinguish between speech, applause, and tonal audio content
Solution Approach 1:
The detection process is segmented into multiple stages: initial signal type detection, transient detection, and hybrid signal classification. This multi-stage segmentation allows the system to achieve high accuracy in distinguishing between speech, applause, and tonal audio content by breaking down the complex detection task into manageable steps with increasing precision.
Solution Approach 2:
The detector dynamically adjusts its classification behavior based on the detected signal characteristics. For hybrid signals containing both transient and tonal components, the system dynamically determines the appropriate companding exponent, improving classification accuracy while maintaining adaptability to different audio content types.
3Manufacturing precision
If companding is applied to improve speech and applause coding, then sound quality is improved, but overhead in audio encoding increases
Solution Approach 1:
The system applies companding selectively rather than universally. By detecting the signal type and applying companding only when beneficial (for speech, applause, and certain transient signals), the system improves sound quality for relevant content while minimizing encoding overhead by avoiding unnecessary companding operations on other signal types.
Solution Approach 2:
Different companding exponents are applied locally based on the specific signal type detected. Speech signals receive one exponent, applause another, and transient signals yet another. This localized approach optimizes sound quality for each signal type while keeping the encoding overhead manageable through targeted application.
4Ease of operation
If detectors are designed to be simple with low complexity, then ease of operation is improved, but detection accuracy decreases
Solution Approach 1:
The detection process is segmented into multiple simple stages rather than one complex operation. Each stage performs a specific function (initial detection, transient detection, hybrid classification), making each individual step simple and easy to implement while achieving high overall accuracy through the cumulative effect of multiple stages.
Data Source
AI summary
Embodiments are directed to a companding method and system for reducing coding noise in an audio codec. A method of processing an audio signal includes the following operations. A system receives an audio signal. The system determines that a first frame of the audio signal includes a sparse transient signal. The system determines that a second frame of the audio signal includes a dense transient signal. The system compresses/expands (compands) the audio signal using a companding rule that applies a first companding exponent to the first frame of the audio signal and applies a second companding exponent to the second frame of the audio signal, each companding exponent being used to derive a respective degree of dynamic range compression and expansion for a corresponding frame. The system then provides the companded audio signal to a downstream device.


