Adaptive Audio Companding for Dense Transient Noise Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio codecs face challenges in effectively reducing quantization noise, particularly pre-echo artifacts, due to the difficulty in predicting the type of companding needed for specific audio signals, and existing detectors are not accurate enough to distinguish between speech, applause, and tonal audio content.

Innovation Solution

A signal-dependent companding system that classifies audio signals as pure sinusoidal, hybrid, or pure transient using threshold values and applies selective companding operations in the QMF domain, using temporal and spectral sharpness measures to determine appropriate companding modes for different types of content, thereby reducing audio distortion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If companding is applied to reduce quantization noise, then temporal noise shaping is improved, but it is difficult to predict the type of companding needed for specific signal types

Engineering Contradiction:
Improvetemporal noise shapingVSAvoidsignal type prediction
Core Design Contradiction:
Manufacturing precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The system dynamically adapts the companding exponent based on the detected signal type. Different companding exponents are applied for speech, music, and transient signals, making the noise shaping process dynamic rather than static. This resolves the contradiction by enabling precise temporal noise shaping while adapting to different signal characteristics through real-time detection and adjustment.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The invention changes the companding exponent parameter based on signal type detection. By adjusting this key parameter according to whether the signal is speech, music, or transient, the system achieves optimal temporal noise shaping for each signal type without requiring complex prediction algorithms.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If existing detectors are used to identify speech and applause, then detection capability is provided, but accuracy is insufficient to distinguish between speech, applause, and tonal audio content

Engineering Contradiction:
Improvedetection capabilityVSAvoidsignal classification accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The detection process is segmented into multiple stages: initial signal type detection, transient detection, and hybrid signal classification. This multi-stage segmentation allows the system to achieve high accuracy in distinguishing between speech, applause, and tonal audio content by breaking down the complex detection task into manageable steps with increasing precision.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The detector dynamically adjusts its classification behavior based on the detected signal characteristics. For hybrid signals containing both transient and tonal components, the system dynamically determines the appropriate companding exponent, improving classification accuracy while maintaining adaptability to different audio content types.

Inventive Principle:
Principle #15Dynamics

3Manufacturing precision

If companding is applied to improve speech and applause coding, then sound quality is improved, but overhead in audio encoding increases

Engineering Contradiction:
Improvesound qualityVSAvoidencoding overhead
Core Design Contradiction:
Manufacturing precisionVSLoss of information

Solution Approach 1:

The system applies companding selectively rather than universally. By detecting the signal type and applying companding only when beneficial (for speech, applause, and certain transient signals), the system improves sound quality for relevant content while minimizing encoding overhead by avoiding unnecessary companding operations on other signal types.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

Different companding exponents are applied locally based on the specific signal type detected. Speech signals receive one exponent, applause another, and transient signals yet another. This localized approach optimizes sound quality for each signal type while keeping the encoding overhead manageable through targeted application.

Inventive Principle:
Principle #3Local quality

4Ease of operation

If detectors are designed to be simple with low complexity, then ease of operation is improved, but detection accuracy decreases

Engineering Contradiction:
Improvedetector complexityVSAvoidsignal detection accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The detection process is segmented into multiple simple stages rather than one complex operation. Each stage performs a specific function (initial detection, transient detection, hybrid classification), making each individual step simple and easy to implement while achieving high overall accuracy through the cumulative effect of multiple stages.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11830507B2Coding dense transient events with companding
Publication Date: 2023.11.28 DOLBY INTERNATIONAL AB
  • US11830507B2 patent drawing
  • US11830507B2 patent drawing
  • US11830507B2 patent drawing

AI summary

Embodiments are directed to a companding method and system for reducing coding noise in an audio codec. A method of processing an audio signal includes the following operations. A system receives an audio signal. The system determines that a first frame of the audio signal includes a sparse transient signal. The system determines that a second frame of the audio signal includes a dense transient signal. The system compresses/expands (compands) the audio signal using a companding rule that applies a first companding exponent to the first frame of the audio signal and applies a second companding exponent to the second frame of the audio signal, each companding exponent being used to derive a respective degree of dynamic range compression and expansion for a corresponding frame. The system then provides the companded audio signal to a downstream device.