Scalable Audio Codec Layered Compression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio encoding technologies struggle to provide scalable and high-quality audio compression that adapts to varying bitrates and storage constraints, particularly in portable devices, where higher bitrates are not feasible due to limited storage capacity.

Innovation Solution

The development of a scalable audio codec that employs perceptual transform coding, where a base layer is encoded and partially decoded to compute residual coefficients, which are then encoded into an enhancement layer, allowing for multiple layers to scale bitstream size and quality, enabling lossless or near-lossless audio reconstruction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If higher bitrates are used for high quality sound reproduction, then audio quality is improved, but storage capacity is exceeded in portable devices

Engineering Contradiction:
Improveaudio qualityVSAvoidstorage capacity
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The audio bitstream is segmented into multiple layers: a base layer providing core audio quality and enhancement layers providing incremental quality improvements. This segmentation allows selective decoding of layers based on available storage capacity, enabling high audio quality when storage is abundant while maintaining acceptable quality when storage is limited, thus resolving the contradiction between audio quality and storage capacity requirements

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The codec dynamically adjusts the number of decoded layers based on available storage capacity. The system can adaptively select to decode only the base layer, base layer plus one or more enhancement layers, or all layers, providing dynamic quality adjustment that matches available storage resources, thereby resolving the fixed contradiction between quality and storage

Inventive Principle:
Principle #15Dynamics

2Quantity of substance

If a single bitrate is selected for compression, then file size is controlled, but adaptability to different audio applications is reduced

Engineering Contradiction:
Improvefile sizeVSAvoidaudio application coverage
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The compressed audio data is segmented into a base layer and multiple enhancement layers, each representing different quality levels. This segmentation enables the same compressed bitstream to serve multiple audio applications: portable devices can use only the base layer for small file sizes, while audiophile applications can decode all layers for high quality, thus achieving both file size control and broad adaptability

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The scalable codec structure provides universal compatibility across different audio applications. A single encoded bitstream can be decoded at multiple quality levels suitable for different devices and use cases, making the system universally applicable from portable media players to high-fidelity audio systems without requiring separate encodings for each application

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS8386271B2Lossless and near lossless scalable audio codec
Publication Date: 2013.02.26 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8386271B2 patent drawing
  • US8386271B2 patent drawing
  • US8386271B2 patent drawing

AI summary

A scalable audio codec encodes an input audio signal as a base layer at a high compression ratio and one or more residual signals as an enhancement layer of a compressed bitstream, which permits a lossless or near lossless reconstruction of the input audio signal at decoding. The scalable audio codec uses perceptual transform coding to encode the base layer. The residual is calculated in a transform domain, which includes a frequency and possibly also multi-channel transform of the input audio. For lossless reconstruction, the frequency and multi-channel transforms are reversible.