Layered Audio Coding with Gain Profiles for Continuous Spatial Speech

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems for teleconferencing lack the ability to provide a spatially layered, encoded audio signal that offers a continuous and varied mix of sound fields and monophonic layers over time, failing to maintain a perceptually continuous listening experience.

Innovation Solution

An audio encoding system comprising a spatial analyzer, adaptive rotation stage, and analysis stage that decomposes audio signals into rotated signals and generates a time-variable gain profile, enabling efficient encoding and decoding of spatially layered audio signals, allowing for continuous and adaptive playback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If spatially layered audio encoding is implemented, then audio quality and listening experience are improved, but system complexity increases

Engineering Contradiction:
Improveaudio qualityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The audio signal is divided into multiple layers including a monophonic layer and directional metadata layers. This segmentation allows the system to process and transmit different audio components separately, improving overall audio quality while managing complexity through hierarchical processing

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces spatial dimensions to audio encoding by adding directional metadata that describes the spatial characteristics of audio sources. This dimensional enhancement transforms traditional monophonic audio into spatial audio without requiring complete system redesign

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If continuous audio transmission is maintained, then perceptual continuity is improved, but bandwidth consumption increases

Engineering Contradiction:
Improveperceptual continuityVSAvoidbandwidth
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential directional metadata from the full audio signal. Instead of transmitting complete multi-channel audio data, only the critical spatial parameters are extracted and transmitted, maintaining perceptual continuity while significantly reducing bandwidth requirements

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system changes the representation parameters of audio data by using compact directional metadata formats. This parameter transformation allows continuous audio perception to be achieved with reduced data transmission by focusing on essential spatial parameters rather than full waveform data

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If layered audio encoding is implemented, then adaptability to different playback configurations is improved, but encoding complexity increases

Engineering Contradiction:
Improveplayback configuration adaptabilityVSAvoidencoding complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The encoded audio stream with directional metadata is designed to be universally compatible with multiple playback configurations including monophonic, stereo, and multi-channel systems. The same encoded data adapts to different playback environments without requiring separate encoding processes

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The spatial encoding and directional metadata extraction are performed in advance during the encoding phase. This preliminary processing prepares the audio data for various playback scenarios beforehand, eliminating the need for complex real-time processing at the playback end

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP2898509B1Audio coding with gain profile extraction and transmission for speech enhancement at the decoder
Publication Date: 2016.10.12 DOLBY INTERNATIONAL AB
  • EP2898509B1 patent drawingFigure 1
  • EP2898509B1 patent drawingFigure 2
  • EP2898509B1 patent drawingFigure 3~4

AI summary

The invention provides a layered audio coding format with a monophonic layer and at least one sound field layer. A plurality of audio signals is decomposed, in accord- ance with decomposition parameters controlling the quantitative properties of an or- thogonal energy-compacting transform, into rotated audio signals. Further, a time- variable gain profile specifying constructively how the rotated audio signals may be processed to attenuate undesired audio content is derived. The monophonic layer may comprise one of the rotated signals and the gain profile. The sound field layer may comprise the rotated signals and the decomposition parameters. In one embodiment, the gain profile comprises a cleaning gain profile with the main purpose of eliminating non-speech components and/or noise. The gain profile may also com- prise mutually independent broadband gains. Because signals in the audio coding format can be mixed with a limited computational effort, the invention may advanta- geously be applied in a tele-conferencing application.