Layered Audio Coding with Gain Profiles for Continuous Spatial Speech
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems for teleconferencing lack the ability to provide a spatially layered, encoded audio signal that offers a continuous and varied mix of sound fields and monophonic layers over time, failing to maintain a perceptually continuous listening experience.
Innovation Solution
An audio encoding system comprising a spatial analyzer, adaptive rotation stage, and analysis stage that decomposes audio signals into rotated signals and generates a time-variable gain profile, enabling efficient encoding and decoding of spatially layered audio signals, allowing for continuous and adaptive playback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If spatially layered audio encoding is implemented, then audio quality and listening experience are improved, but system complexity increases
Solution Approach 1:
The audio signal is divided into multiple layers including a monophonic layer and directional metadata layers. This segmentation allows the system to process and transmit different audio components separately, improving overall audio quality while managing complexity through hierarchical processing
Solution Approach 2:
The patent introduces spatial dimensions to audio encoding by adding directional metadata that describes the spatial characteristics of audio sources. This dimensional enhancement transforms traditional monophonic audio into spatial audio without requiring complete system redesign
2Reliability
If continuous audio transmission is maintained, then perceptual continuity is improved, but bandwidth consumption increases
Solution Approach 1:
The patent extracts only the essential directional metadata from the full audio signal. Instead of transmitting complete multi-channel audio data, only the critical spatial parameters are extracted and transmitted, maintaining perceptual continuity while significantly reducing bandwidth requirements
Solution Approach 2:
The system changes the representation parameters of audio data by using compact directional metadata formats. This parameter transformation allows continuous audio perception to be achieved with reduced data transmission by focusing on essential spatial parameters rather than full waveform data
3Adaptability or versatility
If layered audio encoding is implemented, then adaptability to different playback configurations is improved, but encoding complexity increases
Solution Approach 1:
The encoded audio stream with directional metadata is designed to be universally compatible with multiple playback configurations including monophonic, stereo, and multi-channel systems. The same encoded data adapts to different playback environments without requiring separate encoding processes
Solution Approach 2:
The spatial encoding and directional metadata extraction are performed in advance during the encoding phase. This preliminary processing prepares the audio data for various playback scenarios beforehand, eliminating the need for complex real-time processing at the playback end
Data Source
Figure 1
Figure 2
Figure 3~4
AI summary
The invention provides a layered audio coding format with a monophonic layer and at least one sound field layer. A plurality of audio signals is decomposed, in accord- ance with decomposition parameters controlling the quantitative properties of an or- thogonal energy-compacting transform, into rotated audio signals. Further, a time- variable gain profile specifying constructively how the rotated audio signals may be processed to attenuate undesired audio content is derived. The monophonic layer may comprise one of the rotated signals and the gain profile. The sound field layer may comprise the rotated signals and the decomposition parameters. In one embodiment, the gain profile comprises a cleaning gain profile with the main purpose of eliminating non-speech components and/or noise. The gain profile may also com- prise mutually independent broadband gains. Because signals in the audio coding format can be mixed with a limited computational effort, the invention may advanta- geously be applied in a tele-conferencing application.