Spatial Audio DTX Encoding for Immersive Background Noise
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies lack efficient methods for transmitting and rendering immersive conversational speech with spatial localization, as existing DTX systems are not designed for parametric spatial audio codecs, leading to incompatibility with low bit-rate communication and loss of immersivity in background noise generation.
Innovation Solution
A DTX system for parametric spatial audio, specifically for DirAC, that classifies frames as active or inactive, generating parametric descriptions for inactive frames to maintain spatial coherence, using soundfield parameter generators and synthetic noise generation to reduce bit-rate while preserving spatial immersion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If Discontinuous Transmission (DTX) mode is used to reduce bit-rate for inactive frames, then transmission efficiency is improved, but spatial coherence and immersivity are lost in background noise generation
Solution Approach 1:
The patent applies preliminary action by generating and transmitting Silence Insertion Descriptor (SID) frames that contain parametric spatial information (direction of arrival, diffuseness) before the actual inactive frames occur. These pre-transmitted parameters enable the decoder to reconstruct spatially coherent background noise during DTX periods, maintaining spatial coherence while achieving bit-rate reduction.
Solution Approach 2:
The patent changes parameters by transmitting compressed parametric representations of spatial audio characteristics (direction, diffuseness, energy) instead of full audio signals during inactive frames. This parameter-based approach allows efficient transmission of spatial coherence information at reduced bit-rates, enabling accurate reconstruction of the spatial sound field during DTX mode.
2Measurement precision
If parametric spatial audio coding (DirAC) is used to represent sound field, then spatial resolution is improved, but data rate increases significantly
Solution Approach 1:
The patent extracts only the essential spatial parameters (direction of arrival, diffuseness, energy distribution) from the full audio signal for transmission during inactive frames. By separating and transmitting only these critical spatial characteristics rather than complete audio data, the system achieves high spatial resolution with significantly reduced data rates during DTX mode.
Solution Approach 2:
The patent applies partial action by transmitting spatial parameter information at a reduced frequency during inactive frames compared to active speech frames. SID frames are sent periodically rather than at every frame interval, providing sufficient spatial resolution for background noise reconstruction while minimizing data transmission requirements.
3Productivity
If Comfort Noise Generation (CNG) is used for inactive frames, then bit-rate is reduced, but spatial immersion is lost
Solution Approach 1:
The patent introduces an intermediary mechanism by using SID frames as carriers of spatial parameter information during inactive frames. These SID frames mediate between the encoder and decoder, transmitting compressed spatial characteristics that enable the CNG to generate background noise with preserved spatial immersion, thus preventing information loss while maintaining bit-rate reduction.
Solution Approach 2:
The patent applies universality by designing the SID frame structure to carry multiple types of spatial information (direction, diffuseness, energy) in a single compact data structure. This multi-functional parameter set enables comprehensive reconstruction of spatial audio characteristics during DTX mode, maintaining spatial immersion across various inactive frame conditions with efficient bit-rate utilization.
Data Source
AI summary
There are disclosed an apparatus for generating an encoded audio scene, and an apparatus for decoding and/or processing an encoded audio scene; as well as related methods and non-transitory storage units storing instructions which, when executed by a processor, cause the processor to perform a related method. An apparatus for processing an encoded audio scene may include, in a first frame, a first soundfield parameter representation and an encoded audio signal, wherein a second frame is an inactive frame, the apparatus including: an activity detector for detecting that the second frame is the inactive frame; a synthetic signal synthesizer for synthesizing a synthetic audio signal for the second frame using the parametric description for the second frame; an audio decoder for decoding the encoded audio signal for the first frame; and a spatial renderer for spatially rendering the audio signal for the first frame using the first soundfield parameter representation and using the synthetic audio signal for the second frame, or a transcoder for generating a meta data assisted output format including the audio signal for the first frame, the first soundfield parameter representation for the first frame, the synthetic audio signal for the second frame, and a second soundfield parameter representation for the second frame.


