Spatial Audio DTX Encoding for Immersive Background Noise

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies lack efficient methods for transmitting and rendering immersive conversational speech with spatial localization, as existing DTX systems are not designed for parametric spatial audio codecs, leading to incompatibility with low bit-rate communication and loss of immersivity in background noise generation.

Innovation Solution

A DTX system for parametric spatial audio, specifically for DirAC, that classifies frames as active or inactive, generating parametric descriptions for inactive frames to maintain spatial coherence, using soundfield parameter generators and synthetic noise generation to reduce bit-rate while preserving spatial immersion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If Discontinuous Transmission (DTX) mode is used to reduce bit-rate for inactive frames, then transmission efficiency is improved, but spatial coherence and immersivity are lost in background noise generation

Engineering Contradiction:
Improvetransmission efficiencyVSAvoidspatial coherence
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies preliminary action by generating and transmitting Silence Insertion Descriptor (SID) frames that contain parametric spatial information (direction of arrival, diffuseness) before the actual inactive frames occur. These pre-transmitted parameters enable the decoder to reconstruct spatially coherent background noise during DTX periods, maintaining spatial coherence while achieving bit-rate reduction.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes parameters by transmitting compressed parametric representations of spatial audio characteristics (direction, diffuseness, energy) instead of full audio signals during inactive frames. This parameter-based approach allows efficient transmission of spatial coherence information at reduced bit-rates, enabling accurate reconstruction of the spatial sound field during DTX mode.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If parametric spatial audio coding (DirAC) is used to represent sound field, then spatial resolution is improved, but data rate increases significantly

Engineering Contradiction:
Improvespatial resolutionVSAvoiddata rate
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential spatial parameters (direction of arrival, diffuseness, energy distribution) from the full audio signal for transmission during inactive frames. By separating and transmitting only these critical spatial characteristics rather than complete audio data, the system achieves high spatial resolution with significantly reduced data rates during DTX mode.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by transmitting spatial parameter information at a reduced frequency during inactive frames compared to active speech frames. SID frames are sent periodically rather than at every frame interval, providing sufficient spatial resolution for background noise reconstruction while minimizing data transmission requirements.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If Comfort Noise Generation (CNG) is used for inactive frames, then bit-rate is reduced, but spatial immersion is lost

Engineering Contradiction:
Improvebit-rate reductionVSAvoidspatial immersion
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent introduces an intermediary mechanism by using SID frames as carriers of spatial parameter information during inactive frames. These SID frames mediate between the encoder and decoder, transmitting compressed spatial characteristics that enable the CNG to generate background noise with preserved spatial immersion, thus preventing information loss while maintaining bit-rate reduction.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies universality by designing the SID frame structure to carry multiple types of spatial information (direction, diffuseness, energy) in a single compact data structure. This multi-functional parameter set enables comprehensive reconstruction of spatial audio characteristics during DTX mode, maintaining spatial immersion across various inactive frame conditions with efficient bit-rate utilization.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12586595B2Apparatus, method and computer program for encoding an audio signal or for decoding an encoded audio scene
Publication Date: 2026.03.24 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US12586595B2 patent drawing
  • US12586595B2 patent drawing
  • US12586595B2 patent drawing

AI summary

There are disclosed an apparatus for generating an encoded audio scene, and an apparatus for decoding and/or processing an encoded audio scene; as well as related methods and non-transitory storage units storing instructions which, when executed by a processor, cause the processor to perform a related method. An apparatus for processing an encoded audio scene may include, in a first frame, a first soundfield parameter representation and an encoded audio signal, wherein a second frame is an inactive frame, the apparatus including: an activity detector for detecting that the second frame is the inactive frame; a synthetic signal synthesizer for synthesizing a synthetic audio signal for the second frame using the parametric description for the second frame; an audio decoder for decoding the encoded audio signal for the first frame; and a spatial renderer for spatially rendering the audio signal for the first frame using the first soundfield parameter representation and using the synthetic audio signal for the second frame, or a transcoder for generating a meta data assisted output format including the audio signal for the first frame, the first soundfield parameter representation for the first frame, the synthetic audio signal for the second frame, and a second soundfield parameter representation for the second frame.