Audio Stream Clustering for Spatial Conference Rendering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current multi-party audio conferences typically render voices as monaural audio streams, making it difficult for listeners to distinguish between multiple speakers, and existing solutions fail to effectively manage bandwidth limitations in transmitting spatialized audio signals.

Innovation Solution

A conference controller is used to create and manage 2D or 3D audio scenes by assigning upstream audio signals to specific spatial locations, reducing the number of downstream audio signals transmitted by mixing and selecting audio signals based on activity and angular separation, and generating metadata for spatialized rendering.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If spatialisation techniques are used to simulate different people talking from different rendered locations, then intelligibility of speech is improved, but bandwidth consumption increases

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidbandwidth consumption
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The audio scene is segmented into a limited number of discrete spatial channels (e.g., 7 channels arranged in a specific geometric pattern). Instead of transmitting continuous spatial audio data for each participant, the system divides the audio space into distinct segments or channels, each carrying mixed audio content from multiple participants. This segmentation allows spatial rendering with reduced bandwidth requirements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Multiple upstream audio signals from different participants are merged and mixed into a smaller number of downstream spatial channels. The conference controller combines audio content from multiple sources into each spatial channel, applying appropriate spatialization parameters. This merging reduces the total number of audio streams that need to be transmitted while maintaining the ability to render participants at different spatial locations.

Inventive Principle:
Principle #5Merging (Combining)

2Manufacturing precision

If multiple upstream audio signals are transmitted individually to each listener, then spatial rendering accuracy is maintained, but network bandwidth is exceeded

Engineering Contradiction:
Improvespatial rendering accuracyVSAvoidnetwork bandwidth
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The system transitions from transmitting audio signals in the time domain (individual streams) to transmitting them in the spatial domain (angular/azimuthal distribution). By organizing audio content according to spatial dimensions and angular positions rather than participant count, the system achieves efficient bandwidth utilization while maintaining spatial rendering accuracy. The audio signals are distributed across spatial channels defined by angular separation.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If all conference participants are placed at different spatial locations, then speaker differentiation is maximized, but device complexity increases

Engineering Contradiction:
Improvespeaker differentiationVSAvoidscene management complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system changes the parameters used to define spatial locations from participant-based identifiers to geometric/angular parameters. Each spatial channel is defined by specific angular coordinates and geometric relationships rather than being tied to individual participants. This parameter transformation simplifies scene management by using consistent geometric patterns that can accommodate any number of participants through parameter adjustment rather than structural reconfiguration.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP2829049B1Clustering of audio streams in a 2d/3d conference scene
Publication Date: 2021.05.26 DOLBY LABORATORIES LICENSING CORP
  • EP2829049B1 patent drawingFigure 1a~1b
  • EP2829049B1 patent drawingFigure 2~3a
  • EP2829049B1 patent drawingFigure 3b~4

AI summary

The present document relates to methods and systems for setting up and managing two-dimensional or three-dimensional scenes for audio conferences. A conference controller (111, 175) configured to place L upstream audio signals (123, 173) within a 2D or 3D conference scene to be rendered to a listener (211) is described. The conference controller (111, 175) is configured to set up a X-point conference scene; assign L upstream audio signals (123, 173) to X talker locations (212); determine a maximum number N of downstream audio signals (124, 174) to be transmitted to the listener (211); determine N downstream audio signals (124, 174) from the L assigned upstream audio signals (123, 173); determine N updated talker locations for the N downstream audio signals (124, 174); and generate metadata identifying the updated talker locations and enabling an audio processing unit (121, 171) to generate a spatialized audio signal.