Audio Stream Clustering for Spatial Conference Rendering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current multi-party audio conferences typically render voices as monaural audio streams, making it difficult for listeners to distinguish between multiple speakers, and existing solutions fail to effectively manage bandwidth limitations in transmitting spatialized audio signals.
Innovation Solution
A conference controller is used to create and manage 2D or 3D audio scenes by assigning upstream audio signals to specific spatial locations, reducing the number of downstream audio signals transmitted by mixing and selecting audio signals based on activity and angular separation, and generating metadata for spatialized rendering.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If spatialisation techniques are used to simulate different people talking from different rendered locations, then intelligibility of speech is improved, but bandwidth consumption increases
Solution Approach 1:
The audio scene is segmented into a limited number of discrete spatial channels (e.g., 7 channels arranged in a specific geometric pattern). Instead of transmitting continuous spatial audio data for each participant, the system divides the audio space into distinct segments or channels, each carrying mixed audio content from multiple participants. This segmentation allows spatial rendering with reduced bandwidth requirements.
Solution Approach 2:
Multiple upstream audio signals from different participants are merged and mixed into a smaller number of downstream spatial channels. The conference controller combines audio content from multiple sources into each spatial channel, applying appropriate spatialization parameters. This merging reduces the total number of audio streams that need to be transmitted while maintaining the ability to render participants at different spatial locations.
2Manufacturing precision
If multiple upstream audio signals are transmitted individually to each listener, then spatial rendering accuracy is maintained, but network bandwidth is exceeded
Solution Approach 1:
The system transitions from transmitting audio signals in the time domain (individual streams) to transmitting them in the spatial domain (angular/azimuthal distribution). By organizing audio content according to spatial dimensions and angular positions rather than participant count, the system achieves efficient bandwidth utilization while maintaining spatial rendering accuracy. The audio signals are distributed across spatial channels defined by angular separation.
3Measurement precision
If all conference participants are placed at different spatial locations, then speaker differentiation is maximized, but device complexity increases
Solution Approach 1:
The system changes the parameters used to define spatial locations from participant-based identifiers to geometric/angular parameters. Each spatial channel is defined by specific angular coordinates and geometric relationships rather than being tied to individual participants. This parameter transformation simplifies scene management by using consistent geometric patterns that can accommodate any number of participants through parameter adjustment rather than structural reconfiguration.
Data Source
Figure 1a~1b
Figure 2~3a
Figure 3b~4
AI summary
The present document relates to methods and systems for setting up and managing two-dimensional or three-dimensional scenes for audio conferences. A conference controller (111, 175) configured to place L upstream audio signals (123, 173) within a 2D or 3D conference scene to be rendered to a listener (211) is described. The conference controller (111, 175) is configured to set up a X-point conference scene; assign L upstream audio signals (123, 173) to X talker locations (212); determine a maximum number N of downstream audio signals (124, 174) to be transmitted to the listener (211); determine N downstream audio signals (124, 174) from the L assigned upstream audio signals (123, 173); determine N updated talker locations for the N downstream audio signals (124, 174); and generate metadata identifying the updated talker locations and enabling an audio processing unit (121, 171) to generate a spatialized audio signal.