Distributed Smartphone Audio Conferencing with Acoustic Zone Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current web-conferencing applications face issues with sound quality and spatial rendering, particularly when multiple participants are in the same room, leading to poor audio capture and playback, with participants often needing to mute devices or rely on non-spatially rendered audio.
Innovation Solution
A method for hosting teleconferences among multiple client devices by grouping them into acoustic spaces, processing audio streams to enhance sound quality and spatial rendering, involving techniques like signal processing, beamforming, and echo cancellation, to improve audio capture and playback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If multiple client devices are used in the same acoustic space, then audio capture coverage is improved, but audio quality and spatial rendering deteriorate due to redundancy and echo
Solution Approach 1:
The system segments the acoustic space into multiple zones and assigns different client devices to capture audio from different segments. Each device focuses on a specific spatial region, dividing the overall audio capture task to avoid redundancy and echo while maintaining comprehensive coverage.
Solution Approach 2:
Different client devices are assigned different roles based on their local spatial position and characteristics. Devices closer to specific participants or sound sources prioritize capturing those local audio signals, while devices in other positions focus on different areas, creating localized optimization throughout the acoustic space.
2Manufacturing precision
If all client devices capture and render audio, then spatial rendering is improved, but echo and redundancy increase
Solution Approach 1:
The system extracts and identifies the active sound source from multiple audio streams, then selectively processes only the relevant audio signals. By taking out the primary audio source and prioritizing its capture while suppressing other redundant signals, the system achieves spatial rendering without echo.
Solution Approach 2:
The audio capture and rendering configuration dynamically adapts based on the active sound source location. The system continuously monitors which device is closest to the active speaker and adjusts the audio routing accordingly, making the system flexible rather than static in its audio handling approach.
3Area of stationary object
If client devices are distributed across acoustic spaces, then audio coverage is improved, but system complexity increases
Solution Approach 1:
Each client device autonomously determines its own role and function based on its spatial position and the location of active sound sources. Devices self-configure their audio capture and rendering behavior without requiring complex centralized control, reducing overall system complexity while maintaining distributed audio coverage.
Data Source
AI summary
Described is a method of hosting a teleconference among a plurality of client devices arranged in two or more acoustic spaces, each client device having an audio capturing capability and/or an audio rendering capability, the method comprising: grouping the plurality of client devices into two or more groups based on their belonging to respective acoustic spaces, receiving first audio streams from the plurality of client devices, generating second audio streams from the first audio streams for rendering by respective client devices among the plurality of client devices, based on the grouping of the plurality of client devices into the two or more groups, and outputting the generated second audio streams to respective client devices. Further described are corresponding computation devise, computer programs, and computer-readable storage media.


