Distributed Smartphone Audio Conferencing with Acoustic Zone Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current web-conferencing applications face issues with sound quality and spatial rendering, particularly when multiple participants are in the same room, leading to poor audio capture and playback, with participants often needing to mute devices or rely on non-spatially rendered audio.

Innovation Solution

A method for hosting teleconferences among multiple client devices by grouping them into acoustic spaces, processing audio streams to enhance sound quality and spatial rendering, involving techniques like signal processing, beamforming, and echo cancellation, to improve audio capture and playback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Area of stationary object

If multiple client devices are used in the same acoustic space, then audio capture coverage is improved, but audio quality and spatial rendering deteriorate due to redundancy and echo

Engineering Contradiction:
Improveacoustic space coverageVSAvoidaudio quality
Core Design Contradiction:
Area of stationary objectVSManufacturing precision

Solution Approach 1:

The system segments the acoustic space into multiple zones and assigns different client devices to capture audio from different segments. Each device focuses on a specific spatial region, dividing the overall audio capture task to avoid redundancy and echo while maintaining comprehensive coverage.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different client devices are assigned different roles based on their local spatial position and characteristics. Devices closer to specific participants or sound sources prioritize capturing those local audio signals, while devices in other positions focus on different areas, creating localized optimization throughout the acoustic space.

Inventive Principle:
Principle #3Local quality

2Manufacturing precision

If all client devices capture and render audio, then spatial rendering is improved, but echo and redundancy increase

Engineering Contradiction:
Improvespatial renderingVSAvoidecho
Core Design Contradiction:
Manufacturing precisionVSObject-generated harmful factors

Solution Approach 1:

The system extracts and identifies the active sound source from multiple audio streams, then selectively processes only the relevant audio signals. By taking out the primary audio source and prioritizing its capture while suppressing other redundant signals, the system achieves spatial rendering without echo.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The audio capture and rendering configuration dynamically adapts based on the active sound source location. The system continuously monitors which device is closest to the active speaker and adjusts the audio routing accordingly, making the system flexible rather than static in its audio handling approach.

Inventive Principle:
Principle #15Dynamics

3Area of stationary object

If client devices are distributed across acoustic spaces, then audio coverage is improved, but system complexity increases

Engineering Contradiction:
Improveaudio coverageVSAvoidsystem complexity
Core Design Contradiction:
Area of stationary objectVSDevice complexity

Solution Approach 1:

Each client device autonomously determines its own role and function based on its spatial position and the location of active sound sources. Devices self-configure their audio capture and rendering behavior without requiring complex centralized control, reducing overall system complexity while maintaining distributed audio coverage.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11991315B2Audio conferencing using a distributed array of smartphones
Publication Date: 2024.05.21 DOLBY LABORATORIES LICENSING CORP
  • US11991315B2 patent drawing
  • US11991315B2 patent drawing
  • US11991315B2 patent drawing

AI summary

Described is a method of hosting a teleconference among a plurality of client devices arranged in two or more acoustic spaces, each client device having an audio capturing capability and/or an audio rendering capability, the method comprising: grouping the plurality of client devices into two or more groups based on their belonging to respective acoustic spaces, receiving first audio streams from the plurality of client devices, generating second audio streams from the first audio streams for rendering by respective client devices among the plurality of client devices, based on the grouping of the plurality of client devices into the two or more groups, and outputting the generated second audio streams to respective client devices. Further described are corresponding computation devise, computer programs, and computer-readable storage media.