Spatial Audio Rendering for Video Conferencing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current communication systems, such as videoconferencing and teleconferencing, face challenges in providing spatial audio, leading to reduced collaboration quality, speech intelligibility, and overall user experience due to spatial collisions caused by low angular separation of estimated directions of arrival of multiple speakers.

Innovation Solution

A communication system that includes a first computing device with an array of audio output devices and a processor, which receives transmitted speech data and metadata describing the estimated direction of arrival of speech from a microphone array at a second computing device, and renders audio at an array of audio devices to eliminate spatial collisions by repositioning audio based on the estimated directions of arrival.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If audio from multiple speakers is rendered using estimated directions of arrival, then spatial audio positioning is improved, but spatial collisions occur due to low angular separation causing reduced speech intelligibility

Engineering Contradiction:
Improvespatial audio positioning accuracyVSAvoidspeech intelligibility
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent transitions from 2D angular separation to 3D spherical coordinate representation of speaker positions. By using azimuth and elevation angles together with distance information, speakers that are angularly close in 2D can be spatially separated in 3D space, eliminating collisions while preserving speech intelligibility through accurate spatial positioning.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system dynamically adjusts audio rendering parameters based on real-time estimated directions of arrival. When speakers are detected with low angular separation, the system adaptively modifies their spatial positioning and audio characteristics to prevent collisions, maintaining both spatial accuracy and speech intelligibility under varying conference conditions.

Inventive Principle:
Principle #15Dynamics

2Reliability

If audio rendering maintains accurate spatial positioning of multiple speakers, then collaboration quality is improved, but spatial collisions reduce overall user experience

Engineering Contradiction:
Improvecollaboration qualityVSAvoidspatial collisions
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent extracts collision detection and resolution as separate processing steps from the audio rendering pipeline. By identifying overlapping spatial positions beforehand and resolving them through coordinate transformation and positional adjustment, the system eliminates harmful spatial collisions while preserving accurate spatial positioning for high-quality collaboration.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs preliminary spatial collision detection before audio rendering by analyzing estimated directions of arrival and predicting potential overlaps. It proactively adjusts speaker positioning and applies corrective rendering techniques to prevent spatial collisions before they degrade user experience, maintaining collaboration quality throughout the conference.

Inventive Principle:
Principle #9Preliminary anti-action

Data Source

PatentUS11317232B2Eliminating spatial collisions due to estimated directions of arrival of speech
Publication Date: 2022.04.26 HEWLETT PACKARD DEVELOPMENT COMPANY LP
  • US11317232B2 patent drawing
  • US11317232B2 patent drawing
  • US11317232B2 patent drawing

AI summary

A communication system may include, in an example, a first computing device communicatively coupled, via a network, to at least a second computing device maintained at a geographically distinct location than the first computing device; the first computing device including: an array of audio output devices and a processor to receive transmitted speech data and metadata describing an estimated direction of arrival (DOA) of speech from a plurality of speakers at an array of microphones at the second computing device and render audio at the array of audio output devices associated with the first computing device by eliminating spatial collision during rendering; said spatial collision arising due to the low angular separation of the estimated DOA of a plurality of speakers.