Spatial Audio Conferencing System for Video Calls

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current teleconferencing systems lack the ability to effectively provide spatial audio cues, making it difficult for participants to discern who is speaking during multi-user conferences, leading to a less immersive and less engaging user experience.

Innovation Solution

The implementation of a Spatial Audio Conferencing System (SACS) that uses spatial audio encoding and decoding techniques, such as Dolby Atmos or DTS:X, to assign specific audio sources to positions within a virtual 2D or 3D space within the conference interface, allowing for accurate localization of sound sources based on participant positions and orientations, enhancing the audio experience by providing clear cues on who is speaking.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional teleconferencing systems are used, then the system complexity is low, but the audio localization precision is insufficient making it difficult to discern who is speaking

Engineering Contradiction:
Improveaudio localization precisionVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies parameter changes by transforming audio signals from traditional mono/stereo formats into spatial audio formats (Dolby Atmos, DTS:X) that encode directional information. The system modifies audio parameters to include spatial coordinates, elevation angles, and azimuth angles, enabling precise localization of sound sources to specific positions in a virtual 3D space corresponding to participant locations.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements dimensionality change by transitioning from traditional 1D/2D audio presentation to 3D spatial audio. The system creates a virtual three-dimensional audio space where sound sources are positioned according to participant locations, adding spatial dimensionality that enables listeners to accurately identify who is speaking based on directional audio cues.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If spatial audio encoding and decoding techniques are implemented, then the audio localization precision is improved, but the device complexity increases

Engineering Contradiction:
Improvesound source localization accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses an intermediary approach by introducing a spatial audio processing module that acts as a mediator between the audio capture stage and the audio output stage. This intermediary component handles the complex encoding and decoding of spatial audio, managing the computational burden centrally rather than distributing it across all devices, thereby reducing overall system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes audio parameters from traditional formats to spatial audio formats that encode directional information. By transforming the audio signal representation to include spatial coordinates and directional data, the system achieves accurate sound source localization while managing processing complexity through efficient parameter transformation rather than complex signal manipulation.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12088762B2Systems and methods for videoconferencing with spatial audio
Publication Date: 2024.09.10 VERIZON PATENT & LICENSING INC
  • US12088762B2 patent drawing
  • US12088762B2 patent drawing
  • US12088762B2 patent drawing

AI summary

A system may provide for the generation of spatial audio for audiovisual conferences, video conferences, etc. (referred to herein simply as “conferences”). Spatial audio may include audio encoding and/or decoding techniques in which a sound source may be specified at a location, such as on a two-dimensional plane and/or within a three-dimensional field, and/or in which a direction or target for a given sound source may be specified. A conference participant's position within a conference user interface (“UI”) may be set as the source of sound associated with the conference participant, such that different conference participants may be associated with different sound source positions within the conference UI.