Audio Directionality Control for Multi-Display Video Conferencing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In telepresence conferencing systems, audio directionality often lags behind video switching, causing abrupt jumps and distracting artifacts when a new speaking participant is detected, disrupting the virtual meeting experience.

Innovation Solution

Implementing a method with audio mixers that utilize gain vectors and finite state machines to manage audio directionality across multiple loudspeakers, ensuring smooth transitions by pre-scaling audio and gradually directionalizing speech associated with the correct video display, even during short talk spurts or background noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If audio directionality tracks video display location, then audio spatial accuracy is improved, but audio transitions become abrupt when video switching is delayed

Engineering Contradiction:
Improveaudio directionality accuracyVSAvoidaudio transition smoothness
Core Design Contradiction:
Measurement precisionVSStability of the object's composition

Solution Approach 1:

The system performs preliminary action by pre-scaling audio signals and pre-positioning them in the audio scene before video switching occurs. When a new speaker is detected, the audio is already prepared and positioned, eliminating the delay between audio directionality change and video display update. This prevents abrupt transitions while maintaining accurate audio-spatial correspondence.

Inventive Principle:
Principle #10Preliminary action

2Stability of the object's composition

If video switching is delayed to prevent thrashing, then video stability is improved, but audio directionality lags behind video display

Engineering Contradiction:
Improvevideo stabilityVSAvoidaudio-video synchronization delay
Core Design Contradiction:
Stability of the object's compositionVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-positioning audio signals in the audio scene before video switching occurs. When a new speaker is detected, the audio is already prepared and positioned, eliminating the delay between audio directionality change and video display update. This decouples the audio timing from the delayed video switching while maintaining accurate audio-spatial correspondence.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements dynamic audio scene management where audio directionality can change independently and more rapidly than video switching. The audio mixer dynamically adjusts audio positions based on speaker detection, allowing audio to respond immediately while video maintains its stability threshold, creating independent but coordinated audio-video dynamics.

Inventive Principle:
Principle #15Dynamics

3Speed

If audio directionality changes rapidly with each speaker, then audio responsiveness is improved, but audio transitions cause distracting artifacts

Engineering Contradiction:
Improveaudio directionality response speedVSAvoidaudio artifacts and disorientation
Core Design Contradiction:
SpeedVSObject-affected harmful factors

Solution Approach 1:

The system applies beforehand cushioning by pre-scaling and pre-positioning audio signals before switching occurs. This preparation creates a smooth transition path that cushions the audio scene against abrupt changes. The gradual fading and positioning of audio signals prevents harsh transitions and distracting artifacts while maintaining responsive audio directionality tracking.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Data Source

PatentUS8289362B2Audio directionality control for a multi-display switched video conferencing system
Publication Date: 2012.10.16 CISCO TECHNOLOGY INC
  • US8289362B2 patent drawing
  • US8289362B2 patent drawing
  • US8289362B2 patent drawing

AI summary

In one embodiment, a method includes setting a target value for each audio source received from a plurality of remote participants to a telepresence conference, the gain coefficient array feeding a mixer associated with a loudspeaker associated with a display. A gain increment value is then set for each audio source, the gain increment value being equal to a difference between the target value and a current gain coefficient, the difference being divided by N, where N is an integer greater than one that represents a number of increments. Then, for each audio source, and for each of N iterations, the gain increment value is added to a current gain coefficient to produce a new current gain coefficient that is loaded into the mixer, such that after the N iterations the new current gain coefficient is equal to the target value. It is emphasized that this abstract is provided to comply with the rules requiring an abstract that will allow a searcher or other reader to quickly ascertain the subject matter of the technical disclosure.