Spatial Audio in Multipoint Videoconferencing Layouts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current multipoint videoconferencing systems cannot associate a participant's voice with their location on the display, leading to a suboptimal user experience due to the lack of spatial audio cues.

Innovation Solution

The method provides spatially resolved audio by differentiating audio streams and attenuating or delaying them based on the position of the speaking endpoint within the videoconference layout, ensuring that audio appears to emanate from the correct location on the screen.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If mixed audio is transmitted in a 2x2 layout with dynamically changing participants, then the audio can include all current speakers, but the audio cannot deliver any impression on the location of the image source on the screen

Engineering Contradiction:
Improvespatial location informationVSAvoidaudio processing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent applies local quality by differentiating audio processing for different spatial locations. Each audio stream is processed individually with specific attenuation and delay parameters based on its corresponding video position, rather than treating all audio uniformly. This allows the system to preserve spatial location information by applying location-specific processing characteristics to each participant's audio.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the mixed audio into separate audio streams corresponding to individual participants. Instead of processing a single mixed audio signal, the system separates and processes each participant's audio independently, applying spatial processing parameters based on their position in the video layout. This segmentation enables precise control over the spatial characteristics of each audio source.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If spatially resolved audio is implemented by differentiating and attenuating audio streams based on position, then audio location perception is improved, but the audio processing complexity increases

Engineering Contradiction:
Improveaudio location perceptionVSAvoidaudio stream processing
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements parameter changes by dynamically adjusting audio processing parameters (attenuation and delay) based on the spatial position of each participant. The system calculates and applies different gain and time delay values for each audio stream according to its corresponding video position, enabling accurate spatial audio perception through continuous parameter adaptation.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies preliminary action by pre-calculating and storing attenuation and delay parameters for different video positions before actual audio processing. The system prepares spatial processing parameters in advance based on the video layout, so that when audio needs to be processed, the appropriate parameters are already available, reducing real-time processing complexity.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If audio streams are differentiated and processed individually for each speaker position, then spatial correspondence between audio and video is achieved, but the processing time and computational resources increase

Engineering Contradiction:
Improvespatial correspondence accuracyVSAvoidaudio processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-calculating spatial processing parameters based on video positions before audio processing occurs. The system prepares attenuation and delay values in advance according to the video layout configuration, so that when audio streams need to be processed, the parameters are already determined, significantly reducing real-time processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements parameter changes by using predetermined parameter sets for different video positions. Instead of calculating complex spatial parameters in real-time for each audio stream, the system selects from pre-defined parameter sets corresponding to different positions, maintaining high spatial correspondence accuracy while reducing computational time.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS7612793B2Spatially correlated audio in multipoint videoconferencing
Publication Date: 2009.11.03 HEWLETT PACKARD DEVELOPMENT COMPANY LP
  • US7612793B2 patent drawing
  • US7612793B2 patent drawing
  • US7612793B2 patent drawing

AI summary

The disclosed method provides audio location perception to an endpoint in a multipoint videoconference by providing a plurality of audio streams to the endpoint, wherein each of the audio streams corresponds to one of the loudspeakers at the endpoint. The audio streams are differentiated so as to emphasize broadcasting of the audio streams through one or more loudspeakers closest to a position of a speaking endpoint in a videoconference layout that is displayed at the endpoint. For example, the audio broadcast at a loudspeaker that is at a far-side of the screen might be attenuated or time delayed compared to audio broadcast at a loudspeaker that is located at a near-side of the display. The disclosure also provides a multipoint control unit (MCU) that processes audio signals from two or more endpoints according to the positions in a layout of the endpoints and then transmits processed audio streams to the endpoints in a way that allows endpoints to broadcast spatially correlated audio.