Audio Signal Processing for Videoconference Spatial Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current teleconference systems are expensive and cumbersome to install, and they fail to provide a natural spatial alignment of audio and video signals, limiting their use to enterprise-level settings and offering a disappointing experience for casual users due to the lack of precise audio directionality and high installation costs.

Innovation Solution

A method and system that dynamically adapt audio signals to match the virtual location and size of participants on a display during a videoconference, using standard components and image processing to align audio and video signals, allowing for a natural spatial restitution of audio and improved immersion without the need for precise microphone and speaker placement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If professional telepresence systems with precise microphone and speaker placement are used, then spatial alignment of audio and video signals is improved, but installation cost and complexity increase

Engineering Contradiction:
Improvespatial alignment precisionVSAvoidinstallation complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical approach of precisely positioning microphones and speakers in specific geometric arrangements with a signal processing approach. The system captures audio from multiple microphones and processes them through amplitude panning and delay adjustments to create virtual speaker positions, eliminating the need for precise physical placement of audio equipment.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates virtual copies of audio sources by processing signals from physical microphones to simulate audio coming from specific locations. Through amplitude panning and delay manipulation, the system generates virtual speaker positions that correspond to the visual layout on screen, without requiring physical speakers to be placed at those exact locations.

Inventive Principle:
Principle #26Copying

2Manufacturing precision

If multiple loudspeakers are used to create natural spatial impression, then audio spatial perception is improved, but system cost and complexity increase

Engineering Contradiction:
Improvespatial perception accuracyVSAvoidnumber of loudspeakers
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent segments the audio signal processing by assigning different processed audio channels to different loudspeakers. The system divides the multi-channel audio signal into separate streams, each adjusted for amplitude and delay, and routes them to corresponding loudspeakers to create spatial audio effects without requiring a large number of speakers.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes audio signal parameters including amplitude and delay for different loudspeaker channels. By adjusting these parameters dynamically, the system creates the perception of audio coming from specific spatial locations using a standard multi-channel loudspeaker configuration, without needing additional specialized equipment.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If manual audio remixing is performed to align audio sources with display positions, then spatial coherence is improved, but processing time increases

Engineering Contradiction:
Improveaudio-video alignment accuracyVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary setup by establishing the geometric relationship between the camera, display, and loudspeaker positions before the videoconference begins. This pre-configuration allows the system to automatically calculate and apply the appropriate amplitude panning and delay adjustments for each audio channel without requiring manual remixing during the actual conference.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements an automatic system that self-adjusts audio positioning based on pre-configured geometric parameters. The system automatically calculates the optimal amplitude and delay settings for each loudspeaker channel based on the known positions of the camera and display, eliminating the need for manual audio remixing by operators.

Inventive Principle:
Principle #25Self-service

4Manufacturing precision

If fixed microphone and speaker positions are prescribed, then spatial reproduction quality is improved, but adaptability to different installations decreases

Engineering Contradiction:
Improvespatial reproduction qualityVSAvoidinstallation flexibility
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent uses adjustable parameters including amplitude panning coefficients and delay times that can be modified based on the specific installation configuration. The system allows users to input their own geometric parameters (camera position, display position, loudspeaker positions) and automatically adjusts the audio processing parameters to optimize spatial reproduction for that specific setup.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements a dynamic system where audio processing parameters are not fixed but can be adjusted based on the installation environment. The system dynamically calculates and applies appropriate amplitude and delay settings for each loudspeaker channel based on the configured geometric relationship between audio and video components, allowing adaptation to various installation scenarios.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP2352290B1Method and apparatus for matching audio and video signals during a videoconference
Publication Date: 2012.11.21 SWISSCOM AG
  • EP2352290B1 patent drawingFigure 1
  • EP2352290B1 patent drawingFigure 2
  • EP2352290B1 patent drawingFigure 3A~3B

AI summary

An apparatus comprising: a TV set (1) as display; a set of loudspeakers (2); one or several microphones (4) freely disposable around said TV set; one or several cameras (3) freely disposable around said TV set; an equipment (14) integrating in a single housing image and audio processing means and a broadband Internet access card with a plurality of interfaces for said TV set, said cameras, said microphones and said amplifier; wherein said processing means are arranged for obtaining a virtual location of an audio source and for manipulating audio signals so as to give the impression that an audio signal originates from said virtual location; wherein said processing means are further arranged for manipulating said audio signal, so as to dynamically adapt the audio width to the displayed video signal.