Positional Audio Metadata Generation for Video Conference Spatial Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In video conference sessions, the spatial positioning of audio output often does not match the visual output, leading to difficulties for remote participants in identifying the speaker, resulting in a less immersive experience due to reliance on incomplete or absent visual cues.

Innovation Solution

The generation of positional audio metadata that accurately reflects the sound source locations by using directional microphones and dividing the video output into tracking sectors, allowing for precise audio panning based on detected head positions and sound sources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If spatial positioning of audio output is not matched with visual output, then device complexity is reduced, but remote participant immersion and speaker identification deteriorate

Engineering Contradiction:
Improveaudio spatial positioning systemVSAvoidspeaker identification accuracy
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent introduces an intermediary system that captures audio signals from multiple directional microphones, determines sound source positions, and generates positional audio metadata. This intermediary processing layer bridges the gap between simple audio capture and accurate spatial audio reproduction, enabling speaker identification without requiring complex end-device modifications.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces mechanical/spatial audio positioning with a computational approach. Instead of physically positioning audio sources in space, the system uses algorithms to analyze microphone signals, determine sound source locations, and render spatial audio through metadata that guides playback devices. This substitution reduces hardware complexity while maintaining spatial accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If visual cues are used to identify speakers, then speaker identification is possible, but immersion deteriorates due to reliance on incomplete visual information

Engineering Contradiction:
Improvevisual cue completenessVSAvoidparticipant immersion
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The patent segments the audio signal processing into distinct functional components: directional microphone arrays capture spatial audio information, sound source position determination algorithms process the captured signals, and positional audio metadata generation creates structured spatial information. This segmentation enables accurate speaker identification independent of visual information, enhancing immersion by providing complete audio cues.

Inventive Principle:
Principle #1Segmentation

3Reliability

If audio positions are accurately matched with visual cues, then participant immersion improves, but device complexity increases

Engineering Contradiction:
Improveaudio-visual spatial matchingVSAvoidpositional audio metadata generation system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts spatial audio information from the audio signal at the source endpoint, separating the complex processing tasks from the playback endpoint. By taking out the sound source position determination and metadata generation functions to the transmitting end, the receiving end only needs to reproduce audio based on provided metadata, significantly reducing overall system complexity while maintaining accurate audio-visual spatial matching.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11115625B1Positional audio metadata generation
Publication Date: 2021.09.07 CISCO TECHNOLOGY INC
  • US11115625B1 patent drawing
  • US11115625B1 patent drawing
  • US11115625B1 patent drawing

AI summary

At a video conference endpoint including a camera, a microphone array, and one or more microphone assemblies, the video conference endpoint may divide a video output of the camera into one or more tracking sectors and detect a head position for each participant in the video output. The video conference endpoint may determine within which tracking sector each detected head position is located. The video conference endpoint may determine active sound source positions of the actively speaking participants based on sound being detected or captured by the microphone array and microphone assemblies, and may determine within which tracking sector the active sound source positions are located. For each tracking sector that contains an active sound source position, the video conference endpoint may update the positional audio metadata for that particular tracking sector based on the active sound source positions and the detected head positions located in that tracking sector.