Cross-View Conference Audio Positioning With Visual Talker Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In cross-view conference arrangements with multiple cameras and directional microphones, accurately assigning audio from participants to directional audio channels is challenging due to the complexity of capturing and matching audio with the correct visual perspective.

Innovation Solution

A conference system that uses video cameras and directional microphones to detect active talkers, positionally classify directional audio, and code it into positional audio channels to match the visual view, ensuring accurate audio assignment to participants.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If multiple directional microphones are positioned around the room to capture audio from participants, then audio capture coverage is improved, but audio assignment accuracy to directional channels deteriorates

Engineering Contradiction:
Improveaudio capture coverageVSAvoidaudio assignment accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The system uses video feedback by detecting heads in the visual view to determine which directional audio should be assigned to which audio channel. The audio assignment is dynamically adjusted based on the visual detection results, creating a closed-loop feedback mechanism that resolves the ambiguity of audio channel assignment when multiple directional microphones are active simultaneously.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The video camera and head detection algorithm serve as an intermediary between the multiple directional microphones and the audio output channels. Instead of directly processing audio signals from multiple microphones, the system uses visual information as an intermediary to mediate the assignment of audio sources to appropriate directional channels.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If directional microphones form directional beams to receive audio from specific areas, then audio directionality is improved, but complexity of matching audio with visual perspective deteriorates

Engineering Contradiction:
Improveaudio directionalityVSAvoidaudio-visual matching complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system creates a visual copy or representation of the physical room layout by detecting head positions in the video view. This visual map is then used to guide the audio routing, where the detected head positions serve as a simplified copy of the actual participant locations, making the audio-visual matching process more manageable despite the complexity of multiple directional beams.

Inventive Principle:
Principle #26Copying

3Measurement precision

If audio is positionally classified to match heads in the view, then participant identification accuracy is improved, but processing complexity deteriorates

Engineering Contradiction:
Improveparticipant identification accuracyVSAvoidaudio processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary head detection and positioning using video analysis before finalizing the audio assignment. By pre-establishing the spatial locations of participants through visual detection, the system simplifies the subsequent audio processing step, as the target positions for audio routing are already determined from the visual data.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20260059256A1Method of controlling directional sound pickup in cross-view conference meetings
Publication Date: 2026.02.26 CISCO TECHNOLOGY INC
  • US20260059256A1 patent drawing
  • US20260059256A1 patent drawing
  • US20260059256A1 patent drawing

AI summary

A method performed by a conference system having video cameras positioned around a room to capture views of areas of the room, the method comprising: receiving directional audio from directional microphones positioned adjacent to the areas and configured to form directional beams to receive the directional audio from the areas; detecting an active talker in an area based on the directional audio; capturing a view of the area with a video camera; detecting one or more heads across the view; positionally classifying the directional audio received by the directional beams adjacent to the area to visually match the one or more heads in the view to produce positionally classified audio; coding the positionally classified audio into positional audio channels; and transmitting the view and the positional audio channels.