Cross-View Conference Audio Positioning With Visual Talker Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In cross-view conference arrangements with multiple cameras and directional microphones, accurately assigning audio from participants to directional audio channels is challenging due to the complexity of capturing and matching audio with the correct visual perspective.
Innovation Solution
A conference system that uses video cameras and directional microphones to detect active talkers, positionally classify directional audio, and code it into positional audio channels to match the visual view, ensuring accurate audio assignment to participants.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If multiple directional microphones are positioned around the room to capture audio from participants, then audio capture coverage is improved, but audio assignment accuracy to directional channels deteriorates
Solution Approach 1:
The system uses video feedback by detecting heads in the visual view to determine which directional audio should be assigned to which audio channel. The audio assignment is dynamically adjusted based on the visual detection results, creating a closed-loop feedback mechanism that resolves the ambiguity of audio channel assignment when multiple directional microphones are active simultaneously.
Solution Approach 2:
The video camera and head detection algorithm serve as an intermediary between the multiple directional microphones and the audio output channels. Instead of directly processing audio signals from multiple microphones, the system uses visual information as an intermediary to mediate the assignment of audio sources to appropriate directional channels.
2Measurement precision
If directional microphones form directional beams to receive audio from specific areas, then audio directionality is improved, but complexity of matching audio with visual perspective deteriorates
Solution Approach 1:
The system creates a visual copy or representation of the physical room layout by detecting head positions in the video view. This visual map is then used to guide the audio routing, where the detected head positions serve as a simplified copy of the actual participant locations, making the audio-visual matching process more manageable despite the complexity of multiple directional beams.
3Measurement precision
If audio is positionally classified to match heads in the view, then participant identification accuracy is improved, but processing complexity deteriorates
Solution Approach 1:
The system performs preliminary head detection and positioning using video analysis before finalizing the audio assignment. By pre-establishing the spatial locations of participants through visual detection, the system simplifies the subsequent audio processing step, as the target positions for audio routing are already determined from the visual data.
Data Source
AI summary
A method performed by a conference system having video cameras positioned around a room to capture views of areas of the room, the method comprising: receiving directional audio from directional microphones positioned adjacent to the areas and configured to form directional beams to receive the directional audio from the areas; detecting an active talker in an area based on the directional audio; capturing a view of the area with a video camera; detecting one or more heads across the view; positionally classifying the directional audio received by the directional beams adjacent to the area to visually match the one or more heads in the view to produce positionally classified audio; coding the positionally classified audio into positional audio channels; and transmitting the view and the positional audio channels.


