Context-Based Audio-Visual Framing for Teleconferencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current videoconferencing systems using timer-based solutions for selecting and framing audio-visual data from multiple angles during videoconferences are not entirely satisfactory, leading to suboptimal participant focus and engagement.
Innovation Solution
A method that involves receiving video and audio data frames, detecting faces and sound sources, updating an audio-visual map with facial and talker weight values, and selecting sub-frames for transmission based on thresholds to prioritize active participants, ensuring clear focus on speakers and relevant subjects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If timer-based solutions are used to automatically select and frame views, then automation is improved, but participant focus and engagement deteriorate
Solution Approach 1:
The system continuously monitors audio data to detect active speakers and uses this feedback to dynamically adjust video framing. The audio-visual map maintains weight values that reflect current speaker activity, enabling the system to respond in real-time to changing meeting dynamics rather than relying on fixed timer-based schedules.
Solution Approach 2:
The system automatically identifies and frames active speakers through audio analysis and face detection without requiring manual intervention. The audio-visual map self-updates weight values based on detected faces and sound sources, enabling autonomous decision-making for view selection that adapts to meeting participation patterns.
2Adaptability or versatility
If multiple video streams from different angles are captured, then adaptability is improved, but device complexity increases
Solution Approach 1:
The audio-visual map serves multiple functions: it tracks participant locations, detects active speakers, maintains facial weight values, and guides video framing decisions. This single data structure handles diverse tasks that would otherwise require separate systems, reducing overall complexity while supporting multi-angle capture and dynamic view selection.
3Measurement precision
If facial weight values and talker weight values are tracked, then measurement precision is improved, but data processing requirements increase
Solution Approach 1:
The system maintains weight values specifically for detected faces and sound sources rather than processing all video data uniformly. By focusing computational resources on identified targets (faces with facial weight values, sound sources with talker weight values), the system achieves precise speaker detection while minimizing unnecessary data processing.
Data Source
AI summary
A method for determining camera framing in a teleconferencing system comprises a process loop which includes acquiring an audio-visual frame from a captured a video data frame; detecting objects and extracting image features of the objects within the video data frame, ingesting the audio-visual frame into a context-based audio-visual map in an intelligent manner, and selecting targets from within the map for inclusion in an audio-video stream for transmission to a remote endpoint.


