Video Conference Layout for Spatial Audio Participant Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video conference systems do not effectively utilize spatial audio capture capabilities to enhance user interface display and audio rendering, leading to a lack of intuitive understanding of participant positions and directions.
Innovation Solution
The system identifies user devices with spatial audio capture capabilities and displays their video data in enlarged windows with wider background regions, based on the degree of freedom tracking (3DoF or 6DoF), allowing for enhanced spatial audio rendering and visual tracking of participant movements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If standard video conference display is used, then all participants are displayed uniformly, but users cannot intuitively understand spatial audio directions and participant positions
Solution Approach 1:
The patent applies local quality by differentiating the display format based on each participant's spatial audio capture capability. Participants with spatial audio capability receive enlarged format windows with wider background regions, while others receive standard format windows. This localized differentiation provides spatial direction information only where available, enabling users to intuitively understand participant positions and audio directions without overwhelming the entire interface with spatial information for all participants.
2Measurement precision
If spatial audio capture capability is implemented, then audio direction tracking is improved, but device complexity increases
Solution Approach 1:
The patent segments the participant list into two distinct groups: those with spatial audio capture capability and those without. This segmentation allows the system to apply enhanced spatial tracking and enlarged display format only to the relevant subset of participants, rather than implementing complex spatial audio processing for all participants uniformly. This reduces overall system complexity while maintaining high measurement precision for those who have the capability.
3Adaptability or versatility
If enlarged format windows are displayed for spatial audio participants, then spatial audio rendering is enhanced, but display space for other participants is reduced
Solution Approach 1:
The patent resolves this contradiction by applying local quality differentiation - enlarged format windows are displayed only for participants with spatial audio capture capability, while standard format windows are used for participants without this capability. This ensures that enhanced spatial audio rendering is provided where the capability exists, while minimizing the impact on overall display area by limiting the enlarged format to only those necessary participants.
Data Source
AI summary
An apparatus and method is disclosed. An example embodiment of the method may comprise receiving audio and video data from a plurality of user devices as part of a conference call, the video data representing a user of the respective user device and identifying one or more of the user devices as having a spatial audio capture capability. Another operation may comprise displaying, or causing display of, the video data from the user devices in different respective windows of a user interface. The respective windows for the identified one or more user devices may be displayed in an enlarged format so as to have a wider background region than for windows for the user devices without a spatial capture capability.


