Multi-device Spatial Browsing for Conference Capture
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional conference recording and viewing technologies fail to provide sufficient information and control to users, particularly in interpreting speaker-oriented interactions and focusing on non-speaking participants, due to limitations in capturing and displaying spatial relationships and nuances of communication.
Innovation Solution
A multi-device capture and spatial browsing system that detects and enlists available cameras and microphones to create a composite media stream, allowing users to navigate a 3-dimensional representation of a conference, with features like automatic reorientation of the 3D view to emphasize the current speaker and manual navigation controls for detailed interaction analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If conventional fixed or isolated views of each participant are used, then device complexity is reduced, but information completeness and user understanding of communication nuances deteriorate
Solution Approach 1:
The patent merges multiple camera views and spatial information into a single panoramic display that shows all participants and their spatial relationships simultaneously. This allows viewers to see who is looking at whom, who is interrupting whom, and other communication nuances without requiring complex multi-device setups at each viewing location.
Solution Approach 2:
The patent transitions from 2D isolated participant views to a 360-degree panoramic view that adds spatial dimensionality. This allows the display to represent the physical arrangement of participants around a table, preserving directional information about who is looking at or speaking to whom, thereby capturing communication nuances that flat 2D views lose.
2Ease of operation
If thumbnail-size videos of all attendees are displayed, then device complexity is reduced, but ease of operation for focusing on non-speaking participants deteriorates
Solution Approach 1:
The patent implements dynamic control of context views where users can adjust the size, position, and focus of participant displays in real-time. The interface allows viewers to expand non-speaking participants' views when needed while maintaining the overall panoramic context, making it easy to focus on specific participants without requiring a complex fixed layout.
3Measurement precision
If omnidirectional camera and specially positioned IP camera are used, then measurement precision of spatial relationships is improved, but device complexity and cost increase
Solution Approach 1:
The patent uses standard webcams and cameras that participants already have on their devices, making them multi-functional for both capturing video content and determining spatial relationships. This eliminates the need for specialized omnidirectional cameras or precisely positioned IP cameras, as the existing camera arrays can be calibrated to provide sufficient spatial information for the panoramic display.
Data Source
AI summary
Multi-device capture and spatial browsing of conferences is described. In one implementation, a system detects cameras and microphones, such as the webcams on participants' notebook computers, in a conference room, group meeting, or table game, and enlists an ad-hoc array of available devices to capture each participant and the spatial relationships between participants. A video stream composited from the array is browsable by a user to navigate a 3-dimensional representation of the meeting. Each participant may be represented by a video pane, a foreground object, or a 3-D geometric model of the participant's face or body displayed in spatial relation to the other participants in a 3-dimensional arrangement analogous to the spatial arrangement of the meeting. The system may automatically re-orient the 3-dimensional representation as needed to best show a currently interesting event.


