Multi-device Spatial Browsing for Conference Capture

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional conference recording and viewing technologies fail to provide sufficient information and control to users, particularly in interpreting speaker-oriented interactions and focusing on non-speaking participants, due to limitations in capturing and displaying spatial relationships and nuances of communication.

Innovation Solution

A multi-device capture and spatial browsing system that detects and enlists available cameras and microphones to create a composite media stream, allowing users to navigate a 3-dimensional representation of a conference, with features like automatic reorientation of the 3D view to emphasize the current speaker and manual navigation controls for detailed interaction analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If conventional fixed or isolated views of each participant are used, then device complexity is reduced, but information completeness and user understanding of communication nuances deteriorate

Engineering Contradiction:
Improvecommunication nuancesVSAvoidrecording and viewing system
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent merges multiple camera views and spatial information into a single panoramic display that shows all participants and their spatial relationships simultaneously. This allows viewers to see who is looking at whom, who is interrupting whom, and other communication nuances without requiring complex multi-device setups at each viewing location.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transitions from 2D isolated participant views to a 360-degree panoramic view that adds spatial dimensionality. This allows the display to represent the physical arrangement of participants around a table, preserving directional information about who is looking at or speaking to whom, thereby capturing communication nuances that flat 2D views lose.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Ease of operation

If thumbnail-size videos of all attendees are displayed, then device complexity is reduced, but ease of operation for focusing on non-speaking participants deteriorates

Engineering Contradiction:
Improvefocus control on participantsVSAvoiduser interface
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent implements dynamic control of context views where users can adjust the size, position, and focus of participant displays in real-time. The interface allows viewers to expand non-speaking participants' views when needed while maintaining the overall panoramic context, making it easy to focus on specific participants without requiring a complex fixed layout.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If omnidirectional camera and specially positioned IP camera are used, then measurement precision of spatial relationships is improved, but device complexity and cost increase

Engineering Contradiction:
Improvespatial relationshipsVSAvoidcamera system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses standard webcams and cameras that participants already have on their devices, making them multi-functional for both capturing video content and determining spatial relationships. This eliminates the need for specialized omnidirectional cameras or precisely positioned IP cameras, as the existing camera arrays can be calibrated to provide sufficient spatial information for the panoramic display.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9065976B2Multi-device capture and spatial browsing of conferences
Publication Date: 2015.06.23 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9065976B2 patent drawing
  • US9065976B2 patent drawing
  • US9065976B2 patent drawing

AI summary

Multi-device capture and spatial browsing of conferences is described. In one implementation, a system detects cameras and microphones, such as the webcams on participants' notebook computers, in a conference room, group meeting, or table game, and enlists an ad-hoc array of available devices to capture each participant and the spatial relationships between participants. A video stream composited from the array is browsable by a user to navigate a 3-dimensional representation of the meeting. Each participant may be represented by a video pane, a foreground object, or a 3-D geometric model of the participant's face or body displayed in spatial relation to the other participants in a 3-dimensional arrangement analogous to the spatial arrangement of the meeting. The system may automatically re-orient the 3-dimensional representation as needed to best show a currently interesting event.