Distributed Meeting Capture via Multimedia Hub

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video conferencing systems are expensive and complex to set up, often require a dedicated room, and struggle with bandwidth limitations, resulting in poor identification of active talkers due to limited facial detail and multi-viewpoint capture challenges.

Innovation Solution

A distributed meeting capture system using multiple personal devices with cameras and microphones connected via a local multimedia hub for intelligent view switching and audio/video stream processing, which includes acoustic echo cancellation and pre-processing to determine active speakers and recompose streams for efficient network transmission.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If a single outbound audio and video stream is used from each endpoint, then bandwidth requirements are reduced, but the ability to identify active talkers and provide facial detail is insufficient

Engineering Contradiction:
Improvebandwidth consumptionVSAvoidfacial detail recognition
Core Design Contradiction:
Loss of energyVSMeasurement precision

Solution Approach 1:

The system segments the video conference into multiple virtual viewpoints, each capturing different participants or regions. Instead of transmitting a single wide-angle view, the system creates multiple segmented video streams from different virtual camera positions, allowing remote participants to see facial details of active talkers while managing bandwidth through selective transmission of only relevant viewpoint streams.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces an intermediary processing layer that acts as a virtual camera controller. This intermediary analyzes audio-visual data from all participants and dynamically determines which participant should be viewed from which virtual viewpoint. The intermediary then synthesizes and transmits only the necessary viewpoint streams, mediating between the need for detailed facial views and bandwidth constraints.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If multiple personal devices are used for distributed meeting capture, then multi-viewpoint capture capability is improved, but device complexity and setup requirements increase

Engineering Contradiction:
Improvemulti-viewpoint capture capabilityVSAvoidsystem configuration complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system makes each personal device universal by enabling it to function as multiple virtual cameras simultaneously. Through software processing, a single physical device can capture and generate multiple viewpoint perspectives (e.g., front view, side views, focused views of different participants). This eliminates the need for multiple physical cameras or complex multi-device setups, as one device performs the work of many.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system creates virtual copies of camera viewpoints through software processing. Instead of requiring physical copies of cameras positioned at different locations, the system generates virtual camera feeds by processing video data from existing devices. These virtual copies provide multiple perspectives without adding physical hardware complexity, allowing distributed meeting capture using standard personal devices.

Inventive Principle:
Principle #26Copying

3Productivity

If intelligent view switching is implemented to identify active talkers, then communication effectiveness is improved, but processing requirements and system complexity increase

Engineering Contradiction:
Improvecommunication effectivenessVSAvoidprocessing system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system implements self-service by enabling personal devices to autonomously perform view switching based on local audio-visual analysis. Each device can independently detect active talkers through audio levels and lip motion in video feeds, then automatically switch virtual camera viewpoints without requiring centralized control or complex external processing systems. This distributes the processing intelligence across devices, reducing overall system complexity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system uses feedback mechanisms where audio and video data from participants continuously inform view switching decisions. By monitoring audio activity levels and lip motion in real-time, the system receives feedback about who is speaking and automatically adjusts virtual camera viewpoints accordingly. This closed-loop feedback system enables intelligent view switching without requiring complex manual configuration or external control systems.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8451315B2System and method for distributed meeting capture
Publication Date: 2013.05.28 HEWLETT PACKARD DEVELOPMENT COMPANY LP
  • US8451315B2 patent drawing
  • US8451315B2 patent drawing
  • US8451315B2 patent drawing

AI summary

Embodiments of the present invention disclose a system and method for distributed meeting capture. According to one embodiment, the system includes a plurality of personal devices configured to capture video data and audio data associated with at least one operating user. A media hub includes a plurality of I/O ports and is configured to receive video and audio data from the plurality of personal devices. In addition, the media hub is configured to collect the video data and/or audio data from the plurality of personal devices and output at least one audio-visual data stream for facilitating video conferencing over a network.