Context-Based Audio-Visual Framing for Teleconferencing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current videoconferencing systems using timer-based solutions for selecting and framing audio-visual data from multiple angles during videoconferences are not entirely satisfactory, leading to suboptimal participant focus and engagement.

Innovation Solution

A method that involves receiving video and audio data frames, detecting faces and sound sources, updating an audio-visual map with facial and talker weight values, and selecting sub-frames for transmission based on thresholds to prioritize active participants, ensuring clear focus on speakers and relevant subjects.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If timer-based solutions are used to automatically select and frame views, then automation is improved, but participant focus and engagement deteriorate

Engineering Contradiction:
Improveautomatic view selectionVSAvoidparticipant focus
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The system continuously monitors audio data to detect active speakers and uses this feedback to dynamically adjust video framing. The audio-visual map maintains weight values that reflect current speaker activity, enabling the system to respond in real-time to changing meeting dynamics rather than relying on fixed timer-based schedules.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system automatically identifies and frames active speakers through audio analysis and face detection without requiring manual intervention. The audio-visual map self-updates weight values based on detected faces and sound sources, enabling autonomous decision-making for view selection that adapts to meeting participation patterns.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If multiple video streams from different angles are captured, then adaptability is improved, but device complexity increases

Engineering Contradiction:
Improvemulti-angle captureVSAvoidsystem configuration
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The audio-visual map serves multiple functions: it tracks participant locations, detects active speakers, maintains facial weight values, and guides video framing decisions. This single data structure handles diverse tasks that would otherwise require separate systems, reducing overall complexity while supporting multi-angle capture and dynamic view selection.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If facial weight values and talker weight values are tracked, then measurement precision is improved, but data processing requirements increase

Engineering Contradiction:
Improvespeaker detection accuracyVSAvoiddata volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system maintains weight values specifically for detected faces and sound sources rather than processing all video data uniformly. By focusing computational resources on identified targets (faces with facial weight values, sound sources with talker weight values), the system achieves precise speaker detection while minimizing unnecessary data processing.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11676369B2Context based target framing in a teleconferencing environment
Publication Date: 2023.06.13 HEWLETT PACKARD DEVELOPMENT COMPANY LP
  • US11676369B2 patent drawing
  • US11676369B2 patent drawing
  • US11676369B2 patent drawing

AI summary

A method for determining camera framing in a teleconferencing system comprises a process loop which includes acquiring an audio-visual frame from a captured a video data frame; detecting objects and extracting image features of the objects within the video data frame, ingesting the audio-visual frame into a context-based audio-visual map in an intelligent manner, and selecting targets from within the map for inclusion in an audio-video stream for transmission to a remote endpoint.