Dynamic Audio-Visual Positioning for Video Conference Mutual Presence

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video conference and telepresence technologies often fail to maintain focus and engagement due to challenges in presenting audio and video data effectively, leading to inefficient interactions among participants.

Innovation Solution

The implementation of a system that includes a presentation device with multiple imaging devices and audio transducers, which dynamically adjusts image and audio positioning to align with the user's gaze and environment, blending backgrounds and accentuating users to enhance focus and engagement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If traditional video conference systems present audio and video data, then participants can communicate remotely, but focus and engagement are lost due to ineffective presentation

Engineering Contradiction:
Improvefocus and engagementVSAvoidinteraction effectiveness
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system dynamically adjusts audio and video presentation based on user gaze detection and environmental context. The presentation device modifies audio spatial positioning and video rendering in real-time to align with where the user is looking and the local environment, creating an adaptive experience that maintains focus and engagement throughout the interaction.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system customizes the audio-visual presentation for each user's specific location and gaze direction. By detecting where a user is looking and blending virtual audio/video content with the local physical environment, the system creates locally optimized presentation quality that enhances engagement for each individual user rather than using a one-size-fits-all approach.

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If multiple imaging devices and audio transducers are deployed, then mutual presence can be simulated, but device complexity increases

Engineering Contradiction:
Improvemutual presence simulationVSAvoidsystem configuration
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The presentation device integrates multiple imaging devices and audio transducers into a single unified system that performs multiple functions: capturing user gaze, rendering audio spatially, displaying video content, and blending virtual与现实 environments. This multi-functional integration enables mutual presence simulation without requiring separate systems for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system uses an intermediary processing layer that coordinates between multiple imaging devices and audio transducers. This intermediary layer manages the complexity by synthesizing data from various sensors and coordinating their output to create a cohesive audio-visual experience, shielding users from the underlying system complexity while maintaining adaptability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11785179B1Image and audio data processing to create mutual presence in a video conference
Publication Date: 2023.10.10 AMAZON TECH INC
  • US11785179B1 patent drawing
  • US11785179B1 patent drawing
  • US11785179B1 patent drawing

AI summary

Systems and methods to improve video conferencing may include various improvements and modifications to image and audio data to simulate mutual presence among users. For example, automated positioning and presentation of image data of users via presentation devices may present users at a position, size, and scale to simulate mutual presence of users within a same physical space. In addition, rendering image data of users from virtual camera positions via presentation devices may simulate mutual eye gaze and mutual presence of users. Further, blending a background, or accentuating a user, within image data that is presented via presentation devices may match visual characteristics of the environment or the user to simulate mutual presence of users. Moreover, dynamic steering of individual audio transducers of an audio transducer array, and simulating audio point sources using a combination of speakers, may simulate direct communication and mutual presence of users.