Videoconference Eye Contact via Virtual Avatar Gaze

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video conferencing systems fail to replicate the experience of real-world human conversations due to the absence of eye contact between participants, leading to a lack of immersive and natural interactions.

Innovation Solution

A system comprising N sensor devices and display devices, with a host computing device generating a virtual space where each participant's physical state, including eye contact, is captured and adapted into a virtual representation, allowing for realistic human-to-human interactions by simulating eye contact and other non-verbal cues.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional video conferencing systems are used with single camera capture and simultaneous video transmission, then system simplicity and ease of operation are maintained, but the immersion and naturalness of human-to-human interactions deteriorate due to missing eye contact

Engineering Contradiction:
Improveease of operationVSAvoidimmersion and naturalness of interactions
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent introduces a virtual representation (avatar) as an intermediary between the participant's actual appearance and the video feed shown to others. The avatar serves as a mediator that can independently control eye contact behavior, allowing the system to maintain operational simplicity while achieving natural eye contact through the virtual representation that follows the participant's gaze direction

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

Instead of making the actual video feed follow the participant's eye movements (which would be complex), the patent inverts the approach by making the virtual representation follow the eye movements. This reversal simplifies the system architecture while achieving the same effect of natural eye contact in the video conference

Inventive Principle:
Principle #13The other way round (Inversion)

2Reliability

If virtual representations with adapted physical states are introduced to simulate eye contact, then immersion and naturalness of interactions are improved, but device complexity increases due to additional processing requirements

Engineering Contradiction:
Improveimmersion and naturalness of interactionsVSAvoiddevice complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts the eye contact functionality from the complex video processing pipeline and implements it as a separate, dedicated virtual representation component. By isolating the eye contact simulation in the avatar's gaze control, the system adds the necessary complexity only where needed while keeping the rest of the video conferencing system simple and efficient

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates a simplified copy (virtual representation) of the participant that replicates only the essential feature needed for eye contact. This copy approach allows the system to achieve natural eye contact without processing and transmitting the full complexity of actual video feeds with eye movement tracking, reducing overall device complexity

Inventive Principle:
Principle #26Copying

3Measurement precision

If sensor devices capture physical state of body parts for virtual representation adaptation, then measurement precision of eye contact is improved, but device complexity and cost increase

Engineering Contradiction:
Improvemeasurement precisionVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent makes the existing camera serve multiple functions: it captures both the participant's video feed and the participant's eye movements. By enabling the single camera to perform dual detection tasks, the system achieves precise eye contact measurement without adding separate sensor devices, thereby avoiding increased device complexity and cost

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP4432650A1System for a videoconference with n participants, method of using the system for conducting a videoconference with n participants, host computing device, client computing device, method for implementing human-to-human interactions in a videoconference with n participants, host method, client method and computer program product
Publication Date: 2024.09.18 UNIVERSITY OF HEIDELBERG
  • EP4432650A1 patent drawingFigure 1
  • EP4432650A1 patent drawingFigure 2
  • EP4432650A1 patent drawingFigure 3

AI summary

An aspect of the present invention relates to a system (10) for a videoconference with N participants (P1, P2, P3, P4), N being an integer greater than or equal to 2, the system comprising: N sensor devices (121, 122, 123, 124), wherein each sensor device of the N sensor devices (121, 122, 123, 124) is assigned to one participant of the N participants (P1, P2, P3, P4), N display devices (141, 142, 143, 144), wherein each display device of the N display devices (141, 142, 143, 144) is assigned to one participant of the N participants (P1, P2, P3, P4), and at least one host computing device (16) configured to generate a virtual space (18), wherein the virtual space (18) comprises a virtual representation (V1, V2, V3, V4) for each of the N participants (P1, P2, P3, P4), and wherein to each of the N participants (P1, P2, P3, P4) one virtual representation (V1, V2, V3, V4) is assigned, wherein, for each participant (P1): the assigned sensor device (121) is configured to capture a physical state of at least one body part, such as an eye, of the participant (P1), the physical state of the at least one body part of the participant (P1) is determined and the assigned virtual representation (V1) is adapted based on the determining of the physical state of the at least one body part of the participant (P1), and the assigned display device (141) is configured to display, to the participant (P1), the virtual representation (V2, V3, V4) of at least one of the remaining participants (P2, P3, P4) of the videoconference in the virtual space (18). Further aspect relate to a method of using the system for conducting a videoconference with N participants, a host computing device, a client computing device, a method for implementing human-to-human interactions in a videoconference with N participants, a host method, a client method and a computer program product.