Virtual Camera Gaze Correction for Realistic Video Conferencing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video conferencing methods require extensive hardware or special display devices to achieve realistic eye contact between users, which is not feasible with minimal hardware expenditure and computing power.

Innovation Solution

A video conferencing method that uses a processing unit to process video image data by recognizing the head of a user, applying artificial neural networks to generate latency vectors representing gaze direction and pose, and virtually shifting the image recording device's perspective to create a realistic eye contact effect without additional hardware, using latency spaces and machine learning to adjust head poses and gaze directions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple cameras or special display devices are used to achieve realistic eye contact, then the eye contact effect is improved, but the hardware complexity and cost increase

Engineering Contradiction:
Improveeye contact effectVSAvoidhardware complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent creates a virtual copy of the camera positioned at the display device location. This virtual camera captures the user's image from the perspective of the display, allowing the user to appear as if they are looking at the other party's eyes. This software-based virtual camera replaces the need for multiple physical cameras or special display hardware, achieving realistic eye contact effect while minimizing hardware requirements.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical/optical system of multiple physical cameras and special display devices with a computational image processing system. By using image capture, virtual camera positioning, and digital image manipulation, the system achieves the same eye contact effect without the complexity of multiple hardware components. The mechanical arrangement of cameras and displays is substituted with software-based image synthesis and transformation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If multiple cameras are used to capture images through the display, then eye contact is achieved, but the device complexity and hardware expenditure increase

Engineering Contradiction:
Improveeye contact effectVSAvoidhardware expenditure
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

Instead of purchasing and installing multiple physical cameras, the patent creates a virtual copy of the camera function through software. This virtual camera is positioned computationally at the location where the display device is physically located, allowing the system to capture and process images as if from that perspective. This approach eliminates the need for additional hardware cameras while achieving the same functional result.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent substitutes the physical mechanical system of multiple cameras with a computational image processing approach. The system captures a single image from the user's perspective, then uses software to simulate the viewpoint of the display device and generate the appropriate corrected image. This replacement of mechanical hardware with computational methods significantly reduces hardware expenditure while maintaining the eye contact effect.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If image processing is performed to adjust gaze direction, then realistic eye contact is achieved, but computing power requirements increase

Engineering Contradiction:
Improveeye contact effectVSAvoidcomputing power
Core Design Contradiction:
ReliabilityVSPower

Solution Approach 1:

The patent performs preliminary action by capturing the user's image from their natural viewing position and pre-processing it to account for the display geometry. The system calculates the virtual camera position and prepares the image transformation parameters in advance, so that when the image needs to be displayed, the correction can be applied efficiently. This preliminary setup reduces the real-time computational burden during actual video conferencing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a virtual camera model that replicates the optical and geometric properties of a physical camera positioned at the display. This virtual camera model pre-encodes the transformation relationships between the user's actual viewpoint and the display perspective. By using this pre-established virtual model, the system avoids complex real-time calculations and can efficiently generate corrected images with lower computing power requirements.

Inventive Principle:
Principle #26Copying

Data Source

PatentEP4637132A1Video conference method and video conference system
Publication Date: 2025.10.22 CASABLANCA AI GMBH
  • EP4637132A1 patent drawingFigure 1~2
  • EP4637132A1 patent drawingFigure 3
  • EP4637132A1 patent drawingFigure 4

AI summary

The invention relates to a video conferencing method in which the video image data recorded by a first image recording device (3) are processed by a processing unit (14) and transmitted to a second display device (8). In the processed video image data, a target viewing direction of a first user (5) appears as if the first image recording device (3) were arranged on a straight line (18) passing through an eye of the first user and through an eye of a second user (9) displayed on a first display device (4). During the processing of the video image data, a source latency vector of a latency space is obtained in an encoder (33), which represents the pose of the head and/or the viewing direction (16).A target latency vector of the latency space is calculated from the source latency vector and a target gaze direction and/or target pose of the head such that the target latency vector represents the target gaze direction and/or the target pose of the head. Then, in a decoder (34), an intermediate representation of the head is obtained using a head model based on the target latency vector, and this intermediate representation is converted by a warp unit (43) into an output representation of the head using the source latency vector, the target latency vector, and the source appearance parameters.