Virtual Camera Gaze Correction for Realistic Video Conferencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional video conferencing methods require extensive hardware or special display devices to achieve realistic eye contact between users, which is not feasible with minimal hardware expenditure and computing power.
Innovation Solution
A video conferencing method that uses a processing unit to process video image data by recognizing the head of a user, applying artificial neural networks to generate latency vectors representing gaze direction and pose, and virtually shifting the image recording device's perspective to create a realistic eye contact effect without additional hardware, using latency spaces and machine learning to adjust head poses and gaze directions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple cameras or special display devices are used to achieve realistic eye contact, then the eye contact effect is improved, but the hardware complexity and cost increase
Solution Approach 1:
The patent creates a virtual copy of the camera positioned at the display device location. This virtual camera captures the user's image from the perspective of the display, allowing the user to appear as if they are looking at the other party's eyes. This software-based virtual camera replaces the need for multiple physical cameras or special display hardware, achieving realistic eye contact effect while minimizing hardware requirements.
Solution Approach 2:
The patent replaces the mechanical/optical system of multiple physical cameras and special display devices with a computational image processing system. By using image capture, virtual camera positioning, and digital image manipulation, the system achieves the same eye contact effect without the complexity of multiple hardware components. The mechanical arrangement of cameras and displays is substituted with software-based image synthesis and transformation.
2Reliability
If multiple cameras are used to capture images through the display, then eye contact is achieved, but the device complexity and hardware expenditure increase
Solution Approach 1:
Instead of purchasing and installing multiple physical cameras, the patent creates a virtual copy of the camera function through software. This virtual camera is positioned computationally at the location where the display device is physically located, allowing the system to capture and process images as if from that perspective. This approach eliminates the need for additional hardware cameras while achieving the same functional result.
Solution Approach 2:
The patent substitutes the physical mechanical system of multiple cameras with a computational image processing approach. The system captures a single image from the user's perspective, then uses software to simulate the viewpoint of the display device and generate the appropriate corrected image. This replacement of mechanical hardware with computational methods significantly reduces hardware expenditure while maintaining the eye contact effect.
3Reliability
If image processing is performed to adjust gaze direction, then realistic eye contact is achieved, but computing power requirements increase
Solution Approach 1:
The patent performs preliminary action by capturing the user's image from their natural viewing position and pre-processing it to account for the display geometry. The system calculates the virtual camera position and prepares the image transformation parameters in advance, so that when the image needs to be displayed, the correction can be applied efficiently. This preliminary setup reduces the real-time computational burden during actual video conferencing.
Solution Approach 2:
The patent creates a virtual camera model that replicates the optical and geometric properties of a physical camera positioned at the display. This virtual camera model pre-encodes the transformation relationships between the user's actual viewpoint and the display perspective. By using this pre-established virtual model, the system avoids complex real-time calculations and can efficiently generate corrected images with lower computing power requirements.
Data Source
Figure 1~2
Figure 3
Figure 4
AI summary
The invention relates to a video conferencing method in which the video image data recorded by a first image recording device (3) are processed by a processing unit (14) and transmitted to a second display device (8). In the processed video image data, a target viewing direction of a first user (5) appears as if the first image recording device (3) were arranged on a straight line (18) passing through an eye of the first user and through an eye of a second user (9) displayed on a first display device (4). During the processing of the video image data, a source latency vector of a latency space is obtained in an encoder (33), which represents the pose of the head and/or the viewing direction (16).A target latency vector of the latency space is calculated from the source latency vector and a target gaze direction and/or target pose of the head such that the target latency vector represents the target gaze direction and/or the target pose of the head. Then, in a decoder (34), an intermediate representation of the head is obtained using a head model based on the target latency vector, and this intermediate representation is converted by a warp unit (43) into an output representation of the head using the source latency vector, the target latency vector, and the source appearance parameters.