3D Face Model Encoding for Eye Contact and Lighting Realism

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing immersive telepresence systems face challenges in achieving proper pose and eye contact of rendered faces and illuminating them with an appropriate lighting model in virtual environments.

Innovation Solution

Encoding and decoding semantic description data representative of a 3D geometric and photometric face model using neural networks to synthesize faces with accurate head pose and lighting in immersive videos, reducing data transmission requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If 2D video data is transmitted for immersive telepresence, then data transmission volume is high, but rendering quality and realism are insufficient

Engineering Contradiction:
Improverendering qualityVSAvoiddata transmission volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential semantic information from 2D video frames (3D face model parameters, head pose, iris location, lighting conditions) and transmits this compressed representation instead of full video data. This extraction approach maintains rendering quality while dramatically reducing data transmission volume.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transforms video data from pixel-space representation to parametric 3D model representation, changing the data format from 2D image arrays to structured 3D geometric and photometric parameters. This parameter transformation enables efficient compression and high-quality reconstruction.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If head pose and iris location are adjusted for eye contact, then immersive experience is improved, but processing complexity increases

Engineering Contradiction:
Improveeye contact capabilityVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary extraction of head pose and iris location parameters from video frames during the encoding phase. These pre-processed parameters are then directly applied during decoding to achieve eye contact adjustment, eliminating the need for complex real-time processing at the receiving end.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses a parametric 3D face model as an intermediary representation that simplifies the adjustment of head pose and iris location. Instead of directly manipulating 2D video pixels, the system adjusts parameters of the 3D model, which automatically renders the appropriate eye contact effect.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If lighting is replaced to match virtual environment, then realism is improved, but computational requirements increase

Engineering Contradiction:
Improvelighting realismVSAvoidcomputational energy
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent extracts lighting parameters (light direction, intensity, color temperature) from the video frames during encoding and transmits them as compact data. This extraction eliminates the need for complex lighting analysis and re-rendering at the receiving end, reducing computational energy while maintaining lighting realism.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates a parametric copy of the lighting conditions captured in the video frames and applies this copied lighting model to the 3D face model. This copying approach preserves the original lighting characteristics without requiring energy-intensive re-rendering from scratch.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20260065583A1Methods and apparatuses for immersive videoconference
Publication Date: 2026.03.05 INTERDIGITAL CE PATENT HOLDINGS SAS
  • US20260065583A1 patent drawing
  • US20260065583A1 patent drawing
  • US20260065583A1 patent drawing

AI summary

Methods and apparatuses for encoding/decoding semantic description data representative of a 3D geometric and photometric face model for immersive telepresence are provided. In an embodiment, video data comprising a face of a user is encoded by extracting semantic description data representative of a 3D geometric and photometric model of the face of the user. In another embodiment, an immersive video is decoded from the semantic description data by, determining a head pose of the face of a remote user in an immersive video; determining a parametric model of a lighting environment of the immersive video; synthesizing the face of the remote user with the head pose and the parametric model; and generating a modified immersive video comprising an image of the synthesized face of the user in the immersive video. In an embodiment, the generation of the immersive video is made recurrent by taking at input the synthesized face and the immersive video at a previous frame.