Virtual Reality User Image Generation From Expression Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users wearing virtual reality (VR) headsets often lack a webcam for video conferencing and are reluctant to appear on calls with their headset visible, leading to the need for efficient avatar solutions that are not universally compatible across different communication platforms.
Innovation Solution
A software-based method generates and renders a user's image on a 2D display using expression tracking and a generative model, allowing for photorealistic avatars without a full 3D animation pipeline, and can be integrated with various communication systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If a full 3D animation pipeline is used to generate photorealistic avatars, then the visual realism is improved, but the computational cost and system complexity increase
Solution Approach 1:
The patent extracts only the essential components needed for photorealistic avatar generation from the full 3D animation pipeline. Instead of implementing complete 3D modeling, rigging, and animation systems, the invention uses a generative model that directly synthesizes photorealistic images from input data, eliminating unnecessary pipeline complexity while maintaining visual quality
Solution Approach 2:
The patent replaces the mechanical 3D animation pipeline with a software-based generative model. Instead of using traditional computer graphics rendering engines and 3D models, the system employs machine learning-based image generation that directly produces photorealistic avatars, substituting complex mechanical processing with intelligent software synthesis
2Manufacturing precision
If a full 3D animation pipeline is implemented, then photorealistic avatars can be generated, but compatibility across different communication platforms decreases
Solution Approach 1:
The patent creates a universal avatar generation system that produces output compatible with multiple communication platforms. The generative model generates standardized image formats and protocols that can be integrated with various video conferencing systems (Zoom, Teams, etc.), making the photorealistic avatar technology platform-agnostic and broadly applicable
Solution Approach 2:
The patent adjusts key parameters of the generative model to optimize for both photorealism and platform compatibility. By tuning parameters such as image resolution, frame rate, and output format, the system achieves high visual quality while ensuring broad compatibility across different communication platforms and devices
3Productivity
If expression tracking and generative models are used instead of full 3D pipelines, then computational cost is reduced, but the challenge remains of generating realistic expressions without 3D data
Solution Approach 1:
The patent introduces expression tracking data as an intermediary between the input video stream and the generative model. The tracking system extracts facial landmark positions and expression parameters, which serve as conditional inputs to guide the generative model in producing realistic facial expressions that match the user's actual emotions, bridging the gap between 2D processing and 3D-like realism
Data Source
Figure 1
Figure 2A
Figure 2B
AI summary
Images are generated and rendered on a first device having a two-dimensional display. The first device receives from a second device expression data indicative of a current facial expression of a user of the second device, where the second device has a three-dimensional display. The expression data is input to a generative model trained on an enrollment image indicative of a baseline image of the user's face. Facial image information is received from the generative model that is usable to render a two-dimensional image of the current facial expression on the first device. The facial image information is sent to the first device for rendering of the two-dimensional image of the current facial expression on the two-dimensional display in context of an on-going session of a communications system.