VR User Image Generation With Expression Tracking for Video Calls

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users wearing virtual reality (VR) devices often lack a separate webcam for video conferencing, and they are unlikely to appear on calls wearing the VR device.

Innovation Solution

A software-based method to generate and inject images in a video stream that includes a rendering of the user without wearing the VR headset, using expression tracking and rendering techniques without a full 3D animation pipeline.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a user wears a VR device for video conferencing, then the user can communicate using the VR device, but the user appears wearing the VR headset which is unacceptable

Engineering Contradiction:
Improvevideo conferencing capabilityVSAvoidunacceptable appearance
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The system creates a virtual copy of the user's face using generative AI models. Instead of showing the actual camera feed that includes the VR headset, the system generates a synthetic image that copies the user's facial features and expressions while removing the headset. This virtual copy is then displayed to other participants, solving the problem of unacceptable appearance while maintaining video conferencing capability.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system introduces an intermediary processing layer between the camera capture and the video transmission. This intermediary uses expression tracking to capture the user's facial movements and feeds them to a generative model that produces the final video output. This intermediary process allows the system to transform the raw camera input into an acceptable virtual representation that excludes the VR headset.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If a separate webcam is used for video conferencing, then the user can appear without the VR device, but the user lacks a separate webcam when only using the VR device

Engineering Contradiction:
Improveappearance without VR deviceVSAvoidhardware requirements
Core Design Contradiction:
Object-affected harmful factorsVSDevice complexity

Solution Approach 1:

The system replaces the mechanical hardware solution (separate webcam) with a software-based generative AI system. Instead of requiring additional physical camera devices, the system uses the existing camera feed combined with expression tracking and generative models to synthesize a virtual video stream that appears as if captured by a separate webcam not wearing the VR headset.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The VR device's own camera and sensors are used to generate the video stream needed for conferencing. The system leverages the expression tracking capabilities already present in the VR device to capture facial movements, eliminating the need for external webcams. The device serves itself by using its own hardware resources to produce the required video output.

Inventive Principle:
Principle #25Self-service

3Object-affected harmful factors

If cartoon avatars or photorealistic avatars with full 3D animation pipeline are used, then the user can appear without the VR device, but the avatar appearance varies between different communications systems

Engineering Contradiction:
Improveappearance without VR deviceVSAvoidconsistency across platforms
Core Design Contradiction:
Object-affected harmful factorsVSAdaptability or versatility

Solution Approach 1:

The system produces a universal video output format that can be consumed by any standard video conferencing application. By generating a realistic video stream rather than system-specific avatar formats, the solution works across Microsoft Teams, Zoom, and other platforms without requiring platform-specific adaptations. The generative model outputs standard video data that maintains consistency across different communication systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system changes the fundamental parameter of avatar representation from discrete 3D model formats to continuous generative video output. This parameter change allows the same underlying technology to produce consistent results across different platforms, as the output is a universal video stream rather than a platform-specific data format. The flexibility of generative AI enables adaptation to different rendering requirements while maintaining core consistency.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250292513A1Virtual reality user image generation
Publication Date: 2025.09.18 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20250292513A1 patent drawing
  • US20250292513A1 patent drawing
  • US20250292513A1 patent drawing

AI summary

Images are generated and rendered on a first device having a two-dimensional display. The first device receives from a second device expression data indicative of a current facial expression of a user of the second device, where the second device has a three-dimensional display. The expression data is input to a generative model trained on an enrollment image indicative of a baseline image of the user's face. Facial image information is received from the generative model that is usable to render a two-dimensional image of the current facial expression on the first device. The facial image information is sent to the first device for rendering of the two-dimensional image of the current facial expression on the two-dimensional display in context of an on-going session of a communications system.