Video Conferencing Avatar Control Through 3D Mesh Gestures

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video conferencing systems lack the ability to seamlessly transition between live video streams and digital avatars, particularly in situations where video quality is compromised, and they do not effectively utilize machine learning to synchronize facial expressions and gestures for a more engaging online presence.

Innovation Solution

A system that utilizes an Avatar Engine to detect facial expressions and gestures in a video stream, generating commands to render a digital avatar on a target device using a 3D mesh model, allowing for synchronized and high-quality representation of a user's presence, even when live video transmission is disrupted.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If live video streams are used for video conferencing, then real-time communication is achieved, but video quality is compromised when network conditions are poor

Engineering Contradiction:
Improvevideo qualityVSAvoidreal-time communication efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent introduces a digital avatar as an intermediary between the user and the video conferencing system. When live video quality degrades, the system seamlessly transitions to displaying a pre-rendered digital avatar that represents the user, maintaining communication effectiveness while avoiding poor quality video transmission.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary actions by capturing high-quality video streams in advance and generating digital avatars before they are needed. These pre-captured video segments and generated avatars are stored and ready for immediate deployment when network conditions deteriorate, ensuring seamless transition without real-time transmission delays.

Inventive Principle:
Principle #10Preliminary action

2Stability of the object's composition

If digital avatars are used to represent users, then consistent online presence is maintained, but the system complexity increases

Engineering Contradiction:
Improveonline presence consistencyVSAvoidsystem complexity
Core Design Contradiction:
Stability of the object's compositionVSDevice complexity

Solution Approach 1:

The patent creates a digital copy (avatar) of the user that replicates their visual presence. This copy is generated from high-quality pre-captured video segments and can be deployed consistently across different platforms and devices, ensuring stable online presence representation without requiring complex real-time processing at each endpoint.

Inventive Principle:
Principle #26Copying

3Manufacturing precision

If machine learning is used to synchronize facial expressions, then avatar realism is improved, but processing time increases

Engineering Contradiction:
Improvefacial expression synchronization accuracyVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs facial expression analysis and avatar synchronization in advance during the video capture phase. Machine learning models process and analyze facial expressions from captured video segments before they are needed, so that when the avatar is deployed, the synchronization is already prepared and can be rendered immediately without real-time processing delays.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250299405A1Video Conferencing Remote Gesture Control
Publication Date: 2025.09.25 ZOOM COMMUNICATIONS INC
  • US20250299405A1 patent drawing
  • US20250299405A1 patent drawing
  • US20250299405A1 patent drawing

AI summary

Transmission of an instance of a three-dimensional (3D) mesh model to a target computer device associated with a first user account is triggered. Changes in a video stream captured at a source computer device associated with a second user account, the second user account represented by a digital rendering at the target computer device according to the instance of the 3D mesh model received by the target computer device are detected. A command based on the detected changes in the video stream captured at the source computer device is identified, the at least one command corresponding to a portion of blendshapes. Transmission of the identified command to the target computer device associated with the first user account is triggered, the target computer device generating a local instantiation of the digital rendering according to the 3D mesh model and the blendshapes.