Video Conferencing Avatar Control Through 3D Mesh Gestures
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video conferencing systems lack the ability to seamlessly transition between live video streams and digital avatars, particularly in situations where video quality is compromised, and they do not effectively utilize machine learning to synchronize facial expressions and gestures for a more engaging online presence.
Innovation Solution
A system that utilizes an Avatar Engine to detect facial expressions and gestures in a video stream, generating commands to render a digital avatar on a target device using a 3D mesh model, allowing for synchronized and high-quality representation of a user's presence, even when live video transmission is disrupted.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If live video streams are used for video conferencing, then real-time communication is achieved, but video quality is compromised when network conditions are poor
Solution Approach 1:
The patent introduces a digital avatar as an intermediary between the user and the video conferencing system. When live video quality degrades, the system seamlessly transitions to displaying a pre-rendered digital avatar that represents the user, maintaining communication effectiveness while avoiding poor quality video transmission.
Solution Approach 2:
The system performs preliminary actions by capturing high-quality video streams in advance and generating digital avatars before they are needed. These pre-captured video segments and generated avatars are stored and ready for immediate deployment when network conditions deteriorate, ensuring seamless transition without real-time transmission delays.
2Stability of the object's composition
If digital avatars are used to represent users, then consistent online presence is maintained, but the system complexity increases
Solution Approach 1:
The patent creates a digital copy (avatar) of the user that replicates their visual presence. This copy is generated from high-quality pre-captured video segments and can be deployed consistently across different platforms and devices, ensuring stable online presence representation without requiring complex real-time processing at each endpoint.
3Manufacturing precision
If machine learning is used to synchronize facial expressions, then avatar realism is improved, but processing time increases
Solution Approach 1:
The system performs facial expression analysis and avatar synchronization in advance during the video capture phase. Machine learning models process and analyze facial expressions from captured video segments before they are needed, so that when the avatar is deployed, the synchronization is already prepared and can be rendered immediately without real-time processing delays.
Data Source
AI summary
Transmission of an instance of a three-dimensional (3D) mesh model to a target computer device associated with a first user account is triggered. Changes in a video stream captured at a source computer device associated with a second user account, the second user account represented by a digital rendering at the target computer device according to the instance of the 3D mesh model received by the target computer device are detected. A command based on the detected changes in the video stream captured at the source computer device is identified, the at least one command corresponding to a portion of blendshapes. Transmission of the identified command to the target computer device associated with the first user account is triggered, the target computer device generating a local instantiation of the digital rendering according to the 3D mesh model and the blendshapes.


