Avatar Expression Transfer Reducing Bandwidth in Video Calls
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Video calls face challenges with high throughput requirements and excessive information transmission, particularly over cellular networks, leading to bandwidth issues and unwanted exposure of personal or distracting backgrounds.
Innovation Solution
The system captures images of a user's face and generates an avatar, transmitting expression information instead of real-time video, allowing for real-time animation on the receiving device without sending actual video frames, thus reducing bandwidth demand and eliminating unwanted background exposure.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If real-time video frames are transmitted between electronic devices, then visual information exchange is achieved, but bandwidth requirements increase and personal background information is exposed
Solution Approach 1:
The patent extracts only the essential facial expression information from the full video feed by capturing facial landmark indicators and animation parameters, transmitting only this extracted data rather than complete video frames. This reduces bandwidth consumption while maintaining the ability to convey visual emotional information through animated avatars.
Solution Approach 2:
The patent creates simplified digital copies of facial expressions using avatar models and animation parameters instead of transmitting actual video frames. The avatar serves as a representative copy that conveys the essential visual information (facial expressions) without requiring the full bandwidth of original video transmission.
2Loss of information
If real-time video frames are transmitted between electronic devices, then visual information exchange is achieved, but transmission data volume increases
Solution Approach 1:
The patent segments the video information transmission into distinct components: facial landmark indicators capturing key facial feature positions and animation parameters describing expression movements. This segmentation allows transmission of only the necessary data elements rather than complete video frames, significantly reducing data volume while preserving facial expression information.
Solution Approach 2:
The patent extracts essential facial expression data from the video stream by identifying and transmitting only facial landmark indicators and animation parameters. This extraction process filters out redundant information (background, non-facial elements) and transmits only the critical data needed for visual expression communication.
3Loss of information
If real-time video is captured and transmitted, then facial expression communication is achieved, but unwanted background information is exposed
Solution Approach 1:
The patent extracts only the facial region information from the complete video feed by capturing facial landmark indicators that specifically map to facial features. This extraction isolates the desired facial expression data while automatically excluding unwanted background information, personal space details, and other non-essential elements from transmission.
Solution Approach 2:
The patent applies local quality analysis by focusing computational and transmission resources specifically on facial regions rather than the entire video frame. Facial landmark indicators are calculated only for facial features, giving high-quality representation to the important area (face) while ignoring or minimizing data from less important areas (background), thus preventing unwanted background exposure.
Data Source
AI summary
Methods, devices, and systems for expression transfer are disclosed. The disclosure includes capturing a first image of a face of a person. The disclosure includes generating an avatar based on the first image of the face of the person, with the avatar approximating the first image of the face of the person. The disclosure includes transmitting the avatar to a destination device. The disclosure includes capturing a second image of the face of the person on a source device. The disclosure includes calculating expression information based on the second image of the face of the person, with the expression information approximating an expression on the face of the person as captured in the second image. The disclosure includes transmitting the expression information from the source device to the destination device. The disclosure includes animating the avatar on a display component of the destination device using the expression information.


