Virtual Character Facial Expression Parameterization for Bandwidth Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video chatting technologies face challenges in maintaining real-time high-quality image and sound transmission without latency, especially in environments with limited communication bandwidth or when processing resources are strained.
Innovation Solution
An information processing device and method that uses a virtual character represented by computer graphics, analyzing facial expressions and voice data to generate animated images, synchronizing these with voice data for real-time output, reducing the processing load and bandwidth requirements by using pre-defined facial expression models and weights.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If high-quality image and sound data are transmitted without compression, then conversation quality is improved, but communication bandwidth requirements increase and latency occurs
Solution Approach 1:
The patent extracts only the essential facial expression information from the original video feed by tracking specific facial landmarks (eyes, eyebrows, mouth) and representing them through a limited set of expression parameters. This allows the system to convey emotional and expressive information without transmitting the full high-resolution video stream, thereby reducing bandwidth requirements while maintaining conversation quality.
Solution Approach 2:
The patent transforms the continuous video data into discrete facial expression parameters by defining specific landmarks and expression types (e.g., smiling, frowning, surprised). This parameterization approach converts high-volume video data into low-volume symbolic representations that can be transmitted efficiently while preserving the essential expressive content needed for natural conversation.
2Reliability
If real-time video processing is performed to maintain natural conversation, then conversation naturalness is improved, but processing resources are strained
Solution Approach 1:
The patent segments the face into distinct anatomical regions (eyes, eyebrows, mouth) and tracks them independently using landmark detection. This segmentation allows the system to process only relevant facial regions rather than analyzing the entire video frame, significantly reducing computational complexity while maintaining the ability to detect natural facial expressions for real-time conversation.
Solution Approach 2:
The patent implements partial action by focusing computational resources only on detecting and tracking essential facial landmarks rather than performing full video analysis. By selectively processing only the critical facial regions and expression parameters needed for natural conversation, the system achieves real-time performance with reduced processing overhead.
3Measurement precision
If facial expressions are analyzed in detail to enhance entertainment experience, then expression accuracy is improved, but processing time increases
Solution Approach 1:
The patent performs preliminary action by pre-defining facial landmarks, expression types, and comparison criteria before actual expression analysis begins. The system establishes the framework for measurement (landmark coordinates, expression categories, threshold values) in advance, allowing rapid real-time classification of facial expressions without requiring complex computations during the actual conversation, thus achieving both accuracy and speed.
Data Source
AI summary
The basic image specifying unit specifies the basic image of a character representing a user of the information processing device. The facial expression parameter generating unit converts the degree of the facial expression of the user to a numerical value. The model control unit determines an output model of the character for respective points of time. The moving image parameter generating unit generates a moving image parameter for generating animated moving image frames of the character for respective points of time. The command specifying unit specifies a command corresponding to the pattern of the facial expression of the user. The playback unit outputs an image based on the moving image parameter and the voice data received from the information processing device of the other user. The command executing unit executes a command based on the identification information of the command.


