Parameter-Based Avatar Generation for Video Chat Traffic Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing communication traffic due to transmission of face images and voices in video chat systems can lead to a significant network communication load, resulting in financial burdens for users, especially in charging systems that rely on data usage.
Innovation Solution
An information processing apparatus that generates parameter information representing a user's state, which is then transmitted over the network, allowing the communication partner's apparatus to generate an image reflecting the user's state, thereby reducing the need for transmitting high-bandwidth data like face images and voices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If face images and voices are transmitted in video chat systems, then communication quality is improved, but communication traffic increases significantly
Solution Approach 1:
The patent extracts only the essential parameters from face images and voice data that are necessary to convey user state information. Instead of transmitting complete multimedia data, the system identifies and transmits only key parameters (such as facial expression parameters, posture parameters, or voice tone parameters) that capture the essential communication content, thereby significantly reducing communication traffic while maintaining communication quality.
Solution Approach 2:
The patent transforms complex multimedia data (face images and voices) into simplified parameter representations. By converting rich media data into compact parameter sets that encode user state information, the system achieves efficient data transmission. The parameter information can be reconstructed or visualized at the receiving end to restore the communication quality without requiring transmission of the original high-bandwidth media files.
2Quantity of substance
If parameter information is generated and transmitted instead of face images, then communication traffic is suppressed, but communication quality may be degraded
Solution Approach 1:
The patent creates a simplified copy or representation of the original face image and voice data in the form of parameter information. This parameter copy contains the essential characteristics needed for communication purposes. At the receiving end, these parameters can be used to reconstruct visual or auditory representations, or directly drive avatar animations, thereby maintaining communication quality while using only the compact parameter data for transmission.
Solution Approach 2:
The patent segments the comprehensive face image and voice data into distinct parameter components (such as facial expression parameters, head pose parameters, or voice intensity parameters). This segmentation allows the system to transmit only the relevant parameters needed for effective communication, reducing overall traffic while preserving the essential communication quality through selective parameter transmission.
Data Source
AI summary
An information processing apparatus according to an embodiment of the present technology includes a generation unit and a first transmission unit. The generation unit generates parameter information that shows states of a user. The first transmission unit transmits the generated parameter information through a network to an information processing apparatus of a communication partner capable of generating an image that reflects the state of the user on the basis of the parameter information.


