3D Avatar Communication via Keyframe Metadata Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current 3D avatar conversation systems require extensive hardware for capturing and transmit large amounts of data, leading to high bandwidth requirements and complexity in managing multiple callers, with limited dynamic adaptation of animations for receivers.
Innovation Solution
A metadata structure and format that enables low-bandwidth 3D avatar communication by transmitting only keypoints and joints from the sender, allowing the receiver to animate the avatar, with dynamic metadata streamed for re-enacting the animation, and supporting various standards like ISO/IEC 23090-8 NBMP for cloud-based media processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If large amounts of 3D point-cloud data are transmitted for real-time XR conversation, then the quality and realism of the conversational experience is improved, but the bandwidth requirements and data transmission complexity increase significantly
Solution Approach 1:
The patent extracts only the essential animation parameters from the complete 3D point-cloud data. Instead of transmitting full point-cloud representations, the system identifies and transmits only the key animation parameters (such as joint positions, rotation angles, and transformation matrices) that are necessary to animate the receiver's local 3D model, thereby dramatically reducing data transmission volume while preserving animation quality
Solution Approach 2:
The patent inverts the traditional approach by having the receiver generate and maintain a local copy of the 3D model, rather than the sender transmitting complete model data. The sender only transmits animation parameters that drive the receiver's local model, reversing the data flow direction and reducing bandwidth requirements
2Measurement precision
If extensive hardware and processing power are used for capturing and transmitting 3D avatar data, then the realism and quality of the conversational service is improved, but the device complexity and cost increase
Solution Approach 1:
The patent uses copying by having the receiver create and maintain a local copy of the 3D avatar model. Instead of requiring the sender to transmit complete high-fidelity model data continuously, the receiver maintains a local replica that can be animated with transmitted parameters, reducing the hardware and bandwidth requirements at the sender side while maintaining quality at the receiver side
3Adaptability or versatility
If complete 3D point-cloud animation data is transmitted to support multiple callers, then the versatility and adaptability of the conversational service is improved, but the bandwidth consumption and system complexity increase
Solution Approach 1:
The patent applies segmentation by dividing the animation data into separate, independent parameter sets for different callers. Each caller's animation parameters (position, orientation, joint transformations) are segmented and transmitted independently, allowing the receiver to efficiently manage multiple avatar instances without transmitting redundant complete model data for each caller
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The embodiments relate to a method comprising establishing a three-dimensional conversational service between a sender and a receiver based on point-cloud and indication on capabilities of supporting three-dimensional point-cloud animation, said indication and said point-cloud having been received from said sender; and receiving media relating to the conversational service and corresponding metadata from the sender, the metadata comprising parameters having dynamically created timed low-level information for re-enacting the animated three-dimensional point-cloud animation on the receiver alongside with the media. In addition, the embodiments relate to an apparatus for implementing the method.