Motion Capture Data Streams for Low Bandwidth Digital Human Sessions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Video conferencing often degrades in quality due to high bandwidth requirements for streaming high-resolution imagery, especially with multiple participants, leading to reduced audio and video quality.
Innovation Solution
Implementing low bitrate digital human communication sessions using motion capture data streams to render three-dimensional (3D) imagery of participants' movements, eliminating the need for video streaming and significantly reducing network bandwidth consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If high-resolution video imagery is streamed to maintain visual quality, then visual clarity is improved, but network bandwidth consumption increases and audio quality degrades
Solution Approach 1:
The patent extracts only the essential motion information from video streams by capturing key pose parameters, facial expressions, and gestures, while discarding redundant pixel data. This allows reconstruction of visual content from minimal data points, dramatically reducing bandwidth requirements while preserving visual quality.
Solution Approach 2:
Instead of transmitting actual video frames, the system creates and transmits simplified digital avatars that copy only the essential motion characteristics of speakers. These avatars are then rendered in real-time at the receiving end, providing visual representation without requiring high-bandwidth video streaming.
2Adaptability or versatility
If multiple participants are included in the conference, then communication versatility is improved, but the number of image streams increases and degrades overall quality
Solution Approach 1:
The system extracts only essential motion parameters from each participant rather than transmitting complete video streams. By capturing key pose, facial expression, and gesture data points, the system maintains support for multiple participants while keeping bandwidth consumption proportional to the number of participants rather than exponentially increasing.
Solution Approach 2:
The patent transforms video data from high-dimensional pixel information to low-dimensional motion parameters. This parameter transformation allows multiple participants to be included in conferences with minimal bandwidth increase, as each participant requires transmission of only a small set of motion capture data points rather than full-resolution video frames.
3Measurement precision
If motion capture data is used to render 3D imagery, then visual quality is improved, but processing complexity increases
Solution Approach 1:
The system performs preliminary processing by capturing and storing motion capture data during the conference. This pre-captured data including pose, facial expressions, and gestures is then used to drive 3D avatar rendering in real-time, reducing the computational burden during live rendering while maintaining high visual quality.
Solution Approach 2:
The patent uses pre-rendered 3D avatar models that copy the essential visual characteristics of participants. These standardized avatar models reduce processing complexity compared to generating unique high-resolution 3D models for each participant, while still providing hyper-realistic visual representation through motion capture animation.
Data Source
AI summary
A first computing device establishes a communication session with a second computing device. The first computing device receives a first motion capture data stream originating from the second computing device during the communication session, the first motion capture data stream quantifying real-time movements of a first user of the second computing device. The first computing device renders, to a display device, imagery of an animation of a three-dimensional (3D) model of the first user based on the first motion capture data stream that depicts the real-time movements of the first user.


