Depth Map Generation for 3D Avatar Creation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Videoconferencing systems face challenges with large data requirements for video streams, particularly when using wireless networks, which can lead to difficulties in data transmission and hardware needs for generating 3D representations of users.
Innovation Solution
A computing device generates a depth map based on a video stream using a depth prediction model, allowing for the creation of a 3D avatar representation of a local user with head movement, eye movement, and facial expressions, using a single color camera, reducing the need for multiple cameras and data transmission.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single color camera is used to generate 3D representations, then hardware requirements are reduced, but the ability to capture accurate depth information deteriorates
Solution Approach 1:
The patent introduces depth prediction models and neural networks as intermediary computational tools that process the 2D video stream from a single color camera to generate depth maps. These intermediaries bridge the gap between the limited capability of a single color camera and the need for accurate depth information, allowing the system to infer depth data that would otherwise require multiple cameras or specialized depth-sensing hardware.
Solution Approach 2:
The patent replaces the mechanical/optical system of multiple cameras or depth-sensing cameras with a computational system using a single color camera combined with AI algorithms. Instead of using multiple physical sensors to capture depth information directly, the system substitutes this mechanical approach with computational image processing and neural network-based depth prediction, thereby reducing hardware complexity while maintaining functional capability.
2Reliability
If full video streams are transmitted for videoconferencing, then communication quality is maintained, but data transmission requirements increase
Solution Approach 1:
The patent extracts only the essential visual information needed for videoconferencing by generating compressed representations such as depth maps, facial landmarks, and key motion data from the full video stream. Instead of transmitting the complete high-resolution video data, the system extracts and transmits only the critical elements required for maintaining communication quality, significantly reducing data transmission volume while preserving the essential visual communication function.
Data Source
AI summary
A method can include receiving, via a camera, a first video stream of a face of a user; determining a location of the face of the user based on the first video stream and a facial landmark detection model; receiving, via the camera, a second video stream of the face of the user; generating a depth map based on the second video stream, the location of the face of the user, and a depth prediction model; and generating a representation of the user based on the depth map and the second video stream.


