3D Face Representation Using Multi-Surface Depth Values
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques fail to accurately represent a user's current appearance in real-time, often using outdated images to generate avatars that do not reflect the user's current facial expressions or changes, such as smiling or beard growth.
Innovation Solution
A method using depth values defined relative to multiple points on a non-planar surface, such as a cylindrical shape, to generate a 3D representation of a user's face, which requires less computation and bandwidth compared to 3D meshes or 3D point clouds, and can be formatted like RGBDA images for efficient integration.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If 3D mesh or 3D point cloud is used to represent user's face, then representation accuracy is improved, but computational complexity and bandwidth requirements increase
Solution Approach 1:
The patent extracts only the essential depth information needed for accurate representation, discarding redundant data structures. Instead of using complete 3D meshes or point clouds, the invention extracts depth values relative to multiple surface points, achieving accurate representation with significantly reduced data complexity and computational requirements.
Solution Approach 2:
The patent applies different representation qualities to different regions of the face by using multiple surface points (e.g., cylindrical surface points) with varying density and precision. This allows higher accuracy where needed while reducing overall data complexity, balancing representation accuracy with computational efficiency.
2Measurement precision
If 3D mesh or 3D point cloud is used to represent user's face, then representation accuracy is improved, but bandwidth requirements increase
Solution Approach 1:
The patent extracts only the necessary depth values relative to multiple surface points, removing unnecessary data from complete 3D mesh or point cloud representations. This extraction achieves accurate user representation while significantly reducing the quantity of data that needs to be transmitted, thereby lowering bandwidth requirements.
3Quantity of substance
If RGBDA images are used to represent user's face, then bandwidth is reduced, but representation accuracy deteriorates
Solution Approach 1:
The patent transitions from single-camera-depth (RGBDA) representation to multi-point surface depth representation. By defining depth values relative to multiple points on a surface (e.g., cylindrical surface) rather than a single camera location, the invention adds dimensional information that improves representation accuracy while maintaining bandwidth efficiency through compact data formatting.
Solution Approach 2:
The patent enhances local representation quality by using multiple surface points with different positions and orientations. This allows capturing fine facial details and expressions more accurately than single-point depth methods, improving overall representation accuracy without proportionally increasing bandwidth consumption.
4Measurement precision
If multiple surface points are used to define depth values, then representation accuracy is improved, but data processing complexity increases
Solution Approach 1:
The patent segments the face representation into multiple independent depth measurements relative to different surface points. This segmentation allows parallel processing of depth values and simplifies the overall computation by breaking down the complex 3D reconstruction task into manageable independent measurements that can be processed efficiently.
Data Source
Figure 1
Figure 2
Figure 3A
AI summary
Various implementations disclosed herein include devices, systems, and methods that generates values for a representation of a face of a user. For example, an example process may include obtaining sensor data (e.g., live data) of a user, wherein the sensor data is associated with a point in time, generating a set of values representing the user based on the sensor data, and providing the set of values, where a depiction of the user at the point in time is displayed based on the set of values. In some implementations, the set of values includes depth values that define three-dimensional (3D) positions of portions of the user relative to multiple 3D positions of points of a projected surface and appearance values (e.g., color, texture, opacity, etc.) that define appearances of the portions of the user.