3D Avatar Reconstruction with Depth Scanning and Voice Sync

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face challenges in capturing and transmitting high-quality, real-time 3D hologram videos of live persons over low bandwidth networks, particularly in maintaining natural movement and texture fidelity during remote visualization.

Innovation Solution

The system employs a depth sensor on a mobile device to scan and reconstruct a 3D model of a person's upper body, using non-rigid object scan technology, and synchronizes voice transfer, with processing capabilities distributed between the mobile device and a remote/cloud server for real-time visualization on remote devices, leveraging edge cloud processing and augmented reality functionalities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If full 3D hologram video data is transmitted in real-time, then visualization quality and realism are improved, but network bandwidth consumption increases

Engineering Contradiction:
Improvevisualization qualityVSAvoidnetwork bandwidth consumption
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent segments the 3D avatar data into multiple components: depth map, color image, and skeleton animation data. These segmented components are processed and transmitted separately, allowing efficient compression and selective transmission based on bandwidth availability while maintaining overall visualization quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a simplified 3D avatar model that copies only the essential visual characteristics of the original person. Instead of transmitting full-resolution hologram video, the system transmits compressed representations (depth maps and texture maps) that can be reconstructed into realistic avatars at the receiving end.

Inventive Principle:
Principle #26Copying

2Measurement precision

If high-resolution depth and color images are captured and transmitted, then texture fidelity is improved, but data transmission volume increases

Engineering Contradiction:
Improvetexture fidelityVSAvoiddata transmission volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent captures high-resolution depth and color images to create accurate texture maps of the avatar, but then compresses these textures for transmission. The high-fidelity texture information is preserved in compressed form, allowing realistic visualization without proportionally high data transmission volumes.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms the captured images into different parameter representations (depth maps, normal maps, texture coordinates) that encode the same visual information more efficiently. This parameter transformation allows high texture fidelity to be achieved with reduced data transmission requirements.

Inventive Principle:
Principle #35Parameter changes

3Loss of time

If real-time processing is performed on mobile device, then latency is reduced, but computational power requirements increase

Engineering Contradiction:
Improveprocessing latencyVSAvoidcomputational power
Core Design Contradiction:
Loss of timeVSPower

Solution Approach 1:

The patent divides processing tasks between the mobile device and remote servers. The mobile device performs lightweight real-time tasks such as capturing depth and color images, extracting skeleton data, and initializing avatar parameters. More computationally intensive tasks like full 3D model reconstruction and rendering are performed remotely, reducing mobile device power consumption while maintaining low latency through efficient task distribution.

Inventive Principle:
Principle #1Segmentation

4Adaptability or versatility

If non-rigid object scan technology is used to capture natural movement, then movement naturalness is improved, but processing complexity increases

Engineering Contradiction:
Improvemovement naturalnessVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent replaces complex mechanical tracking systems with sensor-based capture methods. Depth sensors and color cameras automatically capture the subject's movements without requiring the subject to follow rigid scanning patterns. This substitution of mechanical scanning with optical sensing simplifies the system while enabling natural movement capture.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms complex 3D movement data into simplified skeleton animation parameters and deformation maps. By changing the representation parameters from full volumetric data to key pose and deformation information, the system maintains movement naturalness while reducing processing complexity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11170552B2Remote visualization of three-dimensional (3D) animation with synchronized voice in real-time
Publication Date: 2021.11.09 SAMSUNG ELECTRONICS CO LTD
  • US11170552B2 patent drawing
  • US11170552B2 patent drawing
  • US11170552B2 patent drawing

AI summary

Described herein are methods and systems for remote visualization of three-dimensional (3D) animation. A sensor of a mobile device captures scans of non-rigid objects in a scene, each scan comprising a depth map and a color image. A server receives a first set of scans from the mobile device and reconstructs an initial model of the non-rigid objects using the first set of scans. The server receives a second set of scans. For each scan in the second set of one or more scans, the server determines an initial alignment between the depth map and the initial model. The server converts the depth map into a coordinate system of the initial model, and determines a displacement between the depth map and the initial model. The server deforms the initial model to the depth map using the displacement, and applies a texture to at least a portion of the deformed model.