3D Avatar Rendering for Video Conferencing Bandwidth Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video conferencing technologies are resource-intensive, costly, and can provide an unflattering and unsatisfying experience due to issues like poor lighting, shaky footage, and unflattering camera angles, as well as requiring significant processing power and bandwidth.

Innovation Solution

The system captures image and audio data of a user to generate a 2D or 3D model, rendering audiovisual information to simulate the video conferencing experience, allowing for photorealistic or avatar representations, adaptive voice synthesis, and environmental modifications, reducing data streaming requirements and improving user experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional video conferencing streams image and audio data in real time, then audiovisual communication is enabled, but network bandwidth consumption and processing power requirements increase significantly

Engineering Contradiction:
Improveaudiovisual communication capabilityVSAvoidnetwork bandwidth consumption
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent creates a 3D digital copy (avatar) of the user that can be rendered and transmitted instead of streaming actual video data. The system captures user features and generates a computational model that replicates the user's appearance and movements, allowing the avatar to represent the user in virtual meetings with minimal data transmission requirements

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary action by capturing and storing user feature data in advance to build a 3D avatar model. This pre-captured data includes facial geometry, skin texture, hair characteristics, and body measurements, which are processed beforehand to create a reusable digital twin that can be rapidly rendered during communication sessions

Inventive Principle:
Principle #10Preliminary action

2Reliability

If conventional video conferencing captures and streams video in real time, then live communication is enabled, but processing power and memory consumption increase significantly

Engineering Contradiction:
Improvelive communication capabilityVSAvoidprocessing power consumption
Core Design Contradiction:
ReliabilityVSPower

Solution Approach 1:

Instead of processing and streaming actual video frames, the system uses a pre-built 3D avatar model that can be rendered efficiently. The avatar serves as a computational copy that requires far less processing power to generate and transmit compared to real-time video encoding and decoding operations

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical video processing system (capture, encode, transmit, decode, display) with a computational rendering system. Instead of manipulating actual video data through complex encoding/decoding pipelines, the system uses mathematical models and algorithms to generate visual representations from stored 3D data

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Ease of operation

If conventional video conferencing holds the device at half arm's length below chest level, then the camera can capture the user's face, but the perspective becomes unflattering

Engineering Contradiction:
Improvedevice positioning simplicityVSAvoidvisual quality
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent transitions from 2D camera capture to 3D digital modeling. By creating a three-dimensional avatar that captures the user's true proportions and features from multiple angles during scanning, the system eliminates the perspective distortion inherent in single-angle 2D video capture, providing flattering representations regardless of device positioning

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Reliability

If conventional video conferencing relies on real-time video capture, then live interaction is enabled, but user errors and environmental conditions cause unsatisfying experiences

Engineering Contradiction:
Improvelive interaction capabilityVSAvoidlighting conditions and device stability
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The system performs comprehensive user scanning and data capture in advance under controlled conditions, building a robust 3D avatar model before actual communication sessions. This preliminary data collection includes multiple angles and lighting conditions, creating a stable digital representation that is immune to environmental variations during live interactions

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The avatar serves as a stable digital copy that replicates the user's appearance and characteristics consistently across different environments. Unlike real-time video capture that is subject to lighting changes, camera shake, and angle variations, the rendered avatar provides a consistent, high-quality representation regardless of physical conditions during communication

Inventive Principle:
Principle #26Copying

Data Source

PatentUS9479736B1Rendered audiovisual communication
Publication Date: 2016.10.25 AMAZON TECH INC
  • US9479736B1 patent drawing
  • US9479736B1 patent drawing
  • US9479736B1 patent drawing

AI summary

Systems and approaches are provided to allow for rendered audiovisual communication. An electronic device can be used to capture image information relating to physical features of a user. A model can be generated from the image information, and the model may be used to render audiovisual communication information from image and audio captured in real time. The rendered audiovisual communication data can simulate live video conferencing with substantial performance gains over conventional approaches to video conferencing. When the image capturing component of the electronic device is capable of depth imaging, stereo imaging, or other imaging techniques, the rendered audiovisual communication can be further enhanced with 3-D rendering of the user. Other aspects of audiovisual data, such as speech, background, and lighting conditions can also be rendered or synthesized to improve audiovisual communication.