3D Avatar Rendering via Markerless Arm Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video conference systems lack immersive and interactive 3D environments, failing to accurately represent participants' gaze directions and body movements, which limits the sense of real-time interaction and engagement among participants.

Innovation Solution

A system and method for generating and updating 3D participant representations within a virtual 3D video conference environment, using direction of gaze information to adjust the avatar's view and incorporating neural networks for real-time rendering and synchronization of audio and video, allowing for dynamic and immersive interactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If 2D video conference representations are used, then device complexity is reduced, but immersion and interaction quality deteriorate

Engineering Contradiction:
Improvesystem complexityVSAvoidimmersion quality
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent transitions from 2D video representations to 3D immersive environments by introducing depth perception, spatial positioning, and volumetric rendering. Participants are represented as 3D avatars or holograms within a virtual 3D space, allowing for natural gaze direction, body orientation, and spatial relationships that enhance immersion while maintaining manageable system complexity through modular architecture.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If detailed gaze direction tracking is implemented, then interaction accuracy is improved, but measurement and detection difficulty increases

Engineering Contradiction:
Improvegaze direction accuracyVSAvoiddetection complexity
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The patent integrates multiple detection functions into unified sensors and algorithms that simultaneously track gaze direction, head pose, and body orientation. The gaze tracking system is combined with overall motion capture and spatial positioning, allowing a single integrated system to perform multiple measurement tasks efficiently, thereby reducing overall detection complexity while maintaining high precision.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If real-time 3D rendering is performed, then interaction quality is improved, but processing speed requirements increase

Engineering Contradiction:
Improverendering speedVSAvoidinteraction quality
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent employs pre-rendered 3D models, pre-computed lighting conditions, and pre-established virtual environments that can be quickly instantiated and populated with participant avatars. Rather than rendering everything in real-time, the system prepares 3D assets, textures, and environmental elements in advance, then assembles them dynamically during the conference, achieving high-quality real-time rendering without excessive processing demands.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12170698B2Arm movement mimicking
Publication Date: 2024.12.17 CAVENDISH CAPITAL LLC
  • US12170698B2 patent drawing
  • US12170698B2 patent drawing
  • US12170698B2 patent drawing

AI summary

A method for virtually mimicking an arm movement of participant of a three dimensional (3D) video conference, the method may include (i) obtaining joints movement information about joints movements of an arm of a participant of the 3D video conference, based on a movement of the arm of the participant that was captured by a video taken by a camera of the participant, wherein arm of the participant is free of physical joint movement markers; (ii) generating, an arm skeletal model that represents the joins movement of the arm, based on the joints movement information; (iii) generating, by a machine learning process, a renderable 3D model of the arm that mimics the movement of the arm of the participant; and (iv) responding to the generating of the renderable 3D model, wherein the responding comprises at least one out of (a) rendering an avatar of the participant based on the renderable 3D model, (b) storing renderable 3D model information, or (c) sending the renderable 3D model information to another computerized device related to the 3D video conference.